Commit Graph

216 Commits

Author SHA1 Message Date
pr-eval f5bd6bbe0e fix(agent): count malformed yields after incremental sections
Narrow the invalid-yield guard to !abortSent so array-typed incremental
yield sections no longer suppress the infinite-submit-loop abort; add
regression coverage for incremental yield followed by repeated malformed
terminal yields.
2026-07-20 22:51:53 +02:00
can1357 a87d7d35cd Merge PR #4961: fix(agent): stop malformed subagent yield loops (@roboomp)
# Conflicts:
#	packages/coding-agent/src/prompts/system/workflow-notice.md
#	packages/coding-agent/src/task/executor.ts
#	packages/coding-agent/src/tools/yield.ts
2026-07-20 22:51:53 +02:00
can1357 ad9d272efe Merge PR #5757: fix(task): inherit default fallback for single-model subagents (@jeffscottward) 2026-07-18 20:12:48 +02:00
metaphorics 8d8ad0faed perf(coding-agent): coalesce subagent output reconstruction
Defer recentOutput line reconstruction from every text_delta to the
progress emit boundary. appendRecentOutputTail only extends the capped
raw tail and marks dirty; refreshRecentOutput runs the exact old
split/filter/slice(-8)/reverse algorithm as the first step of every
emitProgressNow snapshot (onProgress + event bus), including coalesced
and finalize/error/cancel flushes. Reset publishes [] immediately;
replace marks dirty; past snapshot arrays stay immutable via spread.

Before (base pool median-of-5):
  w8_d3  61.55 cpu_ms/1k_events
  w32_d3 44.16 cpu_ms/1k_events
After (stable final run on E+G, 7 episodes, trimmed CV gate pass):
  w8_d3  55.78 cpu_ms/1k_events  (1.103×)  trimmed CV 15.1%
  w32_d3 40.72 cpu_ms/1k_events  (1.084×)  trimmed CV 11.1%
Checksums match prior exactness baseline; retained_after_release_kb
1284 / 2864 (no regression vs prior concur).

Op: GConcurEmitBoundary emit-boundary dirty flag
Restores: none
2026-07-18 11:07:51 +09:00
Jeff Scott Ward 6dcc485389 fix(task): inherit default subagent fallback 2026-07-17 17:43:04 -04:00
vmcall d53cf023b0 fix(task): reconciled structured subagents with upstream
- Preserved the plan-mode capability clamp after upstream removed report_finding.
- Updated persisted-revival coverage for mounted xdev tool activation.
- Applied current formatter output to conflicted runtime files.
2026-07-17 17:38:12 +02:00
vmcall d944879f21 feat(task): unified structured subagent execution
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.

Fixes #5279
2026-07-17 17:36:59 +02:00
can1357 1074332098 fix(task): abort background label generation 2026-07-17 04:15:40 +02:00
can1357 9afedb591e feat(coding-agent): added opt-in task prewalk and tightened --tools and xdev behavior
- Added a `task.prewalk` option (default `false`), removed default task `prewalk` flags, and updated prewalk resolution so bunded generic task execution only prewalks when explicitly enabled.
- Enforced strict `--tools` validation in CLI parsing, making unknown tool names fail fast with `CliUsageError` instead of being silently filtered.
- Migrated legacy discovery settings (`tools.discoveryMode`, `tools.essentialOverride`, MCP discovery keys) into updated `tools.xdev` handling with preserved explicit override behavior.
- Hardened xdev/ACP execution flow by capping `docsAll` payloads with overflow listing and remapping `xd://` dispatches/approval gating for correct execute/read behavior and reduced duplicate prompts.
2026-07-15 18:39:36 +02:00
can1357 5ff277349c refactor(coding-agent): consolidated tool surface onto xd:// devices and hub
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
2026-07-15 15:16:29 +02:00
can1357 c32ea55ba3 fix(coding-agent): kept todo active for prewalk subagents and fixed prewalk gate deadlock
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
2026-07-15 05:45:59 +02:00
can1357 425e583ae0 feat(coding-agent): added support for task-agent field and model resolution
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
2026-07-15 00:50:55 +02:00
djdembeck aa470b9e34 fix(coding-agent): forward sessionId to getApiKey in subagent auth fallback
The pre-flight auth check in resolveModelOverrideWithAuthFallback called
getApiKey without a session id. For providers with session-sticky OAuth
credentials, this returned undefined even though the credential was
usable once the subagent session started, causing the auth fallback to
silently replace the configured model with the parent's (#5325).

The subagent's id is now forwarded as the session id so session-sticky
credentials resolve during the pre-flight check. Genuinely broken auth
(stale OAuth, revoked tokens) still falls back as before.

Also propagate model resolution warnings through resolveModelOverride
and log them in the executor so users see why a pattern didn't match.
2026-07-14 00:28:42 -05:00
can1357 33b6774aa1 feat(coding-agent): implemented resumable subagent yielding for tasks
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
2026-07-11 17:27:35 +02:00
can1357 cb2153e9a5 feat(coding-agent-task): implemented agent-centric flat task structure
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
2026-07-11 13:44:04 +02:00
can1357 441037025a feat(coding-agent): updated thinking-level configuration and precedence
- Configured the default `task` subagent to use `auto` thinking.
- Enabled `auto` as a valid thinking-level value in agent frontmatter.
- Adjusted thinking-level precedence to ensure that explicit `:level` suffixes in resolved model patterns override agent-defined defaults.
2026-07-11 13:35:05 +02:00
can1357 75bac085a7 feat(vibe): implemented persistent worker session infrastructure
- Introduced a registry and runtime for managing persistent worker subagent sessions.
- Added lifecycle management capabilities including spawning, dispatching, waiting, and terminating background jobs.
- Implemented TUI visualization tools to track and render worker session states.
- Enabled subagent session continuation via follow-up turn processing.
2026-07-11 05:47:55 +02:00
can1357 6dbbfbe1e0 feat(coding-agent): renamed explore agent to scout
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
2026-07-10 12:51:50 +02:00
can1357 2bf8f43b67 fix(agent): keep incremental yields budget-bound 2026-07-10 12:37:46 +02:00
roboomp 4f689c0b1b fix(agent): validated pending yield commits
Keep assistant yield tool calls pending until YieldTool.execute returns a successful tool result.

Prevent invalid pre-execution yield arguments from bypassing schema retry handling while still suppressing the soft budget abort during validation.

Fixes #5006
2026-07-10 01:47:08 +00:00
roboomp 1bb97efec8 fix(agent): committed yield before budget abort
Persist yield tool-call arguments as soon as an assistant turn commits the yield call, before the soft request budget guard can abort the session.

Add a regression covering a yielding turn that crosses the budget threshold without a tool result event.

Fixes #5006
2026-07-10 01:16:06 +00:00
roboomp f89194d033 fix(agent): ignored malformed yield siblings after success
- Stopped the malformed-yield loop guard once a valid yield has been captured or termination is pending.
- Covered same-turn valid yield calls followed by malformed sibling yield calls.

Fixes #4957
2026-07-09 18:26:01 +00:00
roboomp 65d3881dc9 fix(agent): stopped malformed subagent yield loops
- Failed subagents after repeated invalid yield submissions instead of waiting forever.
- Added regression coverage for malformed yield submit loops.

Fixes #4957
2026-07-09 18:15:56 +00:00
can1357 336e1529dc Merge PR #4271: fix(task): abort underlying MCP call when proxy tool timeout fires (@metaphorics) 2026-07-05 13:39:09 +02:00
roboomp 4409e6cfb5 fix(task): preserved deferred fallback chains
Install subagent retry fallback chains after deferred model patterns resolve so runtime-only candidates keep their ordered fallbacks.\n\nFixes #4421
2026-07-03 09:41:25 +00:00
roboomp 17163a2b9f fix(task): preserved auth fallback for deferred models
Preserve the subagent parent-model auth fallback when an explicit selector resolves only after child runtime provider loading.\n\nFixes #4421
2026-07-03 09:29:11 +00:00
roboomp ddd6ae1182 fix(task): preserved deferred subagent model selectors
Forward unresolved explicit subagent model selectors into child session startup so modelRoles.task cannot disappear during executor preflight and fall through to an unrelated provider default.\n\nFixes #4421
2026-07-03 09:16:30 +00:00
can1357 3830ad353e merge PR #3843 (surviving delta): perf: streaming-reveal/render throughput + core hot-path optimizations (@oldschoola)
# Conflicts:
#	packages/coding-agent/src/config/model-resolver.ts
2026-07-02 10:31:45 +02:00
metaphorics ae21a7a7e7 fix(task): abort underlying MCP call when proxy tool timeout fires
`withAbortTimeout` rejected after `MCP_CALL_TIMEOUT_MS` but left `waitForConnection` / `callTool` running because the timeout never propagated an abort signal to those underlying promises. The fix creates a per-call `AbortController` in `createMCPProxyTools.execute`, combines its signal with the caller signal via `AbortSignal.any`, and passes both the combined signal and the controller to `withAbortTimeout`, which aborts the controller on timeout or caller abort. Verified by `bun build packages/coding-agent/src/task/executor.ts --target=bun --format=esm` and `bun test packages/coding-agent/test/task-executor-mcp-timeout.test.ts`.

Closes #4242
2026-07-02 17:29:26 +09:00
can1357 b4cb78304c feat(coding-agent): introduced configurable soft request budget steering notices
- Introduced the `task.softRequestBudgetNotice` boolean setting to opt into budget steering notices.
- Disabled the wrap-up steering notice by default when a subagent crosses its soft request budget.
- Maintained the 1.5x graceful abort safety guard regardless of whether the steering notice option is enabled.
- Updated the settings schema to document the conditional steering notice behavior.
2026-07-02 02:40:07 +02:00
can1357 63b812320d Merge PR #3903: fix(coding-agent): identify user-invoked skills and expose skill directory (@metaphorics)
# Conflicts:
#	packages/coding-agent/src/extensibility/skills.ts
#	packages/coding-agent/src/modes/acp/acp-agent.ts
#	packages/coding-agent/src/modes/rpc/rpc-mode.ts
#	packages/coding-agent/src/modes/skill-command.ts
2026-07-01 22:16:53 +02:00
can1357 88e3e77f3e Merge PR #3927: fix(agent): reject stale yield labels for override schemas (@roboomp) 2026-07-01 21:50:54 +02:00
roboomp 00ef58f843 fix(agent): steered override-schema subagents
- Marked eval agent schema calls as caller overrides so subagent prompts can revoke native output/yield instructions.\n- Added override-schema prompt guidance telling agents to ignore conflicting native output labels and terminal-yield the caller schema object.\n- Added prompt coverage for the override notice.\n\nRefs #3926
2026-06-30 22:30:15 +00:00
roboomp e4561d64fa fix(coding-agent): restored subagent thinking precedence
Agent frontmatter thinkingLevel now wins over model role suffix thinking when both are configured.

Fixes #3915
2026-06-30 20:15:04 +00:00
can1357 9ccd83a13d feat(coding-agent): made the agent parameter optional with a default value
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
2026-06-30 16:16:39 +02:00
metaphorics fa64097992 fix(coding-agent): identify user-invoked skills and expose skill directory
User-invoked skills (typed /skill:, steered, follow-up, interrupted/
resumed via compaction, ACP, RPC) only appended a bare "Skill: <path>"
line, so the model neither learned the user had invoked that specific
skill nor where the skill directory was. Relative paths in skill bodies
(scripts/, templates/) could not be resolved.

Route all user-invoked paths through a self-identifying, baseDir-aware
prompt template; keep hidden autoload skills on the minimal non-user
format. Interactive skillCommands now carries the loaded Skill object
instead of a bare path so baseDir flows through without reconstruction.
The invocation kind defaults to "user" to keep buildSkillPromptMessage
source-compatible.

Op: correct
Restores: spec:user-invoked-skill-prompt-self-identifies-and-exposes-skill-directory
2026-06-30 23:11:05 +09:00
oldschoola 5abde300a1 Merge remote-tracking branch 'origin/main' into perf/streaming-reveal-throughput 2026-06-29 22:08:03 -07:00
oldschoola a1972d6345 perf: cut hot-path quadratic/allocation costs across 5 subsystems
Five hot-path performance fixes + two low-risk allocation reductions.
No behavior change; all derived counts/orderings are identical.

- session-manager pathTo: leaf->root walk used branch.unshift() per node
  (O(n^2) over branch length); now push + single reverse. Backs
  getBranch(), hit at ~17 sites per turn.

- edit/modes/patch: collapseConsecutiveSharedLines / collapseRepeatedBlocks
  / trimCommonContext built shared-line sets via
  new Set(oldLines.filter(l => newLines.includes(l))) -> O(old*new) per
  hunk. Precompute new Set(newLines) and use .has() -> O(old+new).

- task/executor appendRecentOutputTail: re-split + filter + slice + reverse
  of the full (up to 8KB) recentOutputTail on every text_delta token. Fast
  path extends the current last line in place; full recompute only when a
  newline boundary or truncation changes the window. tailLastLineRepresentable
  flag guards the trailing-whitespace-only-line edge case.

- task/render renderResult: header booleans (3x .some) + footer counts
  (3x .filter) + request total (.reduce) re-scanned details.results ~30x/sec
  via the spinner. Single pass derives aborted/failed/mergeFailed/success
  counts + requestTotal; booleans derived from counts.

- task/render extractIncrementalReviewResult: re-called normalizeYieldData
  internally though both callers had already normalized the same yield data.
  Signature now takes pre-normalized RenderYieldItem[].

Honorable mentions (allocation reduction, no algorithmic change):
- config/model-resolver: hoist case-folded pattern out of matchModel filter
  passes; build the O(n) preference context once per role in
  resolveModelRoleValue and reuse across fallback patterns.
- tools/read countTextLines: count newlines directly instead of allocating
  via split("\n"); hashline formatter reuses the line count instead of
  recomputing.
2026-06-29 22:05:09 -07:00
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
roboomp 2042d3b114 fix(task): scoped provider concurrency cap to each LLM turn
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.

Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).

Fixes #3749
2026-06-28 22:50:59 +02:00
can1357 93f68b75a2 test(coding-agent/modes): updated controller test mocks
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
2026-06-28 16:53:16 +02:00
can1357 51a2a0342f test(coding-agent): implemented guest reconciliation and expanded testing for collaboration
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
2026-06-28 09:52:44 +02:00
can1357 2cf28f972e feat(coding-agent): added logic to assembleYieldResult to
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
2026-06-28 09:05:37 +02:00
can1357 289dd770c8 feat(coding-agent): reworked subagent yields for incremental results
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
2026-06-28 07:27:01 +02:00
can1357 2c20a2d368 Fix subagent yield abort cleanup
(cherry picked from commit e22673537f7cc95920cd4ec75d2a4b36d81d7626)
2026-06-26 23:43:03 +02:00
can1357 84ef014faa Merge PR #3407: fix(eval): dispose one-shot eval subagents after run (@korri123) 2026-06-26 23:27:40 +02:00
can1357 938489f3fd feat(coding-agent): added configurable service tier settings for subagents and advisor
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
2026-06-25 22:35:15 +02:00
can1357 ecd6f608e5 fix(agent): resize provider concurrency limiter in place instead of replacing it
The per-provider subagent limiter (providers.ollama-cloud.maxConcurrency)
created a fresh Semaphore whenever the configured limit changed, orphaning
in-flight slots on the old instance so a runtime or mixed limit value could
exceed the cap. getProviderSemaphore now always hands out one shared limiter
(Infinity when unlimited, so every run is still counted) and resizes it in
place. Semaphore.release() decrements before admitting, and the new
Semaphore.resize() raises the ceiling by admitting queued waiters while
lowering it drains in-flight holders without admitting past the new cap.

Refs #3464
2026-06-25 20:48:38 +02:00
roboomp 7e90e4d081 fix(agent): released ollama-cloud semaphore slot when waiter aborts
Semaphore.acquire now accepts an AbortSignal so a queued waiter that is cancelled (parent task abort, wall-clock budget elapsing) removes itself from the wait queue instead of being resolved by the next release. The provider semaphore in runSubprocess passes the run's abortSignal through, preventing aborted ollama-cloud subagents from permanently draining the provider concurrency budget.

Fixes #3464
2026-06-25 12:14:45 +00:00
roboomp 80862b79da fix(agent): handled ollama-cloud task backoff
Added ollama-cloud subagent concurrency limiting, role fallback-chain inheritance, and visible empty length errors for native Ollama responses.

Fixes #3464
2026-06-25 11:57:51 +00:00