Commit Graph

253 Commits

Author SHA1 Message Date
can1357 94c838faea Merge PR #7539: fix(coding-agent): complete usage-aware fallback integration (@eggpeat) 2026-08-05 01:12:02 +02:00
Brent db97103c32 fix(coding-agent): complete usage-aware fallback integration 2026-08-03 16:34:35 +00:00
metaphorics e42b95bcae fix(task): report non-isolated cleanup state (#7488) 2026-08-03 22:26:24 +09:00
metaphorics d563e25cfe fix(task): settle all late cleanup (#7488) 2026-08-03 21:39:33 +09:00
metaphorics 687f3326b2 fix(task): bound abort cleanup and quarantine late jobs 2026-08-03 20:14:25 +09:00
roboomp 57732e9dcd fix(coding-agent): detach hard-aborted refs so ensureLive can't route into a dead session
Preserving aborted refs on dispose exposed a latent invariant break: the
executor's hard-abort path (finalizeSubagentLifecycle) set status `aborted`
and disposed the session without detaching it. With the ref now retained, it
kept a dangling pointer to the disposed session, and ensureLive returns any
non-null ref.session before its revivability check — so hub focus / transcript
chat could route into a dead session.

- finalizeSubagentLifecycle: detach the session before disposing on the
  terminal hard-abort path, upholding the AgentRef invariant (session === null
  when aborted).
- release(tombstone): detach before dispose too (capture the live session
  first), same invariant.
- unregisterUnlessParked: preserve `aborted` refs only when already detached;
  an aborted ref still holding a live session is a bug and is unregistered
  rather than kept reachable.
- Regression test now asserts ensureLive rejects a tombstoned id as terminal.

Fixes #7250
2026-08-01 09:42:41 +00:00
can1357 bd96752b83 fix(rpc): serialize IRC wake finalization 2026-07-31 19:28:43 +02:00
can1357 e02510f8c9 Merge PR #7108: fix(rpc): restore frames for IRC-revived subagents (@roboomp) 2026-07-31 19:28:42 +02:00
can1357 305ad95a20 fix(task): guard deferred launch work 2026-07-31 19:16:56 +02:00
Wolfgang Schoenberger dc292d4c84 fix(task): track deferred model refresh 2026-07-30 18:34:40 -07:00
Wolfgang Schoenberger 37a4adfcea perf(task): defer subagent launch setup 2026-07-30 15:16:47 -07:00
roboomp d686375b8c fix(rpc): monitored IRC wakes on cold-revived subagents
- Extracted the shared attachIrcWakeTurnMonitor from the executor reviver closure.
- Installed it in the persisted cold-revive path, forwarding the top-level event bus.
- Covered that a resumed process's parked subagent emits wake lifecycle frames.

Fixes #7105
2026-07-30 20:08:21 +00:00
roboomp b2d18b6920 fix(rpc): restored frames for IRC-revived subagents
- Monitored autonomous IRC wake turns with the task executor lifecycle and progress channels.
- Preserved monitoring after idle-TTL parking and session revival.
- Covered RPC subscriptions for both idle and parked keep-alive agents.

Fixes #7105
2026-07-30 19:42:58 +00:00
Kyle McCleary 089a9963f8 feat(security): OMP-native security subsystem (planner handoff) 2026-07-29 18:47:51 -07:00
Nik Divjak c3011fff3c fix(task): let task.softRequestBudget lower bundled subagent budgets
The soft request budget resolved to `SOFT_REQUEST_BUDGET[agent.name] ??
configured`, so the bundled entries for scout and sonic replaced the
configured value outright. Lowering `task.softRequestBudget` to tighten
the guard therefore did nothing for exactly the two agents that spawn
most often: a scout kept its 100-request budget no matter how small the
user set the knob. Only 0 (disable) and raising the value for
non-bundled agents had any effect.

Treat both numbers as upper bounds and take the smaller one. The bundled
entries stay ceilings, so a runaway scout is still stopped at 100 by
default and existing behavior is unchanged for anyone who has not
lowered the setting; a configured 0 still disables the guard entirely.
Resolution moves into `resolveSoftRequestBudget`, which also normalizes
negative and fractional inputs, so the rule is testable without standing
up a subprocess run.

This composes with `task.maxEffort` on a separate axis: effort caps how
hard each request thinks, this caps how many requests a run may spend.

(cherry picked from commit f0db29f8f725f11390b64ca9342300c482ff5c5d)
2026-07-29 23:09:01 +02:00
can1357 8c5dc16344 fix(task): enforced per-spawn effort ceiling across retry fallbacks
- task.maxEffort only clamped the initial thinking level; a retry
  fallback candidate could clamp back up to its model floor and run a
  low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
  ModelControls (constructor, setThinkingLevel, auto classifier,
  restore) and in applyRetryFallbackCandidate; fallback candidates whose
  floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
  attribution added.
- Review follow-up for PR #6794.
2026-07-27 16:09:28 +02:00
Wolfgang Schoenberger bd1605e8f5 feat(task): add per-spawn effort ceiling 2026-07-27 03:39:55 -07:00
roboomp 583ff590f2 fix(prewalk): apply same-model effort downgrades instead of skipping
The prewalk arm/switch guard compared model identity only (modelsAreEqual /
provider+id), discarding the resolved thinkingLevel. A legal same-model target
at a cheaper effort (e.g. prewalk: "@task" resolving to the active model at a
lower level) was dropped as a no-op, so the session ran the expensive effort for
the whole run while still paying the plan/continue nudges — silently on the
session path, logger.debug only on the subagent path.

Compare (provider, id, effective thinking level) via a shared prewalkWouldBeNoop
helper. Effort-only deltas on the same model now switch; a genuine no-op emits a
user-visible notice on the session path and never arms on the subagent path.

Fixes #6659
2026-07-26 02:20:03 +00:00
can1357 db937bd149 feat(coding-agent): added coarse effort parameter to task tool
- Add `effort` (`lo`/`med`/`hi`) parameter to task spawn parameters and prompts.
- Implement `resolveTaskEffortLevel` to map coarse task effort onto model-supported thinking ranges.
- Pass effort configuration through executor options and structured subagent requests.
2026-07-24 15:51:34 +02:00
Alex TYRODE e0928070c2 feat(extensions): expose session service tiers 2026-07-23 21:49:20 +00:00
can1357 5d66eb7f2a Merge PR #5464: fix(coding-agent): persist vibe sessions across restarts (@roboomp) 2026-07-23 18:06:11 +02:00
can1357 c522eceff2 Merge PR #6119: feat: lift subagent async/auto-background limits via owner-routed delivery and quiescence (@korri123)
# Conflicts:
#	packages/coding-agent/src/task/executor.ts
2026-07-23 17:52:52 +02:00
can1357 26726fdcb9 Merge PR #6255: add dynamic multi-root workspace context (@maatheusgois-dd) 2026-07-23 17:30:33 +02:00
can1357 db3a6a1407 Merge PR #6318: fix(tui): show fallback models in Agent Hub (@roboomp) 2026-07-23 11:37:13 +02:00
roboomp 00aed97ed1 fix(tui): flagged fallback badges for observer-only hub rows
- Threaded a resolvedModelIsFallback flag through AgentProgress and SingleResult.
- Set the flag from the executor retry-fallback handlers and settled results.
- Rendered the observer/no-session hub path as fallback -> provider/model.
- Added an observer-only fallback-badge regression test.

Fixes #6316
2026-07-22 19:19:08 +00:00
can1357 9952adbc82 Merge PR #6242: fix(mcp): route task proxies through source tool and mark tools non-strict (@roboomp) 2026-07-22 21:13:21 +02:00
can1357 5477e74bb3 Merge PR #6307: fix(coding-agent): align bundled task prewalk display (@roboomp) 2026-07-22 21:13:18 +02:00
can1357 a1bdf4a352 Merge PR #6214: fix(tui): refresh prewalked subagent model (@roboomp) 2026-07-22 21:13:18 +02:00
roboomp e2dc801d44 fix(coding-agent): aligned task prewalk dashboard state
Shared the bundled task prewalk default between runtime execution and the Agent Control Center. Added dashboard regression coverage for task.prewalk.

Fixes #6306
2026-07-22 17:14:36 +00:00
maatheusgois-dd 6c15874d15 Clear workspace.additionalDirectories setting for isolated worktree subagent runs
Setting options.additionalDirectories to undefined wasn't enough:
createAgentSession also merges settings.get('workspace.additionalDirectories').
Now createSubagentSettings gets an override to clear the setting
when worktree is set, ensuring isolated runs can't edit outside the
worktree.

Co-authored-by: oh-my-pi <https://omp.sh>
2026-07-22 02:52:42 -03:00
maatheusgois-dd 2fa43ad5a6 Omit additionalDirectories for isolated (worktree) task runs
Isolated tasks clone only cwd into the worktree. Forwarding the
parent's additional directories would let the subagent edit absolute
paths under the original extra roots, bypassing isolation. Now
additionalDirectories is undefined when worktree is set.

Co-authored-by: oh-my-pi <https://omp.sh>
2026-07-22 02:44:56 -03:00
maatheusgois-dd ffe50650e1 Propagate workspace roots into subagent sessions
Subagents (task tool) now inherit the parent session's
additionalDirectories via ToolSession → ExecutorOptions →
CreateAgentSessionOptions, so delegated agents see the same
<workspace-roots> block and can read/grep/glob added roots.

Co-authored-by: oh-my-pi <https://omp.sh>
2026-07-22 02:36:38 -03:00
roboomp 04179bf0b8 fix(mcp): route task proxies through source tool and mark tools non-strict
MCP-backed tools never declared an explicit strict value, so OpenAI-family
serializers (post-#4336/#4340) had no false to preserve and models over-filled
mutually exclusive optional fields. Task/subagent proxies also rebuilt a raw
tools/call instead of executing through the source MCPTool, bypassing intent
stripping, placeholder pruning, local-URL resolution, reconnect, abort, and
result metadata; strict servers rejected proxied calls with
unrecognized_keys ["i"].

- MCPTool/DeferredMCPTool now declare `readonly strict = false as const`.
- createMCPProxyTools delegates to the current source tool, re-resolved by raw
  MCP server/tool metadata so reconnect replacements are honored, and keeps the
  Task 60s timeout by combining its abort signal with the caller's.
- Regression coverage: strict flags, proxy parity for i/placeholder shaping,
  declared-i passthrough, and reconnect re-resolution.

Fixes #6208
2026-07-21 22:37:12 +00:00
roboomp d52a759e5e fix(tui): refreshed prewalked subagent model
Tracked active subagent session model changes in progress snapshots so prewalk handoffs replace the starting-model badge.

Added regression coverage for a prewalk handoff and documented the fix.

Fixes #6083
2026-07-21 21:00:40 +00:00
Kormákur 57580aa413 fix(coding-agent): require a fresh yield after async-result deliveries in the quiescence barrier
The barrier in driveSessionToYield was unreachable for terminal yields:
the yield tool's shouldTerminate fired requestAbort("terminate"), so
abortSignal was always aborted before the barrier's loop condition ran,
and a run with pending owner jobs completed immediately with whatever
the pre-async yield said (Codex review on #6119).

- Split "stop the free-running turn after yield" from "terminate the
  run": a terminal yield with pending owner async work now parks the
  run with a recoverable session abort (budget-stop precedent) via
  requestYieldTurnStop; only a quiescent yield terminates.
- An async-result follow-up injected after a recorded yield un-latches
  it (transcript-ordered, in the run monitor) and re-runs the reminder
  ladder, so the run only completes on a yield that postdates every
  delivered result — including results injected during the notice turn.
- A run that never refreshes a superseded yield fails (exit 1) with an
  explicit reason; the stale payload ships only as failed-run salvage
  through the existing failed-after-yield finalize path.
- Rewrote subagent-async-pending.md: the "your current yield stands"
  option contradicted the enforced contract.
- Regression tests: parked yield -> injected result -> fresh yield wins;
  refusal -> stale payload fails; no-async fast path unchanged.
2026-07-20 21:53:38 +00:00
pr-eval f5bd6bbe0e fix(agent): count malformed yields after incremental sections
Narrow the invalid-yield guard to !abortSent so array-typed incremental
yield sections no longer suppress the infinite-submit-loop abort; add
regression coverage for incremental yield followed by repeated malformed
terminal yields.
2026-07-20 22:51:53 +02:00
can1357 a87d7d35cd Merge PR #4961: fix(agent): stop malformed subagent yield loops (@roboomp)
# Conflicts:
#	packages/coding-agent/src/prompts/system/workflow-notice.md
#	packages/coding-agent/src/task/executor.ts
#	packages/coding-agent/src/tools/yield.ts
2026-07-20 22:51:53 +02:00
Kormákur 21a7725a09 feat: lift subagent async/auto-background limits via owner-routed delivery and quiescence
Three-piece architecture so subagents inherit async.enabled and
bash.autoBackground.enabled instead of having both force-disabled:

- Owner-routed delivery: AsyncJobManager gains registerDeliverySink /
  waitForOwnerJobs; every AgentSession registers a sink for its own agent
  id, so background job results inject into the owning agent's run.
  Owned deliveries with no live sink dead-letter (result retained on the
  job row) instead of misrouting into the first top-level session.

- Quiescence barrier: a subagent's final yield with owner jobs still
  running/undelivered is a scheduling pause, not completion. The run
  driver notifies the model once (hub wait/cancel), settles owner work,
  and folds results in as async-result follow-ups; teardown cancels and
  awaits surviving jobs before isolation worktree capture/cleanup.

- Steering soft channel: queued steering no longer hard-aborts
  non-interruptible tools; it aborts interruptible waits and raises a
  cooperative ToolCallContext.steeringSignal. The mid-batch watch runs
  for every batch, and auto-backgroundable bash backgrounds itself on
  steer so incoming messages inject promptly with no work lost.
2026-07-20 15:00:37 +00:00
can1357 ad9d272efe Merge PR #5757: fix(task): inherit default fallback for single-model subagents (@jeffscottward) 2026-07-18 20:12:48 +02:00
metaphorics 8d8ad0faed perf(coding-agent): coalesce subagent output reconstruction
Defer recentOutput line reconstruction from every text_delta to the
progress emit boundary. appendRecentOutputTail only extends the capped
raw tail and marks dirty; refreshRecentOutput runs the exact old
split/filter/slice(-8)/reverse algorithm as the first step of every
emitProgressNow snapshot (onProgress + event bus), including coalesced
and finalize/error/cancel flushes. Reset publishes [] immediately;
replace marks dirty; past snapshot arrays stay immutable via spread.

Before (base pool median-of-5):
  w8_d3  61.55 cpu_ms/1k_events
  w32_d3 44.16 cpu_ms/1k_events
After (stable final run on E+G, 7 episodes, trimmed CV gate pass):
  w8_d3  55.78 cpu_ms/1k_events  (1.103×)  trimmed CV 15.1%
  w32_d3 40.72 cpu_ms/1k_events  (1.084×)  trimmed CV 11.1%
Checksums match prior exactness baseline; retained_after_release_kb
1284 / 2864 (no regression vs prior concur).

Op: GConcurEmitBoundary emit-boundary dirty flag
Restores: none
2026-07-18 11:07:51 +09:00
Jeff Scott Ward 6dcc485389 fix(task): inherit default subagent fallback 2026-07-17 17:43:04 -04:00
vmcall d53cf023b0 fix(task): reconciled structured subagents with upstream
- Preserved the plan-mode capability clamp after upstream removed report_finding.
- Updated persisted-revival coverage for mounted xdev tool activation.
- Applied current formatter output to conflicted runtime files.
2026-07-17 17:38:12 +02:00
vmcall d944879f21 feat(task): unified structured subagent execution
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.

Fixes #5279
2026-07-17 17:36:59 +02:00
can1357 1074332098 fix(task): abort background label generation 2026-07-17 04:15:40 +02:00
can1357 9afedb591e feat(coding-agent): added opt-in task prewalk and tightened --tools and xdev behavior
- Added a `task.prewalk` option (default `false`), removed default task `prewalk` flags, and updated prewalk resolution so bunded generic task execution only prewalks when explicitly enabled.
- Enforced strict `--tools` validation in CLI parsing, making unknown tool names fail fast with `CliUsageError` instead of being silently filtered.
- Migrated legacy discovery settings (`tools.discoveryMode`, `tools.essentialOverride`, MCP discovery keys) into updated `tools.xdev` handling with preserved explicit override behavior.
- Hardened xdev/ACP execution flow by capping `docsAll` payloads with overflow listing and remapping `xd://` dispatches/approval gating for correct execute/read behavior and reduced duplicate prompts.
2026-07-15 18:39:36 +02:00
can1357 5ff277349c refactor(coding-agent): consolidated tool surface onto xd:// devices and hub
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
2026-07-15 15:16:29 +02:00
can1357 c32ea55ba3 fix(coding-agent): kept todo active for prewalk subagents and fixed prewalk gate deadlock
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
2026-07-15 05:45:59 +02:00
can1357 425e583ae0 feat(coding-agent): added support for task-agent field and model resolution
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
2026-07-15 00:50:55 +02:00
roboomp 4bc68b45e2 fix(coding-agent): persisted vibe sessions across restarts
Vibe worker roster lived only in a process-local Map, so a resumed parent
session started with an empty registry and vibe_send failed with
"Unknown vibe session". Persist a versioned, parent-scoped lifecycle
journal (spawn/turn/tombstone events), rehydrate validated idle workers
through the persisted-subagent reviver on resume, and gate the flow with
generation/CAS protection so stale finalizers cannot clobber a
replacement worker. Killed transcripts stay readable but non-revivable;
mode-exit commits tombstones atomically with the mode change and rolls
back cleanly on storage failure.

Ported from @mastertyko's fork branch fix/vibe-session-persistence.

Fixes #5303
2026-07-14 17:58:14 +00:00
djdembeck aa470b9e34 fix(coding-agent): forward sessionId to getApiKey in subagent auth fallback
The pre-flight auth check in resolveModelOverrideWithAuthFallback called
getApiKey without a session id. For providers with session-sticky OAuth
credentials, this returned undefined even though the credential was
usable once the subagent session started, causing the auth fallback to
silently replace the configured model with the parent's (#5325).

The subagent's id is now forwarded as the session id so session-sticky
credentials resolve during the pre-flight check. Genuinely broken auth
(stale OAuth, revoked tokens) still falls back as before.

Also propagate model resolution warnings through resolveModelOverride
and log them in the executor so users see why a pattern didn't match.
2026-07-14 00:28:42 -05:00