Commit Graph

277 Commits

Author SHA1 Message Date
can1357 aeed1e6195 refactor(coding-agent): extracted shared auto-backgrounding and job wait helpers
- Extract foreground wait, notice formatting, and settlement racing logic into a shared module.
- Update async job type definitions to support eval jobs alongside bash and task jobs.
- Add configuration settings for eval auto-background thresholds.
2026-08-20 05:17:47 +02:00
roboomp 3acc57de8c fix(task): initialize extension runtime on subagent revival
Both subagent revivers rebuilt the session but never wired the extension
runtime, leaving it pre-init where every action method throws
ExtensionRuntimeNotInitializedError. An extension with a tool_call handler
touching a runtime action then tripped the fail-closed gate in emitToolCall
and blocked every tool, including the hidden yield, so the revived agent
could neither finish nor exit and looped until killed.

Both the warm lifecycle reviver (executor.ts) and the cold persisted
reviver (persisted-revive.ts) now call the shared initializeExtensions
helper on the rebuilt session, restoring runtime actions, onError, and the
session_start event.

Fixes #8824
2026-08-19 08:44:55 +00:00
roboomp e8b7024b12 fix(hub): prevented stale agent refs from blocking wait
Mirrored registry status from pre-wire session run-state transitions and required live session corroboration before a peer can sustain bare hub waits.

Fixes #8634
2026-08-15 10:19:05 +00:00
roboomp 6fc50c3e41 fix(coding-agent): prevented post-yield TUI stalls
Kept the Bun event loop live across subagent yield drains and delayed parent result flushes. Added a timer-lifecycle regression for the idle flush.

Fixes #8462
2026-08-14 00:01:31 +02:00
can1357 b279db1790 test: refactored test suites to eliminate time-based sleeps and polling loops
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
2026-08-13 19:32:22 +02:00
can1357 e8eed95130 feat(coding-agent): introduced fine grained per agent advisor configuration
- Replaced blanket subagent advisor global settings with fine-grained per-agent configuration and frontmatter support.
- Added dashboard keybindings and inline override editors for managing agent advisor patterns.
- Implemented settings migration logic to convert legacy global options into per-agent settings.
- Updated session persistence and execution layers to restore and enforce per-agent advisor behaviors.
2026-08-13 04:59:50 +02:00
can1357 ca6e13fd47 Merge PR #8218: fix(agent): park subagents on shutdown (@roboomp) 2026-08-11 15:18:17 +02:00
can1357 6592b799b3 Merge PR #7980: fix(session): attribute a run to the model that produced its output (@enieuwy) 2026-08-11 15:06:12 +02:00
can1357 261eb965ac fix(task): preserved direct role alias fallback 2026-08-11 15:06:11 +02:00
can1357 0748651036 Merge PR #7910: fix(task): key subagent fallback chains off the pre-expansion model role (@enieuwy) 2026-08-11 15:06:11 +02:00
roboomp fb626e4ef0 fix(agent): let shutdown supersede a prior budget abort
requestAbort's abortSent branch only upgraded incoming signal reasons, so a
shutdown landing after a soft-budget hard-abort was discarded and abortKind()
stayed budget. finalizeSubagentLifecycle then followed the budget-resumable
path (idle + adopt) even though AgentLifecycleManager.dispose() had run,
leaking the subagent session into SDK/process reuse.

- Upgrade a prior budget abort to shutdown so the run takes the shutdown
  release path; genuine kills (signal/timeout/terminate) stay terminal and
  shutdown is never downgraded to signal.
- Cover the shutdown-races-budget-hard-abort case.

Fixes #8216
2026-08-11 06:51:52 +00:00
roboomp a2ea780571 fix(agent): parked subagents on shutdown
- Distinguished owning-manager shutdown from explicit job cancellation.
- Released shutdown-interrupted subagents without durable kill tombstones.
- Covered parked restart recovery and terminal explicit cancellation.

Fixes #8216
2026-08-11 06:24:31 +00:00
roboomp 506f57fbf2 fix(task): recorded subagent model performance
Shared the parent AgentStorage handle with isolated task settings while keeping setting overrides non-persistent.

Added regression coverage for task samples reaching the shared TPS/TTFT aggregate.

Fixes #8022
2026-08-08 15:59:27 +00:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
enieuwy 77ee3f2e7e fix(task): key subagent fallback chains off the pre-expansion model role
A single-model subagent is pinned to a `subagent:<id>` role whose
`retry.fallbackChains` entry shadows every configured role chain, so the
chain it inherits decides where the child retries. Inheritance resolved
the role by re-deriving it from the child's `modelPatterns` — but every
spawn path expands the role alias into `modelOverride` before calling
`runSubprocess` (`modelPatterns = normalizeModelPatterns(modelOverride ??
agent.model)`), so `@task` never reached the derivation and it returned
`undefined` every time. Every task subagent inherited `chains.default`.

With `modelRoles.task: anthropic/claude-sonnet-5`, `task` chained to
sonnet alone, and `default` chained to sonnet plus a second provider, a
transient stall on sonnet routed the child onto the default chain's
second model — one the operator had deliberately kept out of the `task`
chain — and a quota error there killed a 28-minute run.

#7694 fixed only the shape where an unexpanded alias reaches the
executor, which no production caller produces; its tests supplied a bare
`agent.model: ["@smol"]` with no `modelOverride`. The incident above
happened on v17.2.10, which contains that fix.

Route inheritance off the role identity the spawn path already computes
and passes as `modelRole`. Since that leaves the pattern-derived operand
unreachable, drop it and the parameter it was the only user of.

The vibe worker path had the same defect independently: `#resolveWorker`
expanded `@task`/`@smol` for the bundled `task`/`sonic` workers and kept
no role, so vibe children inherited `default` no matter what the
executor did. It now carries `modelRole` on `ResolvedVibeWorker` and
`VibeRecord` through both the spawn and rehydrate sites.

To stop the two halves drifting apart again — the mistake that caused
this bug — `resolveAgentModelSelection` returns the expanded `patterns`
and the pre-expansion `role` from one call, and both spawn paths take
both from it. `resolveAgentModelSource` is removed: its only use was
being fed to `resolveExplicitModelRole`, and keeping it invites the same
split derivation. `resolveAgentModelPatterns` stays for the UI callers
that legitimately want patterns alone.

Tests cover the producible shapes: the incident's chain layout (role
chain equal to the primary, default chain a superset), role identity
surviving expansion for every alias-routed bundled agent, and the
patterns/role pairing itself. #7694's two tests are re-anchored to a
shape a real caller produces.
2026-08-07 21:16:27 +08:00
can1357 677b6f10f5 Merge PR #7735: fix(coding-agent): return ToolInfo[] from getAllTools extension API (@roboomp) 2026-08-05 22:15:46 +02:00
roboomp 83496b8211 fix(coding-agent): returned ToolInfo[] from getAllTools extension API
The ExtensionAPI getAllTools() wired to session.getAllToolNames(),
returning bare tool-name strings. Upstream @earendil-works/pi-coding-agent
promises ToolInfo[] with sourceInfo, so extensions loaded through the
legacy-pi shim (e.g. gentle-pi) crashed on t.sourceInfo.source at every
session start.

Added SourceInfo/ToolInfo types plus SessionTools.getAllToolInfos(), which
returns { name, description, parameters, sourceInfo } and classifies each
tool as builtin/mcp/sdk/extension. Rewired every getAllTools action site
(interactive, acp, print/rpc, subagent executor) and the example extension.

Fixes #7732
2026-08-05 15:42:06 +00:00
enieuwy b279f06849 fix(task): inherit the aliased role's retry fallback chain for subagents
A single-model subagent has no fallbacks of its own, so it inherits one
and is pinned to a `subagent:<id>` role. That pin is inserted first in
`retry.fallbackChains` so no other role can capture its routing — which
also means it shadows every configured role chain at runtime.

Inheritance was hardcoded to `chains.default`, so a subagent spawned
through a role alias (the bundled scout's `model: "@smol"`) retried on
the default role's chain instead of its own. With `smol` chained to
composer/grok/luna and `default` chained to gpt-5.6-sol, every scout
fell back onto sol.

Resolve the inherited chain from the role identity still present in the
raw pattern (`@smol` -> `smol`), falling back to `default` when that
role configures no chain. An explicitly empty role chain still means
"no fallbacks", mirroring `expandDefaultRetryFallbackChains`. Explicit
model selectors keep inheriting `default`: they carry no role identity,
and a role assigned the same model must not capture the child's routing.
2026-08-05 18:13:42 +08:00
Kyle McCleary 0a0eca3834 fix(coding-agent): preserve override role provenance 2026-08-04 18:32:06 -07:00
Kyle McCleary 5cc4f5c93a Merge main into refactor/agent-hub-fullscreen 2026-08-04 18:15:42 -07:00
Kyle McCleary 3180fd3d7f fix(coding-agent): complete Agent Hub inspector metadata 2026-08-04 17:20:00 -07:00
Kyle McCleary 8e5f619502 fix(coding-agent): harden Agent Hub lifecycle and persistence 2026-08-04 16:29:15 -07:00
can1357 94c838faea Merge PR #7539: fix(coding-agent): complete usage-aware fallback integration (@eggpeat) 2026-08-05 01:12:02 +02:00
Kyle McCleary 91467c2f27 fix(coding-agent): restore Agent Hub lineage metrics 2026-08-03 19:42:11 -07:00
Kyle McCleary 8f1de61e9f refactor(coding-agent): densify Agent Hub metrics 2026-08-03 19:42:11 -07:00
Brent db97103c32 fix(coding-agent): complete usage-aware fallback integration 2026-08-03 16:34:35 +00:00
metaphorics e42b95bcae fix(task): report non-isolated cleanup state (#7488) 2026-08-03 22:26:24 +09:00
metaphorics d563e25cfe fix(task): settle all late cleanup (#7488) 2026-08-03 21:39:33 +09:00
metaphorics 687f3326b2 fix(task): bound abort cleanup and quarantine late jobs 2026-08-03 20:14:25 +09:00
roboomp 57732e9dcd fix(coding-agent): detach hard-aborted refs so ensureLive can't route into a dead session
Preserving aborted refs on dispose exposed a latent invariant break: the
executor's hard-abort path (finalizeSubagentLifecycle) set status `aborted`
and disposed the session without detaching it. With the ref now retained, it
kept a dangling pointer to the disposed session, and ensureLive returns any
non-null ref.session before its revivability check — so hub focus / transcript
chat could route into a dead session.

- finalizeSubagentLifecycle: detach the session before disposing on the
  terminal hard-abort path, upholding the AgentRef invariant (session === null
  when aborted).
- release(tombstone): detach before dispose too (capture the live session
  first), same invariant.
- unregisterUnlessParked: preserve `aborted` refs only when already detached;
  an aborted ref still holding a live session is a bug and is unregistered
  rather than kept reachable.
- Regression test now asserts ensureLive rejects a tombstoned id as terminal.

Fixes #7250
2026-08-01 09:42:41 +00:00
can1357 bd96752b83 fix(rpc): serialize IRC wake finalization 2026-07-31 19:28:43 +02:00
can1357 e02510f8c9 Merge PR #7108: fix(rpc): restore frames for IRC-revived subagents (@roboomp) 2026-07-31 19:28:42 +02:00
can1357 305ad95a20 fix(task): guard deferred launch work 2026-07-31 19:16:56 +02:00
Wolfgang Schoenberger dc292d4c84 fix(task): track deferred model refresh 2026-07-30 18:34:40 -07:00
Wolfgang Schoenberger 37a4adfcea perf(task): defer subagent launch setup 2026-07-30 15:16:47 -07:00
roboomp d686375b8c fix(rpc): monitored IRC wakes on cold-revived subagents
- Extracted the shared attachIrcWakeTurnMonitor from the executor reviver closure.
- Installed it in the persisted cold-revive path, forwarding the top-level event bus.
- Covered that a resumed process's parked subagent emits wake lifecycle frames.

Fixes #7105
2026-07-30 20:08:21 +00:00
roboomp b2d18b6920 fix(rpc): restored frames for IRC-revived subagents
- Monitored autonomous IRC wake turns with the task executor lifecycle and progress channels.
- Preserved monitoring after idle-TTL parking and session revival.
- Covered RPC subscriptions for both idle and parked keep-alive agents.

Fixes #7105
2026-07-30 19:42:58 +00:00
Kyle McCleary 089a9963f8 feat(security): OMP-native security subsystem (planner handoff) 2026-07-29 18:47:51 -07:00
Nik Divjak c3011fff3c fix(task): let task.softRequestBudget lower bundled subagent budgets
The soft request budget resolved to `SOFT_REQUEST_BUDGET[agent.name] ??
configured`, so the bundled entries for scout and sonic replaced the
configured value outright. Lowering `task.softRequestBudget` to tighten
the guard therefore did nothing for exactly the two agents that spawn
most often: a scout kept its 100-request budget no matter how small the
user set the knob. Only 0 (disable) and raising the value for
non-bundled agents had any effect.

Treat both numbers as upper bounds and take the smaller one. The bundled
entries stay ceilings, so a runaway scout is still stopped at 100 by
default and existing behavior is unchanged for anyone who has not
lowered the setting; a configured 0 still disables the guard entirely.
Resolution moves into `resolveSoftRequestBudget`, which also normalizes
negative and fractional inputs, so the rule is testable without standing
up a subprocess run.

This composes with `task.maxEffort` on a separate axis: effort caps how
hard each request thinks, this caps how many requests a run may spend.

(cherry picked from commit f0db29f8f725f11390b64ca9342300c482ff5c5d)
2026-07-29 23:09:01 +02:00
can1357 8c5dc16344 fix(task): enforced per-spawn effort ceiling across retry fallbacks
- task.maxEffort only clamped the initial thinking level; a retry
  fallback candidate could clamp back up to its model floor and run a
  low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
  ModelControls (constructor, setThinkingLevel, auto classifier,
  restore) and in applyRetryFallbackCandidate; fallback candidates whose
  floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
  attribution added.
- Review follow-up for PR #6794.
2026-07-27 16:09:28 +02:00
Wolfgang Schoenberger bd1605e8f5 feat(task): add per-spawn effort ceiling 2026-07-27 03:39:55 -07:00
roboomp 583ff590f2 fix(prewalk): apply same-model effort downgrades instead of skipping
The prewalk arm/switch guard compared model identity only (modelsAreEqual /
provider+id), discarding the resolved thinkingLevel. A legal same-model target
at a cheaper effort (e.g. prewalk: "@task" resolving to the active model at a
lower level) was dropped as a no-op, so the session ran the expensive effort for
the whole run while still paying the plan/continue nudges — silently on the
session path, logger.debug only on the subagent path.

Compare (provider, id, effective thinking level) via a shared prewalkWouldBeNoop
helper. Effort-only deltas on the same model now switch; a genuine no-op emits a
user-visible notice on the session path and never arms on the subagent path.

Fixes #6659
2026-07-26 02:20:03 +00:00
can1357 db937bd149 feat(coding-agent): added coarse effort parameter to task tool
- Add `effort` (`lo`/`med`/`hi`) parameter to task spawn parameters and prompts.
- Implement `resolveTaskEffortLevel` to map coarse task effort onto model-supported thinking ranges.
- Pass effort configuration through executor options and structured subagent requests.
2026-07-24 15:51:34 +02:00
Alex TYRODE e0928070c2 feat(extensions): expose session service tiers 2026-07-23 21:49:20 +00:00
can1357 5d66eb7f2a Merge PR #5464: fix(coding-agent): persist vibe sessions across restarts (@roboomp) 2026-07-23 18:06:11 +02:00
can1357 c522eceff2 Merge PR #6119: feat: lift subagent async/auto-background limits via owner-routed delivery and quiescence (@korri123)
# Conflicts:
#	packages/coding-agent/src/task/executor.ts
2026-07-23 17:52:52 +02:00
can1357 26726fdcb9 Merge PR #6255: add dynamic multi-root workspace context (@maatheusgois-dd) 2026-07-23 17:30:33 +02:00
can1357 db3a6a1407 Merge PR #6318: fix(tui): show fallback models in Agent Hub (@roboomp) 2026-07-23 11:37:13 +02:00
roboomp 00aed97ed1 fix(tui): flagged fallback badges for observer-only hub rows
- Threaded a resolvedModelIsFallback flag through AgentProgress and SingleResult.
- Set the flag from the executor retry-fallback handlers and settled results.
- Rendered the observer/no-session hub path as fallback -> provider/model.
- Added an observer-only fallback-badge regression test.

Fixes #6316
2026-07-22 19:19:08 +00:00
can1357 9952adbc82 Merge PR #6242: fix(mcp): route task proxies through source tool and mark tools non-strict (@roboomp) 2026-07-22 21:13:21 +02:00