Commit Graph

275 Commits

Author SHA1 Message Date
roboomp e8b7024b12 fix(hub): prevented stale agent refs from blocking wait
Mirrored registry status from pre-wire session run-state transitions and required live session corroboration before a peer can sustain bare hub waits.

Fixes #8634
2026-08-15 10:19:05 +00:00
roboomp 6fc50c3e41 fix(coding-agent): prevented post-yield TUI stalls
Kept the Bun event loop live across subagent yield drains and delayed parent result flushes. Added a timer-lifecycle regression for the idle flush.

Fixes #8462
2026-08-14 00:01:31 +02:00
can1357 b279db1790 test: refactored test suites to eliminate time-based sleeps and polling loops
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
2026-08-13 19:32:22 +02:00
can1357 e8eed95130 feat(coding-agent): introduced fine grained per agent advisor configuration
- Replaced blanket subagent advisor global settings with fine-grained per-agent configuration and frontmatter support.
- Added dashboard keybindings and inline override editors for managing agent advisor patterns.
- Implemented settings migration logic to convert legacy global options into per-agent settings.
- Updated session persistence and execution layers to restore and enforce per-agent advisor behaviors.
2026-08-13 04:59:50 +02:00
can1357 ca6e13fd47 Merge PR #8218: fix(agent): park subagents on shutdown (@roboomp) 2026-08-11 15:18:17 +02:00
can1357 6592b799b3 Merge PR #7980: fix(session): attribute a run to the model that produced its output (@enieuwy) 2026-08-11 15:06:12 +02:00
can1357 261eb965ac fix(task): preserved direct role alias fallback 2026-08-11 15:06:11 +02:00
can1357 0748651036 Merge PR #7910: fix(task): key subagent fallback chains off the pre-expansion model role (@enieuwy) 2026-08-11 15:06:11 +02:00
roboomp fb626e4ef0 fix(agent): let shutdown supersede a prior budget abort
requestAbort's abortSent branch only upgraded incoming signal reasons, so a
shutdown landing after a soft-budget hard-abort was discarded and abortKind()
stayed budget. finalizeSubagentLifecycle then followed the budget-resumable
path (idle + adopt) even though AgentLifecycleManager.dispose() had run,
leaking the subagent session into SDK/process reuse.

- Upgrade a prior budget abort to shutdown so the run takes the shutdown
  release path; genuine kills (signal/timeout/terminate) stay terminal and
  shutdown is never downgraded to signal.
- Cover the shutdown-races-budget-hard-abort case.

Fixes #8216
2026-08-11 06:51:52 +00:00
roboomp a2ea780571 fix(agent): parked subagents on shutdown
- Distinguished owning-manager shutdown from explicit job cancellation.
- Released shutdown-interrupted subagents without durable kill tombstones.
- Covered parked restart recovery and terminal explicit cancellation.

Fixes #8216
2026-08-11 06:24:31 +00:00
roboomp 506f57fbf2 fix(task): recorded subagent model performance
Shared the parent AgentStorage handle with isolated task settings while keeping setting overrides non-persistent.

Added regression coverage for task samples reaching the shared TPS/TTFT aggregate.

Fixes #8022
2026-08-08 15:59:27 +00:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
enieuwy 77ee3f2e7e fix(task): key subagent fallback chains off the pre-expansion model role
A single-model subagent is pinned to a `subagent:<id>` role whose
`retry.fallbackChains` entry shadows every configured role chain, so the
chain it inherits decides where the child retries. Inheritance resolved
the role by re-deriving it from the child's `modelPatterns` — but every
spawn path expands the role alias into `modelOverride` before calling
`runSubprocess` (`modelPatterns = normalizeModelPatterns(modelOverride ??
agent.model)`), so `@task` never reached the derivation and it returned
`undefined` every time. Every task subagent inherited `chains.default`.

With `modelRoles.task: anthropic/claude-sonnet-5`, `task` chained to
sonnet alone, and `default` chained to sonnet plus a second provider, a
transient stall on sonnet routed the child onto the default chain's
second model — one the operator had deliberately kept out of the `task`
chain — and a quota error there killed a 28-minute run.

#7694 fixed only the shape where an unexpanded alias reaches the
executor, which no production caller produces; its tests supplied a bare
`agent.model: ["@smol"]` with no `modelOverride`. The incident above
happened on v17.2.10, which contains that fix.

Route inheritance off the role identity the spawn path already computes
and passes as `modelRole`. Since that leaves the pattern-derived operand
unreachable, drop it and the parameter it was the only user of.

The vibe worker path had the same defect independently: `#resolveWorker`
expanded `@task`/`@smol` for the bundled `task`/`sonic` workers and kept
no role, so vibe children inherited `default` no matter what the
executor did. It now carries `modelRole` on `ResolvedVibeWorker` and
`VibeRecord` through both the spawn and rehydrate sites.

To stop the two halves drifting apart again — the mistake that caused
this bug — `resolveAgentModelSelection` returns the expanded `patterns`
and the pre-expansion `role` from one call, and both spawn paths take
both from it. `resolveAgentModelSource` is removed: its only use was
being fed to `resolveExplicitModelRole`, and keeping it invites the same
split derivation. `resolveAgentModelPatterns` stays for the UI callers
that legitimately want patterns alone.

Tests cover the producible shapes: the incident's chain layout (role
chain equal to the primary, default chain a superset), role identity
surviving expansion for every alias-routed bundled agent, and the
patterns/role pairing itself. #7694's two tests are re-anchored to a
shape a real caller produces.
2026-08-07 21:16:27 +08:00
can1357 677b6f10f5 Merge PR #7735: fix(coding-agent): return ToolInfo[] from getAllTools extension API (@roboomp) 2026-08-05 22:15:46 +02:00
roboomp 83496b8211 fix(coding-agent): returned ToolInfo[] from getAllTools extension API
The ExtensionAPI getAllTools() wired to session.getAllToolNames(),
returning bare tool-name strings. Upstream @earendil-works/pi-coding-agent
promises ToolInfo[] with sourceInfo, so extensions loaded through the
legacy-pi shim (e.g. gentle-pi) crashed on t.sourceInfo.source at every
session start.

Added SourceInfo/ToolInfo types plus SessionTools.getAllToolInfos(), which
returns { name, description, parameters, sourceInfo } and classifies each
tool as builtin/mcp/sdk/extension. Rewired every getAllTools action site
(interactive, acp, print/rpc, subagent executor) and the example extension.

Fixes #7732
2026-08-05 15:42:06 +00:00
enieuwy b279f06849 fix(task): inherit the aliased role's retry fallback chain for subagents
A single-model subagent has no fallbacks of its own, so it inherits one
and is pinned to a `subagent:<id>` role. That pin is inserted first in
`retry.fallbackChains` so no other role can capture its routing — which
also means it shadows every configured role chain at runtime.

Inheritance was hardcoded to `chains.default`, so a subagent spawned
through a role alias (the bundled scout's `model: "@smol"`) retried on
the default role's chain instead of its own. With `smol` chained to
composer/grok/luna and `default` chained to gpt-5.6-sol, every scout
fell back onto sol.

Resolve the inherited chain from the role identity still present in the
raw pattern (`@smol` -> `smol`), falling back to `default` when that
role configures no chain. An explicitly empty role chain still means
"no fallbacks", mirroring `expandDefaultRetryFallbackChains`. Explicit
model selectors keep inheriting `default`: they carry no role identity,
and a role assigned the same model must not capture the child's routing.
2026-08-05 18:13:42 +08:00
Kyle McCleary 0a0eca3834 fix(coding-agent): preserve override role provenance 2026-08-04 18:32:06 -07:00
Kyle McCleary 5cc4f5c93a Merge main into refactor/agent-hub-fullscreen 2026-08-04 18:15:42 -07:00
Kyle McCleary 3180fd3d7f fix(coding-agent): complete Agent Hub inspector metadata 2026-08-04 17:20:00 -07:00
Kyle McCleary 8e5f619502 fix(coding-agent): harden Agent Hub lifecycle and persistence 2026-08-04 16:29:15 -07:00
can1357 94c838faea Merge PR #7539: fix(coding-agent): complete usage-aware fallback integration (@eggpeat) 2026-08-05 01:12:02 +02:00
Kyle McCleary 91467c2f27 fix(coding-agent): restore Agent Hub lineage metrics 2026-08-03 19:42:11 -07:00
Kyle McCleary 8f1de61e9f refactor(coding-agent): densify Agent Hub metrics 2026-08-03 19:42:11 -07:00
Brent db97103c32 fix(coding-agent): complete usage-aware fallback integration 2026-08-03 16:34:35 +00:00
metaphorics e42b95bcae fix(task): report non-isolated cleanup state (#7488) 2026-08-03 22:26:24 +09:00
metaphorics d563e25cfe fix(task): settle all late cleanup (#7488) 2026-08-03 21:39:33 +09:00
metaphorics 687f3326b2 fix(task): bound abort cleanup and quarantine late jobs 2026-08-03 20:14:25 +09:00
roboomp 57732e9dcd fix(coding-agent): detach hard-aborted refs so ensureLive can't route into a dead session
Preserving aborted refs on dispose exposed a latent invariant break: the
executor's hard-abort path (finalizeSubagentLifecycle) set status `aborted`
and disposed the session without detaching it. With the ref now retained, it
kept a dangling pointer to the disposed session, and ensureLive returns any
non-null ref.session before its revivability check — so hub focus / transcript
chat could route into a dead session.

- finalizeSubagentLifecycle: detach the session before disposing on the
  terminal hard-abort path, upholding the AgentRef invariant (session === null
  when aborted).
- release(tombstone): detach before dispose too (capture the live session
  first), same invariant.
- unregisterUnlessParked: preserve `aborted` refs only when already detached;
  an aborted ref still holding a live session is a bug and is unregistered
  rather than kept reachable.
- Regression test now asserts ensureLive rejects a tombstoned id as terminal.

Fixes #7250
2026-08-01 09:42:41 +00:00
can1357 bd96752b83 fix(rpc): serialize IRC wake finalization 2026-07-31 19:28:43 +02:00
can1357 e02510f8c9 Merge PR #7108: fix(rpc): restore frames for IRC-revived subagents (@roboomp) 2026-07-31 19:28:42 +02:00
can1357 305ad95a20 fix(task): guard deferred launch work 2026-07-31 19:16:56 +02:00
Wolfgang Schoenberger dc292d4c84 fix(task): track deferred model refresh 2026-07-30 18:34:40 -07:00
Wolfgang Schoenberger 37a4adfcea perf(task): defer subagent launch setup 2026-07-30 15:16:47 -07:00
roboomp d686375b8c fix(rpc): monitored IRC wakes on cold-revived subagents
- Extracted the shared attachIrcWakeTurnMonitor from the executor reviver closure.
- Installed it in the persisted cold-revive path, forwarding the top-level event bus.
- Covered that a resumed process's parked subagent emits wake lifecycle frames.

Fixes #7105
2026-07-30 20:08:21 +00:00
roboomp b2d18b6920 fix(rpc): restored frames for IRC-revived subagents
- Monitored autonomous IRC wake turns with the task executor lifecycle and progress channels.
- Preserved monitoring after idle-TTL parking and session revival.
- Covered RPC subscriptions for both idle and parked keep-alive agents.

Fixes #7105
2026-07-30 19:42:58 +00:00
Kyle McCleary 089a9963f8 feat(security): OMP-native security subsystem (planner handoff) 2026-07-29 18:47:51 -07:00
Nik Divjak c3011fff3c fix(task): let task.softRequestBudget lower bundled subagent budgets
The soft request budget resolved to `SOFT_REQUEST_BUDGET[agent.name] ??
configured`, so the bundled entries for scout and sonic replaced the
configured value outright. Lowering `task.softRequestBudget` to tighten
the guard therefore did nothing for exactly the two agents that spawn
most often: a scout kept its 100-request budget no matter how small the
user set the knob. Only 0 (disable) and raising the value for
non-bundled agents had any effect.

Treat both numbers as upper bounds and take the smaller one. The bundled
entries stay ceilings, so a runaway scout is still stopped at 100 by
default and existing behavior is unchanged for anyone who has not
lowered the setting; a configured 0 still disables the guard entirely.
Resolution moves into `resolveSoftRequestBudget`, which also normalizes
negative and fractional inputs, so the rule is testable without standing
up a subprocess run.

This composes with `task.maxEffort` on a separate axis: effort caps how
hard each request thinks, this caps how many requests a run may spend.

(cherry picked from commit f0db29f8f725f11390b64ca9342300c482ff5c5d)
2026-07-29 23:09:01 +02:00
can1357 8c5dc16344 fix(task): enforced per-spawn effort ceiling across retry fallbacks
- task.maxEffort only clamped the initial thinking level; a retry
  fallback candidate could clamp back up to its model floor and run a
  low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
  ModelControls (constructor, setThinkingLevel, auto classifier,
  restore) and in applyRetryFallbackCandidate; fallback candidates whose
  floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
  attribution added.
- Review follow-up for PR #6794.
2026-07-27 16:09:28 +02:00
Wolfgang Schoenberger bd1605e8f5 feat(task): add per-spawn effort ceiling 2026-07-27 03:39:55 -07:00
roboomp 583ff590f2 fix(prewalk): apply same-model effort downgrades instead of skipping
The prewalk arm/switch guard compared model identity only (modelsAreEqual /
provider+id), discarding the resolved thinkingLevel. A legal same-model target
at a cheaper effort (e.g. prewalk: "@task" resolving to the active model at a
lower level) was dropped as a no-op, so the session ran the expensive effort for
the whole run while still paying the plan/continue nudges — silently on the
session path, logger.debug only on the subagent path.

Compare (provider, id, effective thinking level) via a shared prewalkWouldBeNoop
helper. Effort-only deltas on the same model now switch; a genuine no-op emits a
user-visible notice on the session path and never arms on the subagent path.

Fixes #6659
2026-07-26 02:20:03 +00:00
can1357 db937bd149 feat(coding-agent): added coarse effort parameter to task tool
- Add `effort` (`lo`/`med`/`hi`) parameter to task spawn parameters and prompts.
- Implement `resolveTaskEffortLevel` to map coarse task effort onto model-supported thinking ranges.
- Pass effort configuration through executor options and structured subagent requests.
2026-07-24 15:51:34 +02:00
Alex TYRODE e0928070c2 feat(extensions): expose session service tiers 2026-07-23 21:49:20 +00:00
can1357 5d66eb7f2a Merge PR #5464: fix(coding-agent): persist vibe sessions across restarts (@roboomp) 2026-07-23 18:06:11 +02:00
can1357 c522eceff2 Merge PR #6119: feat: lift subagent async/auto-background limits via owner-routed delivery and quiescence (@korri123)
# Conflicts:
#	packages/coding-agent/src/task/executor.ts
2026-07-23 17:52:52 +02:00
can1357 26726fdcb9 Merge PR #6255: add dynamic multi-root workspace context (@maatheusgois-dd) 2026-07-23 17:30:33 +02:00
can1357 db3a6a1407 Merge PR #6318: fix(tui): show fallback models in Agent Hub (@roboomp) 2026-07-23 11:37:13 +02:00
roboomp 00aed97ed1 fix(tui): flagged fallback badges for observer-only hub rows
- Threaded a resolvedModelIsFallback flag through AgentProgress and SingleResult.
- Set the flag from the executor retry-fallback handlers and settled results.
- Rendered the observer/no-session hub path as fallback -> provider/model.
- Added an observer-only fallback-badge regression test.

Fixes #6316
2026-07-22 19:19:08 +00:00
can1357 9952adbc82 Merge PR #6242: fix(mcp): route task proxies through source tool and mark tools non-strict (@roboomp) 2026-07-22 21:13:21 +02:00
can1357 5477e74bb3 Merge PR #6307: fix(coding-agent): align bundled task prewalk display (@roboomp) 2026-07-22 21:13:18 +02:00
can1357 a1bdf4a352 Merge PR #6214: fix(tui): refresh prewalked subagent model (@roboomp) 2026-07-22 21:13:18 +02:00