24 Commits

Author SHA1 Message Date
can1357 906f5b1463 style(coding-agent/registry): formatted persisted mid-spawn test writes as template literals
- Replaced `join("\\n") + "\\n"` concatenations with template-literal interpolation in mid-spawn registry fixtures.
- Updated the three Bun.write calls in `persisted-mid-spawn-stub.test.ts` to use the template form.
2026-08-16 09:19:30 +03:00
can1357 45e4fa7884 Merge PR #8492: fix(coding-agent): reclaim dead parked agent corpses on respawn (@roboomp) 2026-08-16 02:43:34 +02:00
can1357 7f2da16cdb fix(hub): preserved concurrent live scan claims 2026-08-16 02:03:17 +02:00
roboomp 2929670a25 fix(coding-agent): preserved cold-revivable parked agents
Consulted the persisted subagent reviver factory before reclaiming an
unadopted parked ref. A valid cold reviver now preserves the existing agent;
factory failures also fail closed rather than deleting a potentially
recoverable generation. Reclaim invariants are rechecked after the async
probe to close lifecycle races.

Added regression coverage proving an unadopted disk-restored ref remains
registered and messageable through ensureLive.

Fixes #8490
2026-08-14 01:40:37 +00:00
roboomp 5421add83b fix(coding-agent): reclaim dead parked agent corpses on respawn
A parked, session-less agent-registry entry with no reviver permanently
poisoned its agent id for the process lifetime. `registerIfAvailable(input,
null)` refuses any existing entry, so a fresh subagent spawn reusing the id
died post-registration with "already owned by another session generation",
and messages to it failed with "is parked and cannot be revived". Such
corpses are left by an isolated run's park or an interrupted construction,
and there was no reclaim or GC path.

Add `AgentLifecycleManager.reclaimDeadCorpse`, which unregisters a ref only
when it is parked, holds no live session, is not owned by the manager (no
in-memory reviver) and has no park/revive in flight. The fresh-spawn path in
`createAgentSession` calls it on a collision and retries registration, so the
id becomes reusable. The registry's strict CAS contract is untouched; the
corpse's transcript stays readable at history://<id>.

Fixes #8490
2026-08-14 01:33:06 +00:00
Md Tahsin Rahman 41dfc2cd9e fix(hub): skip mid-spawn stubs in persisted scan
SessionManager.open writes title+session before createAgentSession
claims the id. Agent Hub parked that stub, so the spawn CAS failed
with "already owned by another session generation" and the row
could not be revived (no session_init).
2026-08-14 01:34:34 +08:00
can1357 b50ff2c4d6 Merge PR #8123: fix(agent): reject cold revive that finishes after lifecycle dispose (@roboomp) 2026-08-13 01:14:48 +02:00
roboomp dc47d73e9c fix(agent): recreated global lifecycle after disposal
Clear the disposed singleton after teardown so a later top-level session receives a fresh manager with an active registry subscription. Keep the retired manager marked disposed so its in-flight revivals still reject and clean up late sessions.

Add regression coverage for sequential top-level lifecycle owners cold-reviving in the same process.

Fixes #8114
2026-08-09 23:51:52 +00:00
roboomp 8061fb2e0a fix(agent): reject cold revive that finishes after lifecycle dispose
A cold persisted-agent revival could complete after
AgentLifecycleManager.dispose() and still attach its live session. When
teardown lands during the reviver-factory await, the id is not yet tracked
in #adopted, so dispose() skips it and leaves the parked registry ref
intact; the late factory then cold-adopts and #revive attaches a session
plus a TTL timer onto the torn-down manager, leaking a live session graph
and timers past teardown.

Set a #disposed flag at the top of dispose() and recheck it after the
reviver-factory await and after the reviver await. A late revive now
rejects deterministically and disposes any session it already built,
preserving singleflight for concurrent callers.

Fixes #8114
2026-08-09 23:46:40 +00:00
enieuwy ae5965241e fix(session): count every real output block as a served turn
Codex found the shared attribution predicate recognising only tool calls,
text and signed thinking. A native image response often arrives with no
text and no tool call at all, so an image-only turn was read as producing
nothing: attribution stayed on whichever model spoke before it, and the
empty-stop rule could classify a successful generation as empty.

Everything the assistant can emit now counts except two: unsigned thinking,
which is not provider-authenticated and was already excluded, and
Anthropic's `fallback` marker, which records that a request was routed
elsewhere rather than carrying output. Redacted thinking and server-tool
blocks are real work by the same argument as the image.

The `toolUse` arm keeps its stricter rule — an orphaned toolUse stop needs
a tool_use block to anchor a later tool_result, and an image cannot.

Also switches the new test to the namespace import AGENTS.md requires for
node builtins.
2026-08-08 13:16:42 +08:00
enieuwy a06a7f5ef1 refactor(session): anchor attribution to its transcript, drop a dead format branch
Two cleanups from review.

The offline walk matched a served turn to its `model_change` by also
accepting a `:level` thinking suffix, but no writer produces one: every
`appendModelChange` call site records a bare `provider/id` or a
`formatModelStringWithRouting` selector, and only the latter decorates —
with `@upstream`, never a level. Speculative matching in the part of the
change that already reconstructs the most from raw strings. Removed, with
its test narrowed to the format writers actually emit.

Attribution was also retained for the life of the recovery component, so
switching sessions in place could report a model from a different
transcript by name. It is now tagged with the session file it was earned
in and ignored when that no longer matches — the same self-invalidating
anchor `#ensurePersistedMessageKeys` uses, so no mutation call site has
to remember to clear anything. Anchoring on the file rather than the leaf
keeps attribution across compaction and appends, which stay within the
transcript that earned it.
2026-08-08 13:16:11 +08:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
Kyle McCleary 5cc4f5c93a Merge main into refactor/agent-hub-fullscreen 2026-08-04 18:15:42 -07:00
Kyle McCleary 8e5f619502 fix(coding-agent): harden Agent Hub lifecycle and persistence 2026-08-04 16:29:15 -07:00
metaphorics 687f3326b2 fix(task): bound abort cleanup and quarantine late jobs 2026-08-03 20:14:25 +09:00
roboomp 52a0af8625 fix(coding-agent): blocked stale revivals after hard kill
A parked revive holds the same AgentRef while constructing a new session.
If the Hub tombstoned that ref before revive completed, the reviver could
still attach its session and set the terminal ref back to idle because
identity alone remained unchanged.

- Made aborted registry refs terminal: reject revival claims, late session
  attachment, and status transitions out of aborted.
- Made lifecycle revival accept only an untouched detached parked ref or the
  exact running session already claimed by createAgentSession; dispose and
  reject every terminal/stale result.
- Added a delayed-revival regression test and claim-before-kill CAS checks.

Fixes #7250
2026-08-01 09:52:09 +00:00
roboomp 57732e9dcd fix(coding-agent): detach hard-aborted refs so ensureLive can't route into a dead session
Preserving aborted refs on dispose exposed a latent invariant break: the
executor's hard-abort path (finalizeSubagentLifecycle) set status `aborted`
and disposed the session without detaching it. With the ref now retained, it
kept a dangling pointer to the disposed session, and ensureLive returns any
non-null ref.session before its revivability check — so hub focus / transcript
chat could route into a dead session.

- finalizeSubagentLifecycle: detach the session before disposing on the
  terminal hard-abort path, upholding the AgentRef invariant (session === null
  when aborted).
- release(tombstone): detach before dispose too (capture the live session
  first), same invariant.
- unregisterUnlessParked: preserve `aborted` refs only when already detached;
  an aborted ref still holding a live session is a bug and is unregistered
  rather than kept reachable.
- Regression test now asserts ensureLive rejects a tombstoned id as terminal.

Fixes #7250
2026-08-01 09:42:41 +00:00
roboomp d350dd8f84 fix(coding-agent): preserve aborted tombstone across live-session dispose
A live-session hub kill did not stick: release(tombstone) awaited the
wrapped session dispose first, and createAgentSession's unregisterUnlessParked
removed any non-parked ref, so the subsequent detach/setStatus no-oped and the
ref was gone — leaving the reopen resurrection for idle/running agents.

- release(tombstone) now marks the ref `aborted` BEFORE disposing, so the
  dispose guard preserves it; the session is detached afterward.
- unregisterUnlessParked now also spares terminal `aborted` refs (matching
  the documented "hard-killed, terminal" retention and finalizeSubagentLifecycle).
- Regression test now uses a session stub that mirrors the real wrapped
  dispose (unregister unless parked/aborted), so it fails if the tombstone is
  set after dispose.

Fixes #7250
2026-08-01 09:34:43 +00:00
roboomp 65fd8cb0e8 fix(coding-agent): keep hub-killed subagent from resurrecting as parked
The Agent Hub kill path called AgentLifecycleManager.release(), which
disposes the session and then unregisters the ref while leaving the
on-disk <id>.jsonl intact. On the next hub open, registerPersistedSubagents
rescans the transcript tree and its `if (!registry.get(id))` guard cannot
distinguish an explicit kill from a normally-parked agent, so it re-adopts
the killed id as a fresh `parked` row.

Add a `tombstone` option to release() that mirrors finalizeSubagentLifecycle's
genuine-kill path: dispose and detach the session but keep the ref registered
as terminal `aborted` instead of removing it. The kept-registered id makes the
rescan guard skip it, and the transcript stays on disk (still reachable via
history://<id>, per #5261). The hub kill button now passes tombstone: true.

Fixes #7250
2026-08-01 09:26:26 +00:00
can1357 5d66eb7f2a Merge PR #5464: fix(coding-agent): persist vibe sessions across restarts (@roboomp) 2026-07-23 18:06:11 +02:00
roboomp 54f4a1894f fix(registry): coordinate park/dispose with ensureLive and IRC delivery
park() detached the live session only after session.dispose() resolved,
so during the dispose window the registry still exposed ref.session at
idle status. Concurrent ensureLive()/hub-send handed out or injected into
the dying session, reporting success while the message was dropped once
detach committed.

- Replace the #parking Set guard with a #parks map of in-flight park state
  that is cancelable until the session is detached.
- park() now yields a cancel window, then detaches + flips status to
  parked BEFORE dispose(), and coalesces concurrent park calls.
- ensureLive() cancels a pre-detach park (keeps the live session) or waits
  for detach+dispose then performs one coalesced revive; never returns a
  disposing session.
- release()/dispose() drain any in-flight park; the idle re-arm skips while
  a park owns the transition.
- IrcBus.send() gates parkable recipients through ensureLive and derives
  the revived receipt from session identity, so receipts/unread counts
  reflect actual delivery.

Fixes #5633
2026-07-16 01:21:08 +00:00
roboomp 4bc68b45e2 fix(coding-agent): persisted vibe sessions across restarts
Vibe worker roster lived only in a process-local Map, so a resumed parent
session started with an empty registry and vibe_send failed with
"Unknown vibe session". Persist a versioned, parent-scoped lifecycle
journal (spawn/turn/tombstone events), rehydrate validated idle workers
through the persisted-subagent reviver on resume, and gate the flow with
generation/CAS protection so stale finalizers cannot clobber a
replacement worker. Killed transcripts stay readable but non-revivable;
mode-exit commits tombstones atomically with the mode change and rolls
back cleanly on storage failure.

Ported from @mastertyko's fork branch fix/vibe-session-persistence.

Fixes #5303
2026-07-14 17:58:14 +00:00
can1357 ef5e5fd27c fix(coding-agent): fixed parked subagent restoration from persisted sessions
- Fixed cold revival flow so parked subagents are restored from persisted sessions at startup.
- Fixed session-init persistence to include spawns and readSummarize fields for replay accuracy.
- Fixed latest-session lookup by adding peekSessionInit for lock-free persisted contract access.
- Added lifecycle and session tests for cold-revive success, decline, and retry paths.
2026-06-17 00:21:20 +02:00
can1357 1654c759ca feat(coding-agent): added AgentLifecycleManager for idle/parked subagents
Introduces AgentLifecycleManager: when the task executor adopts a finished subagent the manager arms a TTL timer on idle, parks the agent on expiry (disposes the live session, keeps the AgentRef + sessionFile), and revives it on demand via an injected reviver. Only this manager flips parked ↔ idle. AgentRegistry now annotates session: AgentRef["session"] as null exactly when parked/aborted, and sdk.ts wires the lifecycle dispose into main-session teardown, derives agentKind once, and gates ref unregistration on parking so a parked agent stays addressable (history://, revive).
2026-06-10 17:46:02 +02:00