23 Commits

Author SHA1 Message Date
can1357 6b4823181b test: cleaned test suites and documented filtering guidelines
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
2026-08-13 08:28:42 +02:00
can1357 d57656f283 Merge PR #8074: fix(coding-agent): display subagent registration time locally (@anatoli-tsinovoy) 2026-08-11 15:06:13 +02:00
Anatoli Tsinovoy 5d8214ea77 fix(coding-agent): keep lineage timezone visible 2026-08-09 16:49:13 +03:00
Anatoli Tsinovoy d725ab10c1 fix(coding-agent): display subagent registration time locally 2026-08-09 16:35:07 +03:00
enieuwy a1e60c3450 refactor(session): fold the pending-fallback surface into servingModel
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.

Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.

The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.

`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
2026-08-08 13:16:41 +08:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
Kyle McCleary 5cc4f5c93a Merge main into refactor/agent-hub-fullscreen 2026-08-04 18:15:42 -07:00
Kyle McCleary 3180fd3d7f fix(coding-agent): complete Agent Hub inspector metadata 2026-08-04 17:20:00 -07:00
Kyle McCleary 8e5f619502 fix(coding-agent): harden Agent Hub lifecycle and persistence 2026-08-04 16:29:15 -07:00
Kyle McCleary 91467c2f27 fix(coding-agent): restore Agent Hub lineage metrics 2026-08-03 19:42:11 -07:00
Kyle McCleary 8f1de61e9f refactor(coding-agent): densify Agent Hub metrics 2026-08-03 19:42:11 -07:00
Kyle McCleary c30ca0b157 refactor(coding-agent): preserve Agent Hub history 2026-08-03 19:42:11 -07:00
Kyle McCleary bda330bb13 refactor(coding-agent): add Agent Hub inspector 2026-08-03 19:42:11 -07:00
Kyle McCleary 4fb3ecd69a refactor(coding-agent): polish agent hub menu 2026-08-03 19:42:11 -07:00
hancens e388bc14ce perf(agent-hub): bound roster rendering to viewport
Agent Hub previously pre-rendered every registry row and resolved each
row's observer via getSessions().find, which copy-sorts the full observer
map. On large rosters that made open/selection/age-tick O(all rows) and
observer lookup near O(N^2 log N).

- Add SessionObserverRegistry.getSession(id) for O(1) Map lookup
- Lazy-render hub rows around the selection within the terminal line budget
- Keep ordering, selection, status counts, overflow, badges, and task lines
- Add 10k-row regression covering bounded getSession/render work
2026-08-03 21:09:45 +08:00
roboomp 509017eda9 fix(tui): honored observed fallback flag on live hub rows
- Merged the executor-reported fallback flag into the live-session badge path so a Fireworks Fast to base degrade (which arms no session retry state) keeps its provenance.
- Added a live-row regression test for a fallback that populates no retryFallbackModel.

Fixes #6316
2026-07-22 19:24:32 +00:00
roboomp 00aed97ed1 fix(tui): flagged fallback badges for observer-only hub rows
- Threaded a resolvedModelIsFallback flag through AgentProgress and SingleResult.
- Set the flag from the executor retry-fallback handlers and settled results.
- Rendered the observer/no-session hub path as fallback -> provider/model.
- Added an observer-only fallback-badge regression test.

Fixes #6316
2026-07-22 19:19:08 +00:00
can1357 42d81f189f feat(coding-agent): redesigned agent hub layout for task clarity
- Redesigned agent hub entries as two-line cards, separating identity and status from task descriptions.
- Added explicit model and thinking level badges to agent status information.
- Updated agent hub rendering to support multi-line entries and adaptive vertical scrolling based on row height.
- Simplified status display by using glyphs instead of redundant status labels.
- Adjusted rendering logic to prioritize displaying agent task summaries on an independent indented line.
2026-07-13 03:03:26 +02:00
can1357 00e96db590 feat(coding-agent): throttled and constrained agent hud updates
- Added throttling and debouncing to HUD data rendering and observer UI synchronization to coalesce update bursts.
- Constrained the subagent HUD display to a maximum of 8 rows with a truncation notice for hidden sessions.
- Enhanced the session observer registry to categorize update types, enabling more granular UI reconciliation.
- Verified render coalescing and display truncation behavior with comprehensive integration tests using fake timers.
2026-07-06 07:38:08 +02:00
can1357 366313df33 test(coding-agent): improved stdout geometry stubs in tests
- Added no-op setters to process.stdout properties in test geometry stubs.
- Prevented potential errors when code under test attempts to reassign stdout dimensions.
2026-06-25 05:15:40 +02:00
can1357 44c7c6a3f2 fix(coding-agent): prevented terminal layout corruption from multiline content
- Sanitized newlines from agent hub display text to ensure single-line output.
- Enforced line clamping during render to prevent terminal wrapping.
- Added test coverage to verify output truncation and newline removal.
2026-06-18 03:09:29 +02:00
can1357 9e65499e15 test(coding-agent): simplified agent hub row ordering test assertions
- Added a helper that renders hub output and extracts ordered agent IDs.
- Replaced positional string assertions with explicit row-order expectations.
- Awaited asynchronous theme initialization in test setup.
2026-06-15 11:43:33 +02:00
can1357 9566c22c77 fix(coding-agent): preserved row order in open agent hub
- Captured each row's initial position in AgentHubOverlayComponent on first refresh and reused it on later refreshes.
- Changed subsequent sorting to prioritize status then prior row position, so activity updates no longer reordered visible rows.
- Added a regression test that verifies row ordering stays stable when activity changes and new agents append at the end.
2026-06-15 11:08:15 +02:00