- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.
Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.
The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.
`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.
Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.
Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.
Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.
One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.
Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.
A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.
`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
Agent Hub previously pre-rendered every registry row and resolved each
row's observer via getSessions().find, which copy-sorts the full observer
map. On large rosters that made open/selection/age-tick O(all rows) and
observer lookup near O(N^2 log N).
- Add SessionObserverRegistry.getSession(id) for O(1) Map lookup
- Lazy-render hub rows around the selection within the terminal line budget
- Keep ordering, selection, status counts, overflow, badges, and task lines
- Add 10k-row regression covering bounded getSession/render work
- Merged the executor-reported fallback flag into the live-session badge path so a Fireworks Fast to base degrade (which arms no session retry state) keeps its provenance.
- Added a live-row regression test for a fallback that populates no retryFallbackModel.
Fixes#6316
- Threaded a resolvedModelIsFallback flag through AgentProgress and SingleResult.
- Set the flag from the executor retry-fallback handlers and settled results.
- Rendered the observer/no-session hub path as fallback -> provider/model.
- Added an observer-only fallback-badge regression test.
Fixes#6316
- Redesigned agent hub entries as two-line cards, separating identity and status from task descriptions.
- Added explicit model and thinking level badges to agent status information.
- Updated agent hub rendering to support multi-line entries and adaptive vertical scrolling based on row height.
- Simplified status display by using glyphs instead of redundant status labels.
- Adjusted rendering logic to prioritize displaying agent task summaries on an independent indented line.
- Added throttling and debouncing to HUD data rendering and observer UI synchronization to coalesce update bursts.
- Constrained the subagent HUD display to a maximum of 8 rows with a truncation notice for hidden sessions.
- Enhanced the session observer registry to categorize update types, enabling more granular UI reconciliation.
- Verified render coalescing and display truncation behavior with comprehensive integration tests using fake timers.
- Added no-op setters to process.stdout properties in test geometry stubs.
- Prevented potential errors when code under test attempts to reassign stdout dimensions.
- Sanitized newlines from agent hub display text to ensure single-line output.
- Enforced line clamping during render to prevent terminal wrapping.
- Added test coverage to verify output truncation and newline removal.
- Added a helper that renders hub output and extracts ordered agent IDs.
- Replaced positional string assertions with explicit row-order expectations.
- Awaited asynchronous theme initialization in test setup.
- Captured each row's initial position in AgentHubOverlayComponent on first refresh and reused it on later refreshes.
- Changed subsequent sorting to prioritize status then prior row position, so activity updates no longer reordered visible rows.
- Added a regression test that verifies row ordering stays stable when activity changes and new agents append at the end.