- Remove `annotateForStaleness` and `hasFreshBacklog` from the advisor runtime.
- Stop appending staleness warnings to delivered advisor notes when newer primary turns queue.
Internal advisor context resets cleared failureNotified immediately after the quarantine warning, causing the status to return to running and repeated warnings every two quarantines.
Keep the notification latch through internal context re-primes. Clear it only on a successful advisor turn, explicit reset, or seed, and cover both deduplication and recovery in the quarantine regression test.
The advisor drain loop's quarantine branch reset context and re-primed silently with no bound, so an advisor that called an ungranted tool (e.g. bash) had its whole turn discarded before dispatch and its advice never reached the primary. Every other non-recovering failure branch calls notifyFailure -> emitNotice; quarantine was the one path with no main-UI signal, leaving supervision failures visible only in advisor diagnostics.
Count consecutive quarantines and, past MAX_QUARANTINE_RETRIES, surface the failure via notifyFailureOnce and drop the batch instead of looping silently. Reset the counter on any successful turn and on reset().
Fixes#6661
Cursor selects server-native tools (bash, grep, ...) outside the advisor's grant. Those exec-channel blocks are stamped kCursorExecResolved: they already ran server-side through the advisor-scoped CursorExecHandlers bridge, which rejects ungranted tools in-band. quarantineAdvisorUnsafeOutput was flagging them as pre-dispatch hazards and discarding the entire turn, dropping the legitimate advise emitted alongside them.
Skip exec-resolved native blocks in the unavailable-tool check so the scoped bridge stays the grant gate and the advisor can still deliver advice.
Fixes#5900
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:
- #failing latch: waitForCatchup resolves immediately while an advisor is
mid-failure; parked waiters wake the moment a turn fails, before any
async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
dedup snapshot and never propagates into the primary's turn-end callback
(per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
runtime latches quotaExhausted, requeues the batch, and notifies —
cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
before any assistant turn is recorded.
A broken advisor could hold the primary agent on the per-turn catch-up
gate for its full 30s budget while retrying, and an exception thrown from
onTurnEnd propagated into the primary's turn-end callback.
- waitForCatchup resolves immediately while the advisor is mid-failure
(new #failing latch, set at the failure catch BEFORE any async hook,
cleared on the next successful turn or reset/seed).
- Every parked waiter is woken the moment an advisor turn fails.
- The turn-end boundary isolates advisor exceptions per advisor: a
throwing advisor loses its delta, the primary and sibling advisors
continue untouched.
- A failed render (poisoned message, formatter bug) restores the delta
cursor and dedup state, so the delta is re-rendered next turn instead
of silently lost; the size probe itself is guarded and falls back to
the deferred renderer.
Grafted the evaluator's port (ec2c1e632) onto the merged advisor
runtime: terminal provider failures classified non-retriable (and not
context overflow) drop the bounded batch after one attempt with a
single notification; fallback-chain recovery and overflow recovery
retain precedence. Includes the one-prompt regression test and tags the
rollback-retry fixture's synthetic failure as transient.
Semantic merge with #5734 (delivered-prefix reconciliation) and #5748
(fallback chains): kept the coalescing round cap and wip threading,
adopted bounded cursor-preserving maintenance resets and overflow
recovery, and gated late-arrival consumption on coalescing rounds so
both suites' backlog and preserved-updates contracts hold.
Two shared failure modes with a single misbehaving advisor:
- A permanently rejected request (invalid_request_error, e.g. a model the
account no longer supports) retried forever: one notice, then silent
re-attempts on every turn, rebuilding heavy context each cycle. Quota
exhaustion already paused with a notice; this class now hard-stops the
runtime after a permanent rejection or three consecutive backlog-drop
cycles, with a visible notice. An explicit reset (/new, config rebuild,
restart) re-enables it, and waitForCatchup resolves while halted so the
primary agent never parks on a runtime that cannot drain.
- The delta render ran synchronously on the event loop; replaying a
multi-MB transcript after a reset blocked it for 600ms+ per render
(measured 675ms at ~54MB). Large deltas now render in size- and
count-bounded chunks that yield between slices (675ms -> single-digit
ms stalls). Tool call/result pairing survives chunk boundaries via a
shared whole-delta result index in formatSessionHistoryMarkdown; small
per-turn deltas keep the synchronous fast path.
- Switched advisor turns to the next configured model after provider quota or rate-limit failures.
- Emitted fallback applied and succeeded lifecycle events without reporting advisor unavailability after recovery.
- Added an end-to-end advisor quota fallback regression test.
Fixes#5740
- Tracked delivered message identities alongside the numeric cursor.
- Re-primed advisor context when a live transcript prefix diverged.
- Covered accepted empty-stop pruning before the next real user turn.
Fixes#5731
- Replaced the advisor turn validation to raise errors only when a turn has no assistant response at all, instead of treating a content-less `stop` as a failure.
- Kept zero-content silent reviews from entering the retry/rollback/warn path by no longer raising "Advisor unavailable" for completed empty-stop turns.
- Updated advisor runtime tests and changelog language to assert and document that consecutive zero-usage silent turns now succeed without notifications.
The advisor runtime rejected every content-less stop completion as a
failed turn, so a deliberate silent review (the documented verifier
behavior) triggered retries and a spurious "unavailable" warning. Only
treat a content-less stop as a failure when it produced no output signal
(zero output and reasoning tokens), so a token-bearing silent stop counts
as a successful silent review.
Fixes#5493
- Anchored advisor compaction on provider-reported context usage (cached
input + generated output) floored by a full local estimate including the
advisor system prompt and tool schemas, so a near-full cached context is no
longer undercounted by the per-message estimate.
- Rejected stale provider usage retained across advisor compaction via a
runtime-only usage-anchor boundary recorded on the summary message.
- Recovered provider overflow by clearing only the advisor's own context at
the current primary cursor, retrying the bounded failing batch once against
a fresh context without replaying old primary history, and keeping later
updates eligible.
- Threaded the selected dashboard range through the stats Recent Errors UI,
API, and database timestamp filter before ordering and the 50-row limit.
Fixes#5282