- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
- Added account-scoped policy error detection to correctly identify Codex cyber-policy rejections.
- Updated credential storage and retry logic to route denied accounts through sibling rotation instead of bypassing it.
- Ensured coding-agent sessions exhaust all sibling accounts before falling back on cyber denials.
- Added comprehensive test coverage for credential rotation and retry behavior on policy errors.
- Separated deterministic replacement generation, placeholder derivation,
placeholder-range scanning and message-tree transforms out of the 2647-line
module; obfuscator.ts now holds the types and SecretObfuscator.
- ephemeralPlaceholderKey stays a single instance and both global regexes stay
beside the code that resets their lastIndex, so placeholder stability and
the security argument in the moved comments are preserved verbatim.
- Repointed every importer at the real modules rather than leaving a re-export
shim; the public ./secrets barrel exports the same 15 names as before.
A classifier refusal returned before the `onTurnError` hook that owns
model fallback, so `AdvisorRuntime` treated one provider's policy verdict
as terminal: `Refusal (cyber)` on the advisor model disabled the advisor
outright even with a fallback chain configured. Its only recovery was
stripping echoed primary reasoning and resending once, which does nothing
for a refusal about the content itself.
Route a refusal that outlives the strip through the same hook the generic
failure path uses, mirroring its epoch guard, session-transition requeue,
and requeue-on-recovery. The cascade walks the chain to exhaustion and
only reports the advisor unavailable once the host runs out of candidates,
matching what turn-recovery already allows for the primary.
Each cascade visits a model at most once. A switch re-arms
`#includeThinking` through `#syncModelIdentity`, so chain keys that point
back at each other (A to B, B to A) would otherwise strip-and-resend
against the same pair forever. A successful turn or a reset starts a
fresh walk.
Also stop `/advisor status` throwing when a live advisor has no roster
entry: `formatAdvisorStatus` guarded only the inactive case before
dereferencing `stats.advisors[0]`, and `#ensureAdvisors` clears
`#advisorStatuses` before repopulating it, so a status call landing in
that window hit `undefined.contextWindow`.
The advisor sends its whole Session update as a single ever-growing user
message. Provider prompt caches are prefix-based: a single user message whose
text keeps growing invalidates the entire message on every turn, so cache_read
stays pinned at the instructions/tools boundary (observed 14491 tokens in
production, 11066 in tests) instead of growing with the session.
Split the update into multiple user messages — one per source message —
delivered via a single Agent.prompt(AgentMessage[]) call, so the provider
caches each appended message incrementally. Verified end-to-end: cache_read
grows 0 -> 11126 -> 11457 -> 11583 across turns with the split, versus pinned
11066 on the old single-message behavior.
- delta-split.ts: pure renderAdvisorDeltaChunks using chunked
formatSessionHistoryMarkdown (shared toolResultIndex/consumedToolCallIds/
watchedRoleState) so toolCall/result pairing and role collapsing stay
byte-identical to the old single-block render (equivalence-tested).
- session-history-format.ts: add HistoryFormatOptions.watchedRoleState so
chunked renders collapse consecutive same-role messages exactly like the
single-block render.
- runtime.ts: #prepareBatch does a single dedup+render pass; #drain delivers
agent.prompt(preparedMessages) (array), falling back to the string.
- Keep field-selective fingerprint (candidate 1) + wip-marker-at-tail
(candidate 3) as complementary wins.
Tests: advisor suite 209 pass / 0 fail; type check clean; lint clean.
Affected subsets (342 tests) green; full suite hits WSL EMFILE fd limit.
- Remove `annotateForStaleness` and `hasFreshBacklog` from the advisor runtime.
- Stop appending staleness warnings to delivered advisor notes when newer primary turns queue.
Internal advisor context resets cleared failureNotified immediately after the quarantine warning, causing the status to return to running and repeated warnings every two quarantines.
Keep the notification latch through internal context re-primes. Clear it only on a successful advisor turn, explicit reset, or seed, and cover both deduplication and recovery in the quarantine regression test.
The advisor drain loop's quarantine branch reset context and re-primed silently with no bound, so an advisor that called an ungranted tool (e.g. bash) had its whole turn discarded before dispatch and its advice never reached the primary. Every other non-recovering failure branch calls notifyFailure -> emitNotice; quarantine was the one path with no main-UI signal, leaving supervision failures visible only in advisor diagnostics.
Count consecutive quarantines and, past MAX_QUARANTINE_RETRIES, surface the failure via notifyFailureOnce and drop the batch instead of looping silently. Reset the counter on any successful turn and on reset().
Fixes#6661
Cursor selects server-native tools (bash, grep, ...) outside the advisor's grant. Those exec-channel blocks are stamped kCursorExecResolved: they already ran server-side through the advisor-scoped CursorExecHandlers bridge, which rejects ungranted tools in-band. quarantineAdvisorUnsafeOutput was flagging them as pre-dispatch hazards and discarding the entire turn, dropping the legitimate advise emitted alongside them.
Skip exec-resolved native blocks in the unavailable-tool check so the scoped bridge stays the grant gate and the advisor can still deliver advice.
Fixes#5900
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:
- #failing latch: waitForCatchup resolves immediately while an advisor is
mid-failure; parked waiters wake the moment a turn fails, before any
async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
dedup snapshot and never propagates into the primary's turn-end callback
(per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
runtime latches quotaExhausted, requeues the batch, and notifies —
cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
before any assistant turn is recorded.
A broken advisor could hold the primary agent on the per-turn catch-up
gate for its full 30s budget while retrying, and an exception thrown from
onTurnEnd propagated into the primary's turn-end callback.
- waitForCatchup resolves immediately while the advisor is mid-failure
(new #failing latch, set at the failure catch BEFORE any async hook,
cleared on the next successful turn or reset/seed).
- Every parked waiter is woken the moment an advisor turn fails.
- The turn-end boundary isolates advisor exceptions per advisor: a
throwing advisor loses its delta, the primary and sibling advisors
continue untouched.
- A failed render (poisoned message, formatter bug) restores the delta
cursor and dedup state, so the delta is re-rendered next turn instead
of silently lost; the size probe itself is guarded and falls back to
the deferred renderer.
Grafted the evaluator's port (ec2c1e632) onto the merged advisor
runtime: terminal provider failures classified non-retriable (and not
context overflow) drop the bounded batch after one attempt with a
single notification; fallback-chain recovery and overflow recovery
retain precedence. Includes the one-prompt regression test and tags the
rollback-retry fixture's synthetic failure as transient.
Semantic merge with #5734 (delivered-prefix reconciliation) and #5748
(fallback chains): kept the coalescing round cap and wip threading,
adopted bounded cursor-preserving maintenance resets and overflow
recovery, and gated late-arrival consumption on coalescing rounds so
both suites' backlog and preserved-updates contracts hold.
Two shared failure modes with a single misbehaving advisor:
- A permanently rejected request (invalid_request_error, e.g. a model the
account no longer supports) retried forever: one notice, then silent
re-attempts on every turn, rebuilding heavy context each cycle. Quota
exhaustion already paused with a notice; this class now hard-stops the
runtime after a permanent rejection or three consecutive backlog-drop
cycles, with a visible notice. An explicit reset (/new, config rebuild,
restart) re-enables it, and waitForCatchup resolves while halted so the
primary agent never parks on a runtime that cannot drain.
- The delta render ran synchronously on the event loop; replaying a
multi-MB transcript after a reset blocked it for 600ms+ per render
(measured 675ms at ~54MB). Large deltas now render in size- and
count-bounded chunks that yield between slices (675ms -> single-digit
ms stalls). Tool call/result pairing survives chunk boundaries via a
shared whole-delta result index in formatSessionHistoryMarkdown; small
per-turn deltas keep the synchronous fast path.
- Switched advisor turns to the next configured model after provider quota or rate-limit failures.
- Emitted fallback applied and succeeded lifecycle events without reporting advisor unavailability after recovery.
- Added an end-to-end advisor quota fallback regression test.
Fixes#5740
- Tracked delivered message identities alongside the numeric cursor.
- Re-primed advisor context when a live transcript prefix diverged.
- Covered accepted empty-stop pruning before the next real user turn.
Fixes#5731