- Replaced legacy `compaction.strategy` and `remoteEnabled` settings with `compaction.methodOrder` across session maintenance, schema, and tests.
- Added automatic fallback mechanism to try subsequent compaction methods upon failure or unsupported model capabilities.
- Added mouse drag-and-drop reordering support and click handlers to multi-select settings submenus.
- Updated documentation and test suites to reflect ordered compaction strategy preferences and fallback chains.
A ThinkingLoop abort is the loop guard asking for a same-model resample
(it injects a thinking-loop-redirect notice that only makes sense on the
model that looped), not a provider failure. #handleRetryableError routed
it through the generic retryable-error branch, so on attempt 1 it called
noteRetryFallbackCooldown + #tryRetryModelFallback and could switch to
another family from retry.fallbackChains while parking the original
selector on a 5-minute cooldown. A healthy Grok 4.6 planning turn got
replaced by whatever the chain listed next.
Carve ThinkingLoop out of the model-fallback branch and out of the
Fireworks Fast->base degrade so the loop guard always re-samples the same
model; the retry budget still bounds a genuinely stuck stream.
Fixes#8760
Skipped same-provider cross-model fallback candidates when the latest assistant turn contains signed or redacted Anthropic thinking.
Kept same-model retry available for transient failures and added regression coverage.
Fixes#8558
Prevented deterministic 400 errors for immutable thinking blocks from entering same-model retries or configured model fallback.
Added a signed-thinking session regression covering the terminal error and retry UI lifecycle.
Fixes#8558
- Stage transcript initialization inside a detached TranscriptContainer to keep existing messages visible during incremental rendering.
- Add fallback state restoration in InteractiveMode.renderInitialMessages when chat rendering is aborted or fails.
- Update render-initial-messages tests to assert that old transcripts remain visible until replacements are fully committed.
The HerdR flicker fix routed direct HerdR panes onto the in-place
multiplexer resize path, but the randomized render-stress oracle still
modeled them as ED3 replaying direct terminals. The width-epoch ledger
now applies to any in-place-resize scenario (multiplexer + direct HerdR)
via a dedicated resizeRepaintsInPlace trait, keeping HerdR's direct
scrollback semantics intact. Adds a deterministic replay regression
(seed 0xcafed00d, 24 iterations) that fails without the oracle update.
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
Updates test assertions and fixtures to reflect that V4 Pro now exposes the full [low, high, max] effort ladder, switches the incompatible-fallback test to the openrouter deepseek-v4-pro entry, and adds a shared browser lease in the evaluation suite to avoid relaunching Chromium per test under load.
retry_fallback_applied and retry_fallback_succeeded were emitted to the
TUI and RPC subscribers but never reached extensions: AgentSession's
#emitExtensionEvent had no branch mapping either event to
ExtensionRunner.emit, and ExtensionEvent / ExtensionAPI.on lacked the
types and overloads, so registration was also rejected at compile time.
Add the typed events, on() overloads, union members, and the two
forwarding branches so extensions observe the same { from, to, role } /
{ model, role } payloads as TUI and RPC. Matches the contract already
documented in docs/non-compaction-retry-policy.md.
Fixes#8079
- Passed the known advisor role through retry fallback resolution so model and wildcard keys retain precedence while ambiguous role matches cannot select another role.
- Covered shared-model roles with distinct thinking levels and asserted advisor fallback lifecycle ownership.
Fixes#8075
Retry-fallback candidate selection filtered on suppression, effort
ceiling, model resolution, and API key, but never compared a
candidate's context window with the live context. A large-window
primary hitting a retryable error could switch onto a smaller-window
fallback and immediately send a predictably oversized request that the
provider rejects, stalling the run. This is the forward counterpart of
the #7952 cooldown-expiry revert fix.
Generalize the existing retry-fit budget check into
SessionMaintenance.contextFitsModel(model) and consult it from both
retry-fallback selection loops (#tryRetryModelFallback and the
usage-aware loop): skip any candidate whose usable window cannot hold
the current context and advance to the first configured candidate that
fits. The check is independent of compaction.enabled since an oversized
request overflows regardless.
Fixes#8065
Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.
`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
Review found the branch does not compile. `bun check` runs biome before
the per-workspace type check and joins them with `&&`, so a pre-existing
format error on `main` short-circuited the run: `tsgo --noEmit` never
executed, and `bun test` type-strips, so the suite stayed green over eight
type errors. `ServingModel` was used in `turn-recovery.ts` without being
imported, and three `subscribe` closures in the retry-fallback suite lost
the outer narrowing of `session`. Verified now against the workspace check
directly rather than the aggregate.
Codex also found `#fallbackRouted` outliving its session. It said how the
CURRENT model was reached, but nothing reset it when a transcript was
switched or resumed, so a freshly loaded session with no served attribution
yet described its own model with the previous session's routing. It is now
anchored on the session id exactly like the attribution beside it — the two
facts are earned together and expire together — and a session switch
reports the current model without claiming to know how it got there.
The replay-unsafe host stub gained `getSessionId`, which the real
`SessionManager` has always had and the anchoring now calls.
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.
Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.
The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.
`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
Two cleanups from review.
The offline walk matched a served turn to its `model_change` by also
accepting a `:level` thinking suffix, but no writer produces one: every
`appendModelChange` call site records a bare `provider/id` or a
`formatModelStringWithRouting` selector, and only the latter decorates —
with `@upstream`, never a level. Speculative matching in the part of the
change that already reconstructs the most from raw strings. Removed, with
its test narrowed to the format writers actually emit.
Attribution was also retained for the life of the recovery component, so
switching sessions in place could report a model from a different
transcript by name. It is now tagged with the session file it was earned
in and ignored when that no longer matches — the same self-invalidating
anchor `#ensurePersistedMessageKeys` uses, so no mutation call site has
to remember to clear anything. Anchoring on the file rather than the leaf
keeps attribution across compaction and appends, which stay within the
transcript that earned it.
Review of the previous commit found the same defect in four more places,
each a variant of one mistake: the unproven-model gate was expressed as
"a retry fallback is armed and has not served", so it only protected the
one swap path that arms a chain.
- The cooldown-expiry restore to the primary deletes the fallback record
before swapping, so the gate vanished with it and an unproven, just-
restored primary was credited immediately.
- A session's first-ever fallback arm constructs the record only after
the swap, so there was nothing to mark and the guard was a no-op.
- The Fireworks Fast degrade swaps models without arming a chain at all,
so a live row showed the degraded base model as the configured primary.
- The fresh-versus-resumed heuristic read the message list, which
compaction collapses, so a resumed-and-compacted session seeded its
startup candidate as already proven.
Attribution now names the last model that settled a turn here, full stop.
A switch of any kind — into a candidate, back to a restored primary, or a
capability degrade — moves it only once the new model answers. Before
anything has served there is no earlier work to miscredit, so the
configured model is both the only answer and a safe one.
Fallback routing is tracked as its own flag rather than inferred from the
chain record, because the degrade path routes without a chain, and it is
set before each swap so the synchronous `model_changed` fan-out cannot
observe a model mid-relabel. The resumed-session heuristic is gone: it
was only needed to decide whether an unserved candidate could be trusted,
and nothing is trusted before it serves.
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.
Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.
Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.
Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.
One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.
Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.
A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.
`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
- A provider-supplied retry-after now bypasses the transient rate/concurrency
heuristic window instead of being overridden by it (regression from the
subscription-cap retry change).
- Updated event-controller/ui-helpers test doubles for provenance-gated
renderer selection (hasBuiltInTool), aggregated retryErrors on
auto_retry_end, and Bedrock override compat gaining streamIdleTimeoutMs.
A cooldown-expiry model revert runs at a turn boundary. The user-prompt
path reverts then re-checks accumulated context against the restored
model via runPrePromptCompactionIfNeeded, but the automatic
agent.continue() path (#scheduleAgentContinue) reverted and issued the
next request with no such check. When a transient failure had fallen
back to a larger-window model and the conversation then grew past the
original model's window, restoring the primary once its cooldown expired
sent a predictably oversized request to the smaller model.
maybeRestoreRetryFallbackPrimary now reports whether it actually
switched, and the auto-continue path runs the same post-revert
context-fit maintenance (compaction/promotion) the prompt path already
runs, but only when a revert occurred.
Fixes#7952
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
- task.maxEffort only clamped the initial thinking level; a retry
fallback candidate could clamp back up to its model floor and run a
low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
ModelControls (constructor, setThinkingLevel, auto classifier,
restore) and in applyRetryFallbackCandidate; fallback candidates whose
floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
attribution added.
- Review follow-up for PR #6794.
A context rebuild that recreated the failed turn's message object made the
identity-keyed active-context removal miss, so the scheduled retry
continuation rejected the terminal assistant error message locally
("Cannot continue from message role: assistant") before any provider
request. auto_retry_end never fired, retryPromise stayed pending, and the
in-flight prompt() plus the TUI retry indicator hung until a manual
follow-up.
The retry path now strips a still-failed assistant tail positionally after
the backoff (generation-guarded, never in preserveFailedTurn mode), and a
continuation that still fails locally closes the retry saga with a failed
auto_retry_end via the new scheduleAgentContinue onError hook.
Fixes#5382
- Repointed tests off removed gpt-5.2/5.3 codex variants and devin models (e06ac0b787): context-promotion and TTSR tests pin gpt-5.5 -> gpt-5.6-sol via per-test modelOverrides since no bundled codex model has a runtime-effective promotion target anymore; replay-boundary and history-payload suites use gpt-5.5; advisor quota fallback uses devin/swe-1-6-slow with suffix-less selectors per #4579 devin-agent semantics; gateway-reference pins kilo/giga-potato and now asserts the reference carries effortRouting so the cross-provider no-inherit contract stays meaningful.
- Verified cross-provider gateway references do not inherit wire routing after variant collapse (identity/reference.ts:145-154) - stale test, no product bug.
- Exposed the active retry fallback selector from live agent sessions.
- Rendered fallback rows with an explicit marker and resolved provider/model.
- Added an end-to-end fallback-to-Agent-Hub regression assertion.
Fixes#6316
- Carried startup-selected fallback role and primary selector into AgentSession.
- Continued remaining role fallback entries after the startup fallback fails.
- Added regression coverage for chained startup failover.
Fixes#6283
- Added `#prunedTerminalRefusal` field to store the pruned refusal for post-settle consumers.
- Modified `getLastAssistantMessage()` to return the pruned refusal before active-context lookup.
- Reset `#prunedTerminalRefusal` on `agent_start` so a fresh run supersedes the settled refusal.
- Updated test mock helpers to include `getLastAssistantMessage` for consistency.
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:
- #failing latch: waitForCatchup resolves immediately while an advisor is
mid-failure; parked waiters wake the moment a turn fails, before any
async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
dedup snapshot and never propagates into the primary's turn-end callback
(per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
runtime latches quotaExhausted, requeues the batch, and notifies —
cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
before any assistant turn is recorded.
- Implemented parsing of id-prefixed wildcard keys and entries, allowing provider-specific prefixes in retry fallback configuration.
- Added logic to re-prefix failing model IDs and to match id-prefixed keys, with validation of provider existence.
- Updated settings schema description and changelog, and added tests covering the new behavior.
- Retained the advisor's original selector and thinking level while progressing through fallback candidates.
- Restored the configured primary before later advisor turns once its selector cooldown expired.
- Covered quota fallback restoration under the default cooldown-expiry policy.
- Switched advisor turns to the next configured model after provider quota or rate-limit failures.
- Emitted fallback applied and succeeded lifecycle events without reporting advisor unavailability after recovery.
- Added an end-to-end advisor quota fallback regression test.
Fixes#5740