- A provider-supplied retry-after now bypasses the transient rate/concurrency
heuristic window instead of being overridden by it (regression from the
subscription-cap retry change).
- Updated event-controller/ui-helpers test doubles for provenance-gated
renderer selection (hasBuiltInTool), aggregated retryErrors on
auto_retry_end, and Bedrock override compat gaining streamIdleTimeoutMs.
A cooldown-expiry model revert runs at a turn boundary. The user-prompt
path reverts then re-checks accumulated context against the restored
model via runPrePromptCompactionIfNeeded, but the automatic
agent.continue() path (#scheduleAgentContinue) reverted and issued the
next request with no such check. When a transient failure had fallen
back to a larger-window model and the conversation then grew past the
original model's window, restoring the primary once its cooldown expired
sent a predictably oversized request to the smaller model.
maybeRestoreRetryFallbackPrimary now reports whether it actually
switched, and the auto-continue path runs the same post-revert
context-fit maintenance (compaction/promotion) the prompt path already
runs, but only when a revert occurred.
Fixes#7952
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
- task.maxEffort only clamped the initial thinking level; a retry
fallback candidate could clamp back up to its model floor and run a
low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
ModelControls (constructor, setThinkingLevel, auto classifier,
restore) and in applyRetryFallbackCandidate; fallback candidates whose
floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
attribution added.
- Review follow-up for PR #6794.
A context rebuild that recreated the failed turn's message object made the
identity-keyed active-context removal miss, so the scheduled retry
continuation rejected the terminal assistant error message locally
("Cannot continue from message role: assistant") before any provider
request. auto_retry_end never fired, retryPromise stayed pending, and the
in-flight prompt() plus the TUI retry indicator hung until a manual
follow-up.
The retry path now strips a still-failed assistant tail positionally after
the backoff (generation-guarded, never in preserveFailedTurn mode), and a
continuation that still fails locally closes the retry saga with a failed
auto_retry_end via the new scheduleAgentContinue onError hook.
Fixes#5382
- Repointed tests off removed gpt-5.2/5.3 codex variants and devin models (e06ac0b787): context-promotion and TTSR tests pin gpt-5.5 -> gpt-5.6-sol via per-test modelOverrides since no bundled codex model has a runtime-effective promotion target anymore; replay-boundary and history-payload suites use gpt-5.5; advisor quota fallback uses devin/swe-1-6-slow with suffix-less selectors per #4579 devin-agent semantics; gateway-reference pins kilo/giga-potato and now asserts the reference carries effortRouting so the cross-provider no-inherit contract stays meaningful.
- Verified cross-provider gateway references do not inherit wire routing after variant collapse (identity/reference.ts:145-154) - stale test, no product bug.
- Exposed the active retry fallback selector from live agent sessions.
- Rendered fallback rows with an explicit marker and resolved provider/model.
- Added an end-to-end fallback-to-Agent-Hub regression assertion.
Fixes#6316
- Carried startup-selected fallback role and primary selector into AgentSession.
- Continued remaining role fallback entries after the startup fallback fails.
- Added regression coverage for chained startup failover.
Fixes#6283
- Added `#prunedTerminalRefusal` field to store the pruned refusal for post-settle consumers.
- Modified `getLastAssistantMessage()` to return the pruned refusal before active-context lookup.
- Reset `#prunedTerminalRefusal` on `agent_start` so a fresh run supersedes the settled refusal.
- Updated test mock helpers to include `getLastAssistantMessage` for consistency.
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:
- #failing latch: waitForCatchup resolves immediately while an advisor is
mid-failure; parked waiters wake the moment a turn fails, before any
async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
dedup snapshot and never propagates into the primary's turn-end callback
(per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
runtime latches quotaExhausted, requeues the batch, and notifies —
cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
before any assistant turn is recorded.
- Implemented parsing of id-prefixed wildcard keys and entries, allowing provider-specific prefixes in retry fallback configuration.
- Added logic to re-prefix failing model IDs and to match id-prefixed keys, with validation of provider existence.
- Updated settings schema description and changelog, and added tests covering the new behavior.
- Retained the advisor's original selector and thinking level while progressing through fallback candidates.
- Restored the configured primary before later advisor turns once its selector cooldown expired.
- Covered quota fallback restoration under the default cooldown-expiry policy.
- Switched advisor turns to the next configured model after provider quota or rate-limit failures.
- Emitted fallback applied and succeeded lifecycle events without reporting advisor unavailability after recovery.
- Added an end-to-end advisor quota fallback regression test.
Fixes#5740
A stalled or dropped provider stream that surfaces as stopReason:"error"
carrying the bare "Request was aborted" sentinel fell through both retry
gates: #isRetryableReasonlessAbort required stopReason:"aborted", and
#isRetryableError's classifier returns no retriable kinds for the generic
sentinel. The turn died immediately despite retry.enabled.
- Relaxed #isRetryableReasonlessAbort to accept an empty generic-abort
sentinel turn under stopReason "aborted" or "error", tagging it Abort so
#handleRetryableError retries it without model fallback.
- Kept the deliberate-abort guards intact: user interrupts and silent aborts
carry their own markers (not the generic sentinel), and #abortInProgress /
#isDisposed / #streamingEditAbortTriggered still settle without retry.
- Rewrote the stale fallback test that froze the buggy no-retry behavior to
assert retry-and-recover for the error-stop sentinel.
Fixes#5375
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
- Permit model fallback even if the retry budget is exhausted when the current provider is locked by credential rotation or usage limits.
- Reset the retry budget when successfully switching to a fallback model to ensure the new model has a full allowance of retries.
- Added a regression test to verify that credential rotation failure triggers model fallback.
AgentSession defers and coalesces the wire-level agent_end while a
prompt is in flight (#emitSessionEvent), so a multi-attempt retry saga
often surfaces only ONE agent_end to EventController — which can be
the final settle, not an intermediate attempt. Consuming #retryPending
against whichever agent_end arrived first (previous commit) could
therefore discard the real final failure notification.
Switch to gating purely on the retry lifecycle: #retryPending is set
by auto_retry_start and cleared only by auto_retry_end (both
outcomes), never consumed by sendErrorNotification itself. Those
lifecycle events are never deferred, so they reliably bracket the
window a retry is actually outstanding regardless of how agent_end
coalescing lands.
Close the residual gap this creates: #handleRetryableError's
classifier-refusal and Fireworks-fallback-ineligible branches could
short-circuit a saga that already announced auto_retry_start without
ever emitting auto_retry_end, latching #retryPending open forever.
Both branches now emit a final auto_retry_end(false) when a prior
attempt already started the saga. #handleAgentStart also clears
#retryPending defensively so a saga that still somehow never resolves
cannot suppress a later, unrelated turn's notification.
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Allowed implicit default fallback resolution when other role fallback chains are configured.
- Covered the mixed-role first-run fallback case.
Fixes#4533
- Treated the active model as the default retry primary when only retry.fallbackChains.default is configured.
- Covered the first-run case where modelRoles.default is unset but a default fallback chain exists.
Fixes#4533
Passed the settled assistant message into session_stop emission so refusal-as-error turns can be pruned from replay context without hiding their stop details from extension hooks.
Expanded the refusal regression test to assert the session_stop payload still exposes the refusal as last_assistant_message.
Fixes#3591
Removed the early return after refusal pruning so the agent_end tail still reaches `#emitSessionStopEvent`, restoring `session_stop` extension hooks (block/continue/telemetry) for refusal-as-error stops.
Regression test wires an extensionRunner with a session_stop handler and asserts it fires for both the refusal turn and the following clean turn.
Fixes#3591
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
- Updated getApiKey signatures to accept a Model and return ApiKey or ApiKeyResolver.
- Updated stream key handling to resolve credentials per model and use seedApiKeyResolver for retries.
- Added antigravityEndpointMode setting with auto/production/sandbox endpoint selection.
- Added 429/5xx endpoint failover for Gemini stream, usage, search, and image calls.
Matched retry fallback roles against the plain model selector as well as the routed in-flight selector, preserving configured chains for compat-routed OpenRouter and Vercel models.
Added regression coverage for a compat-routed OpenRouter primary using a plain role selector.
Stopped retry fallback selector parsing from treating every @ suffix as upstream routing, preserving exact model ids like google-vertex Claude @default variants.
Resolved fallback candidates from raw selectors during preflight so routed selectors still work without corrupting exact at-suffixed ids.
Resolved retry fallback primaries from the raw selector during cooldown restore so OpenRouter and Vercel upstream pins survive fallback recovery.
Added regression coverage for routed OpenRouter primaries reverting after cooldown expiry.
Normalized max thinking aliases when recording and checking retry fallback cooldown suppressions while preserving live literal :max model IDs.\n\nFixes #2727