- Repointed tests off removed gpt-5.2/5.3 codex variants and devin models (e06ac0b787): context-promotion and TTSR tests pin gpt-5.5 -> gpt-5.6-sol via per-test modelOverrides since no bundled codex model has a runtime-effective promotion target anymore; replay-boundary and history-payload suites use gpt-5.5; advisor quota fallback uses devin/swe-1-6-slow with suffix-less selectors per #4579 devin-agent semantics; gateway-reference pins kilo/giga-potato and now asserts the reference carries effortRouting so the cross-provider no-inherit contract stays meaningful.
- Verified cross-provider gateway references do not inherit wire routing after variant collapse (identity/reference.ts:145-154) - stale test, no product bug.
- Exposed the active retry fallback selector from live agent sessions.
- Rendered fallback rows with an explicit marker and resolved provider/model.
- Added an end-to-end fallback-to-Agent-Hub regression assertion.
Fixes#6316
- Carried startup-selected fallback role and primary selector into AgentSession.
- Continued remaining role fallback entries after the startup fallback fails.
- Added regression coverage for chained startup failover.
Fixes#6283
- Added `#prunedTerminalRefusal` field to store the pruned refusal for post-settle consumers.
- Modified `getLastAssistantMessage()` to return the pruned refusal before active-context lookup.
- Reset `#prunedTerminalRefusal` on `agent_start` so a fresh run supersedes the settled refusal.
- Updated test mock helpers to include `getLastAssistantMessage` for consistency.
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:
- #failing latch: waitForCatchup resolves immediately while an advisor is
mid-failure; parked waiters wake the moment a turn fails, before any
async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
dedup snapshot and never propagates into the primary's turn-end callback
(per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
runtime latches quotaExhausted, requeues the batch, and notifies —
cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
before any assistant turn is recorded.
- Implemented parsing of id-prefixed wildcard keys and entries, allowing provider-specific prefixes in retry fallback configuration.
- Added logic to re-prefix failing model IDs and to match id-prefixed keys, with validation of provider existence.
- Updated settings schema description and changelog, and added tests covering the new behavior.
- Retained the advisor's original selector and thinking level while progressing through fallback candidates.
- Restored the configured primary before later advisor turns once its selector cooldown expired.
- Covered quota fallback restoration under the default cooldown-expiry policy.
- Switched advisor turns to the next configured model after provider quota or rate-limit failures.
- Emitted fallback applied and succeeded lifecycle events without reporting advisor unavailability after recovery.
- Added an end-to-end advisor quota fallback regression test.
Fixes#5740
A stalled or dropped provider stream that surfaces as stopReason:"error"
carrying the bare "Request was aborted" sentinel fell through both retry
gates: #isRetryableReasonlessAbort required stopReason:"aborted", and
#isRetryableError's classifier returns no retriable kinds for the generic
sentinel. The turn died immediately despite retry.enabled.
- Relaxed #isRetryableReasonlessAbort to accept an empty generic-abort
sentinel turn under stopReason "aborted" or "error", tagging it Abort so
#handleRetryableError retries it without model fallback.
- Kept the deliberate-abort guards intact: user interrupts and silent aborts
carry their own markers (not the generic sentinel), and #abortInProgress /
#isDisposed / #streamingEditAbortTriggered still settle without retry.
- Rewrote the stale fallback test that froze the buggy no-retry behavior to
assert retry-and-recover for the error-stop sentinel.
Fixes#5375
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
- Permit model fallback even if the retry budget is exhausted when the current provider is locked by credential rotation or usage limits.
- Reset the retry budget when successfully switching to a fallback model to ensure the new model has a full allowance of retries.
- Added a regression test to verify that credential rotation failure triggers model fallback.
AgentSession defers and coalesces the wire-level agent_end while a
prompt is in flight (#emitSessionEvent), so a multi-attempt retry saga
often surfaces only ONE agent_end to EventController — which can be
the final settle, not an intermediate attempt. Consuming #retryPending
against whichever agent_end arrived first (previous commit) could
therefore discard the real final failure notification.
Switch to gating purely on the retry lifecycle: #retryPending is set
by auto_retry_start and cleared only by auto_retry_end (both
outcomes), never consumed by sendErrorNotification itself. Those
lifecycle events are never deferred, so they reliably bracket the
window a retry is actually outstanding regardless of how agent_end
coalescing lands.
Close the residual gap this creates: #handleRetryableError's
classifier-refusal and Fireworks-fallback-ineligible branches could
short-circuit a saga that already announced auto_retry_start without
ever emitting auto_retry_end, latching #retryPending open forever.
Both branches now emit a final auto_retry_end(false) when a prior
attempt already started the saga. #handleAgentStart also clears
#retryPending defensively so a saga that still somehow never resolves
cannot suppress a later, unrelated turn's notification.
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Allowed implicit default fallback resolution when other role fallback chains are configured.
- Covered the mixed-role first-run fallback case.
Fixes#4533
- Treated the active model as the default retry primary when only retry.fallbackChains.default is configured.
- Covered the first-run case where modelRoles.default is unset but a default fallback chain exists.
Fixes#4533
Passed the settled assistant message into session_stop emission so refusal-as-error turns can be pruned from replay context without hiding their stop details from extension hooks.
Expanded the refusal regression test to assert the session_stop payload still exposes the refusal as last_assistant_message.
Fixes#3591
Removed the early return after refusal pruning so the agent_end tail still reaches `#emitSessionStopEvent`, restoring `session_stop` extension hooks (block/continue/telemetry) for refusal-as-error stops.
Regression test wires an extensionRunner with a session_stop handler and asserts it fires for both the refusal turn and the following clean turn.
Fixes#3591
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
- Updated getApiKey signatures to accept a Model and return ApiKey or ApiKeyResolver.
- Updated stream key handling to resolve credentials per model and use seedApiKeyResolver for retries.
- Added antigravityEndpointMode setting with auto/production/sandbox endpoint selection.
- Added 429/5xx endpoint failover for Gemini stream, usage, search, and image calls.
Matched retry fallback roles against the plain model selector as well as the routed in-flight selector, preserving configured chains for compat-routed OpenRouter and Vercel models.
Added regression coverage for a compat-routed OpenRouter primary using a plain role selector.
Stopped retry fallback selector parsing from treating every @ suffix as upstream routing, preserving exact model ids like google-vertex Claude @default variants.
Resolved fallback candidates from raw selectors during preflight so routed selectors still work without corrupting exact at-suffixed ids.
Resolved retry fallback primaries from the raw selector during cooldown restore so OpenRouter and Vercel upstream pins survive fallback recovery.
Added regression coverage for routed OpenRouter primaries reverting after cooldown expiry.
Normalized max thinking aliases when recording and checking retry fallback cooldown suppressions while preserving live literal :max model IDs.\n\nFixes #2727
Checked structured classifier refusals before the interrupted-output retry guard so provider refusals with explanatory content still use the configured fallback path.
Fixes#2683
- The AgentSession retry fallback test now validates the assistant message before reading it.
- It now verifies the first content block is text before checking the recovered message text.
v15.11.4 introduced stateful previous_response_id chaining on the
official OpenAI endpoint. The in-provider retry classifier matched only
the generic stale-id phrasing ('previous response ... not found |
invalid | expired | stale'), missing the Zero Data Retention 400
'Previous response cannot be used for this organization due to Zero
Data Retention.'. The error therefore bypassed the categorical-disable
path, so the chain was reset (not disabled), the next successful turn
re-armed it, and every other turn 400'd in a loop.
Add a dedicated isOpenAIResponsesZeroDataRetentionError detector and a
markOpenAIResponsesChainZeroDataRetention helper that disables chaining
on the first hit (skipping the three-strike circuit breaker). The
in-call retry now drops 'store: true' from the replay so the request is
semantically valid for ZDR orgs, and reasoning continuity is preserved
by the existing include: ['reasoning.encrypted_content'] flag.
AgentSession.#isStaleOpenAIResponsesReplayError gains the ZDR phrasing
too, so any ZDR error that does bubble past the provider retry resets
the Responses session and retries at zero backoff instead of falling
back to a different model.
Fixes#2341
Kept classifier refusals eligible for model fallback, but restored the retry.maxRetries guard so fallback chains cannot consume extra provider calls after the turn budget is exhausted.
Fixes#2290
Preserved Anthropic stop_details on assistant messages so the agent can distinguish classifier refusals from transport failures.
Taught AgentSession to use configured retry fallback chains for refusal and sensitive stops without same-model retries, then pin the fallback for the conversation.
Fixes#2290
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
Add retry.modelFallback so users can keep automatic retry enabled while preventing retry recovery from switching through configured fallback model chains.
The default remains enabled, preserving existing fallback behavior. When disabled, retry still honors retry-after delays and retry limits while staying on the primary model.
Op: correct
Restores: spec:retry can stay enabled without automatic model switching