Files
oh-my-pi/packages/coding-agent/test/agent-session-retry-fallback.test.ts
T
enieuwy 323a3bc29c feat(core): add first-class retry fallback chains for model/provider failover (#541)
* feat(rules): implemented alwaysApply auto-injection into system prompt

rules with alwaysApply: true were parsed by all providers and used to
exclude the rule from rulebookRules, but the inclusion half was never
built — rule content was silently dropped. now:

- full content is injected directly into the system prompt (before the
  rulebook rules section) in both default and custom prompt templates
- rules remain addressable via rule:// for re-reading
- ttsr rules still take priority (condition + alwaysApply goes to ttsr only)

updated rulebook-matching-pipeline.md to reflect the three-bucket split
(ttsr > always-apply > rulebook) and corrected the rule:// resolution
docs.

* feat(browser): implement screenshot path saving

- Add `browser.screenshotDir` setting (tools tab) for a persistent
  default screenshot directory, configurable via /settings
- Honour the existing `path` parameter in the screenshot action,
  which was declared in the schema but never consumed by the implementation
- Resolution order: params.path (abs) > join(screenshotDir, params.path)
  > join(screenshotDir, screenshot-<timestamp>.png) > /tmp only
- Writes full-resolution buffer to disk (not the API-compressed copy)
- Creates destination directory recursively if it doesn't exist
- details.screenshotPath reflects the actual saved location

Fixes: path param silently ignored since removal in v11.5.0

* feat(browser): expand ~ in screenshotDir and path params

Users can now configure browser.screenshotDir as ~/Downloads or
~/Pictures and use params.path as ~/screenshots/foo.png without
needing to supply fully-qualified paths.

* fix(browser): write screenshot to exactly one location

Previously wrote to /tmp unconditionally then copied to user path,
resulting in two files. Now resolves a single destination upfront:
  1. params.path (absolute, or relative to screenshotDir/cwd)
  2. screenshotDir + auto-timestamp filename
  3. /tmp fallback (original behaviour, unchanged)

Also: user-defined destinations receive the full-res buffer;
/tmp fallback retains the API-compressed copy as before.

* feat(settings): add text input support for plain string settings

Settings with type: "string" and no submenu now render as an editable
text field in the /settings TUI panel instead of being silently skipped.

- settings-defs.ts: add TextInputSettingDef interface; pathToSettingDef
  falls through to { type: "text" } for any plain string schema entry
- settings-selector.ts: add TextInputSubmenu class (mirrors
  ConfigInputSubmenu from plugin-settings.ts); add "text" case to
  #defToItem; add #createTextInput method

This makes browser.screenshotDir visible and editable in the Tools tab.
Empty field on submit clears the setting (falls back to /tmp behavior).

* fix(settings): text input — cursor at end, block tab navigation

- Cursor: call handleInput(ctrl+e) after setValue to jump to end of
  pre-filled string instead of leaving it at position 0
- Tab/arrow guard: add #textInputActive flag; suppress tab-switch and
  left/right routing to tab bar while a TextInputSubmenu is open, so
  arrow keys reach the Input component's cursor movement handlers
  instead of switching settings tabs

* fix(browser): expand ~\ (Windows backslash) in screenshot paths

expandHome now handles both Unix ~/... and Windows ~\... separators,
matching user expectation on all platforms. Addresses Codex review.

* fix(coding-agent): use expandPath for screenshots

* docs: add changelog for screenshot path option

Bug fixes:
- Add browser.screenshotDir to settings-schema.ts (lost during rebase conflict)
- Fix stale expandHome comment to expandPath in settings-selector.ts
- Final screenshotDir description: Directory to save screenshots with ~ support

* fix: biome format corrections for screenshot-path-option branch

* fix: add missing #textInputActive field and screenshotDir default

* fix(browser): align screenshot metadata with saved file contents

When screenshotDir or params.path is set, the full-resolution PNG buffer
is written to disk. Previously mimeType/bytes in details still reflected
the resized payload sent to the model, making metadata inconsistent with
the actual saved file.

Now savedBuffer/savedMimeType track what is written, and details reflects
that. Display output distinguishes 'Saved' vs 'Model' when full-res is
used, and collapses to a single Format/Dimensions line for temp-only.

* fix(browser): resolve params.path relative to cwd regardless of screenshotDir

screenshotDir is a default save location, not an anchor for explicit
paths. A relative params.path should always resolve against cwd so its
semantics are stable and predictable regardless of user settings.

* fix(tui): enforce strict line budget for collapsed tool output

The grep, ast_grep, and ast_edit renderers used group-count-based
collapse that always included the first group unconditionally,
allowing collapsed output to remain visually large when a single
group contained many lines.

Add maxCollapsedLines to renderTreeList that enforces a strict
total-line cap in collapsed mode. Items that exceed the remaining
budget are skipped entirely (no broken fragments). The isLast tree
branch is computed after the budget check to avoid double-last
branches when a summary line follows.

Remove the per-tool getCollapsedMatchLimit / getCollapsedChangeLimit
helpers that are now redundant.

Fixes #455

Made-with: Cursor

* WIP

* WIP

* cleanup

* docs(coding-agent): update changelog for custom model tags and cycle order

* Fix custom model precedence across load and refresh

* feat(ask): add multiline editor support for custom input

* feat(extension-ui): add dialog options and abort signal support to editor

* docs(ask): add multiline editor and timeout behavior guidance

* docs(ask): simplify multiline input documentation

* fix(tui): preserve terminal scrollback during full redraws

Replace destructive \x1b[3J\x1b[2J\x1b[H full-redraw sequence with
scrollback-preserving repaint helpers:

- seedTranscript: first paint with no prior frame, writes full transcript
  without clearing scrollback
- repaintViewport: trusted-frame repaints that scroll viewport-shift
  delta into scrollback before overwriting visible rows in-place
- Height-increase handler that pushes revealed scrollback back before
  reclaiming the display

Replace requestRender(true) state destruction with a one-shot
without resetting #previousLines or cursor bookkeeping.

Track #previousHeight to detect terminal height increases that pull
scrollback lines into the visible area.

fixes #507

* fix(tui): fix exit gaps, content shrink drift, and overlay cursor recovery

- Rewrite stop() to use viewport-relative cursor positioning instead of
  content-length, preventing blank gaps when content is shorter than viewport
- Replace trailing \r\n\x1b[2K clear loops in repaintViewport() and
  height-increase path with \r\n\x1b[J to avoid cursor drift past content
- Add 12 TUI regression tests for exit gaps, content shrink, and overlay
  dismiss cursor recovery
- Add 6 coding-agent controller tests for /new and /tree commands

* fix(tui): redraw sparse height increases atomically

* fix(coding-agent,tui): address review regressions

* fix(tui): reseed after terminal resume

* test(ai): avoid leaking kagi module mocks

* fix(ask): keep multiline custom input in prompt gutter

* fix(ask): preserve multiselect choices on editor dismiss

* fix(ask): honor app interrupt in prompt editor

* fix(coding-agent): preserve new-session approval state

* feat(core): add session-scoped model/provider retry fallback policy

* feat(core): validate retry fallback chains on session startup

* fix(ask): preserve prior answer when custom editor dismissed in single-select

* fix(core): harden retry fallback policy semantics

* fix(core): correct retry fallback edge cases

* test(coding-agent): removed macOS fallback test case from theme detection

- Removed test case for macOS fallback behavior inside Zellij.

* refactor: simplified null checks using optional chaining across TypeScript and Rust modules

- Simplified null/empty checks across TypeScript codebase using optional chaining operator (?.) for improved readability.
- Replaced explicit null checks in validation logic with optional chaining in oauth-discovery, gemini-cli, claude, zai, and lsp modules.
- Updated error handling in Rust command invocation to use double question mark operator (??) for cmd_result.
- Consolidated null validation patterns across tools (bash-skill-urls, browser, gemini-image, resolve) and keybindings using optional chaining.

* chore: bump version to 13.15.1

* fix(core): address PR review on fallback retry behavior

* test(ai): refactored auth storage test to use spyOn for cleaner mocks

- Refactored auth storage test to use vi.spyOn() instead of vi.mock() for cleaner mock management.
- Simplified mock type definitions by leveraging bun:test's Mock type import.

* chore: bump version to 13.15.2

* Fix stale OpenAI Responses replay across session boundaries (#534)

* Fix stale OpenAI Responses replay across session boundaries

Fixes #505

* Fix CI tests for session replay change

* Harden session reload and switch rollback

* Guard session switch snapshots

* fix(coding-agent): preserve responses replay snapshots

* Normalize pasted image formats before attach (#543)

Co-authored-by: iter <itertoolz@gmail.com>

* Allow overriding the Codex web search model (#516)

* Allow codex web search model override

* Handle blank codex web search model

* feat(ai): add gemini-3.1-pro-preview models to google-vertex provider (#521)

Add gemini-3.1-pro-preview and gemini-3.1-pro-preview-customtools to
the google-vertex provider in models.json, matching the existing
google-generative-ai entries.

Fixes #520

Co-authored-by: Muness Castle <munesscastle@artium.ai>

* Make temporary model selector keybinding configurable (#539)

Fixes #533

* chore: bump models.json

* refactor(coding-agent): migrated test mocks to vitest spyOn with Symbol.dispose cleanup

- Migrated test mocking from bun:test mock API to vitest spyOn pattern across 9 test files.
- Extracted mock setup logic into reusable helper functions with Symbol.dispose cleanup pattern.
- Replaced manual beforeEach/afterEach and try-finally blocks with TypeScript 5.2 using declarations.
- Removed 178 lines of boilerplate mock initialization and restoration code from test suite.

* chore: bump version to 13.15.3

* fix(ask): restore prompt-style enter handling

* fix(tui): avoid blank scrollback regressions

* fix(models): keep same-id replacements authoritative

* fix(tui): enforce collapsed line budgets

* fix(models): unify selector role sources

* fix(rules): dedupe always-apply prompt injection

* fix(browser): show saved screenshot path

* style: format merged PR fixes

* refactor(coding-agent): restructured screenshot and prompt utilities into focused helpers

- Extracted screenshot formatting logic into dedicated `formatScreenshot()` function with options support.
- Consolidated prompt source deduplication into `dedupePromptSource()` helper to prevent rule duplication.
- Refactored editor text sanitization to use `replaceTabs()` utility for consistent tab width handling.
- Added test coverage verifying editor respects configured tab width when loading text programmatically.

* refactor(coding-agent): restructured validation and rendering for consistency

- Refactored theme color validation to use single source of truth with THEME_COLOR_RECORD object.
- Simplified model registry to defer per-model overrides to dedicated method and use constant for role IDs.
- Refactored tree list rendering to pre-render items once for consistent line counts across phases.
- Refactored question result formatting to use early returns and consistently include question ID in output.
- Updated hook editor hint text to include ctrl+g external editor option when prompt style is enabled.
- Removed unused isLogicalLineStart property from LayoutLine interface in editor component.

* revert: 535 due to TUI regressions

* test(coding-agent): corrected hook-editor assertion for keybinding render

- Corrected assertion in hook-editor test to verify external editor keybinding is rendered.

* feat(tools): added root path alias to resolve bare / to working directory

- Added root path alias feature to resolve bare `/` to session working directory in path resolution.
- Updated browser tool to use `resolveToCwd()` for consistent workspace-relative path handling.
- Added comprehensive test suite validating root path alias resolution across grep, read, find, ast_grep, and ast_edit tools.

* feat(prompts/tools): clarified hashline block boundary handling with examples

- Improved hashline tool documentation with clearer guidance on block boundary handling and closing delimiter duplication prevention.
- Added concrete example demonstrating correct anchor placement when replacing entire blocks including closing braces.
- Reorganized boundary duplication warnings into actionable self-check guidance with visual comparison steps.

* chore: bump version to 13.16.0

* fix(coding-agent): fixed python kernel startup hangs (#548)

* fix(coding-agent): fixed python kernel startup hangs

* fix(coding-agent): fixed startup timeout regressions

* fix(coding-agent): preserved startup cancellation typing

* perf(pi-natives): optimized memory allocation with MiMalloc integration

- Integrated MiMalloc as global allocator to improve memory allocation performance.

* feat: fff

- Added SearchDb class for stateful shared search database instances enabling persistent file indexing and frecency tracking across grep, glob, and fuzzyFind operations.
- Added optional db parameter to grep(), glob(), and fuzzyFind() functions for database-backed searching with improved performance via cached file indices.
- Replaced grep-searcher with fff-grep and added fff-search dependency for enhanced file discovery and search capabilities with memory-mapped file support.
- Migrated fuzzy file discovery from fd module to fff module with SearchDb integration for stateful caching and improved search performance.
- Exported SearchDb type from @oh-my-pi/pi-natives public API for type-safe usage in grep, glob, and fuzzyFind workflows.

* feat(pi-natives): added unified picker coordination for file search operations

- Added `wait_for_picker_scan()` function to search_db module with cancellation token support for polling picker scan completion.
- Integrated picker-based file search into glob matching logic with `collect_files_from_picker()` helper to reuse shared SearchDb picker results.
- Refactored fff and grep modules to use centralized `wait_for_picker_scan()` wrapper instead of direct FilePicker calls, improving cancellation handling.
- Propagated SearchDb instance through agent session initialization and input controller to enable unified file picker coordination across search operations.

* chore: bump version to 13.16.1

* fix: install zig in CI

* feat(tui): make inline image max-width configurable via tui.maxInlineImageColumns (#551)

* feat(tui): make inline image max-width configurable via tui.maxInlineImageColumns

* fix(tui): handle 0 as unlimited in maxInlineImageColumns; drop || undefined coercion

* feat(browser): auto-detect NixOS and use system Chromium (#550)

Puppeteer's bundled Chromium is a dynamically-linked FHS binary that
cannot run on NixOS. On startup, resolveSystemChromium() checks for
/etc/NIXOS and searches for a usable binary in order:

  1. chromium on PATH
  2. chromium-browser on PATH
  3. ~/.nix-profile/bin/chromium
  4. /run/current-system/sw/bin/chromium

The resolved path is passed as executablePath to puppeteer.launch().
Result is cached per process. On non-NixOS systems the function returns
undefined immediately, leaving Puppeteer's default resolution intact.

* feat(pi-natives): added automatic parenthesis escaping in regex patterns

- Added automatic escaping of unescaped parentheses in regex patterns when group syntax errors occur, enabling literal function call patterns like `fetchAnthropicProvider(` to work as search queries.
- Extracted regex matcher builder into separate function for reusability and error recovery logic.
- Added 2 test cases validating parenthesis escaping behavior for both escaped and literal parentheses.
- Fixed documentation formatting in sanitize_braces comment.

* style: reformat

* chore: bump version to 13.16.2

* fix: only show update banner when npm version is strictly newer (#552)

* fix(ai): corrected OAuth credential updates to replace in-place instead of accumulating soft-deleted rows

- Fixed OAuth credential updates to replace matching credentials in-place rather than creating disabled rows, preventing unbounded accumulation of soft-deleted credentials.
- Modified OAuth credential saving to preserve unrelated identities instead of replacing all credentials for a provider.
- Updated credential identity resolution to use provider context for more accurate email deduplication.
- Implemented upsertAuthCredentialForProvider method to handle credential matching and in-place updates.
- Added 5 test cases covering credential preservation across reauth, multi-account scenarios, and stale cache handling.

* chore: bump version to 13.16.3

* feat: introduced unified range API for hashline edits and model catalog updates

- Simplified hashline edit location API by replacing separate `line` and `block` properties with unified `range` property accepting `{ pos, end }` anchors.
- Renamed hashline helper functions from `hlineref`/`hlinefull` to `href`/`hline` for improved brevity and consistency.
- Added detection for `kysely-codegen` generated files in auto-generated file guard with corresponding test coverage.
- Added 13 new AI model configurations and updated token limits and pricing for existing models across multiple providers.
- Enhanced file type validation in grep native to reject symlinks, FIFOs, sockets, and non-regular files with improved error handling.

* chore: bump version to 13.16.4

* fix: pin rustc-hash to 2.1.1 to avoid SIGILL on CI

rustc-hash 2.1.2 (released today) refactored hash_bytes to use
split_first_chunk, which produces illegal instructions when compiled
with nightly + -C target-cpu=x86-64-v3 on CI runners.

* fix(ci): pin nightly to 2026-03-27 to avoid codegen SIGILL regression

Today's nightly produces illegal instructions when compiled with
-C target-cpu=x86-64-v3. Reverts the unnecessary rustc-hash pin from
the previous commit since the real cause is the nightly compiler.

Also lets Cargo.lock return to rustc-hash 2.1.2 (not the culprit).

* fix(ci): add rustup target fallback for pinned nightly cross-compile

* fix(coding-agent): do not prompt to use grep and find tools if they are disabled (#566)

Co-authored-by: le-cameleon <200889489+le-cameleon@users.noreply.github.com>

* fix(natives): skipped grep special files (#565)

avoided opening fifos and other special filesystem nodes during grep and added fifo regressions in native and coding-agent tests.

* fix: skill baseDir regex fails on Windows backslash paths (#554)

The regex that strips SKILL.md from the path to compute baseDir only
matches forward slashes. On Windows where paths use backslashes, the
replace is a no-op and baseDir equals the full SKILL.md file path.

This breaks sub-path resolution for skills: the subpath gets appended
to the SKILL.md file path instead of the skill directory.

Fix: use character class matching both path separators.

---------

Co-authored-by: deadcode-walker <268043493+deadcode-walker@users.noreply.github.com>
Co-authored-by: Rens Tillmann <rens@super-forms.com>
Co-authored-by: haiyang.zhou <haiyang.zhou@seamoney.com>
Co-authored-by: Leo P <junk@slact.net>
Co-authored-by: Zakhar Kogan <36503576+zaharkogan@users.noreply.github.com>
Co-authored-by: Vu Anh Nguyen <vuanhng00@gmail.com>
Co-authored-by: can1357 <me@can.ac>
Co-authored-by: daandden <64765666+daandden@users.noreply.github.com>
Co-authored-by: iter <72358817+itertea@users.noreply.github.com>
Co-authored-by: iter <itertoolz@gmail.com>
Co-authored-by: Cheol Kang <dev@cheol.me>
Co-authored-by: Muness Castle <931+muness@users.noreply.github.com>
Co-authored-by: Muness Castle <munesscastle@artium.ai>
Co-authored-by: zamo <falby97@proton.me>
Co-authored-by: elikoga <elikowa@gmail.com>
Co-authored-by: BayLee4 <63376748+BayLee4@users.noreply.github.com>
Co-authored-by: le-cameleon <200889489+le-cameleon@users.noreply.github.com>
Co-authored-by: Wiedzmin <56316383+art-wiedzmin@users.noreply.github.com>
2026-03-29 17:50:36 +02:00

390 lines
14 KiB
TypeScript

import { afterEach, beforeEach, describe, expect, it } from "bun:test";
import * as path from "node:path";
import { Agent } from "@oh-my-pi/pi-agent-core";
import { type AssistantMessage, Effort, getBundledModel, type Model } from "@oh-my-pi/pi-ai";
import { AssistantMessageEventStream } from "@oh-my-pi/pi-ai/utils/event-stream";
import { ModelRegistry } from "@oh-my-pi/pi-coding-agent/config/model-registry";
import { Settings } from "@oh-my-pi/pi-coding-agent/config/settings";
import { AgentSession, type AgentSessionEvent } from "@oh-my-pi/pi-coding-agent/session/agent-session";
import { AuthStorage } from "@oh-my-pi/pi-coding-agent/session/auth-storage";
import { SessionManager } from "@oh-my-pi/pi-coding-agent/session/session-manager";
import { TempDir } from "@oh-my-pi/pi-utils";
class MockAssistantStream extends AssistantMessageEventStream {}
function createAssistantMessage(
model: Model,
options: { text?: string; stopReason: "stop" | "error"; errorMessage?: string },
): AssistantMessage {
return {
role: "assistant",
content: options.text ? [{ type: "text", text: options.text }] : [],
api: model.api,
provider: model.provider,
model: model.id,
usage: {
input: 0,
output: 0,
cacheRead: 0,
cacheWrite: 0,
totalTokens: 0,
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 },
},
stopReason: options.stopReason,
errorMessage: options.errorMessage,
timestamp: Date.now(),
};
}
async function _waitFor(predicate: () => boolean, timeoutMs = 1000): Promise<void> {
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
if (predicate()) return;
await Bun.sleep(10);
}
throw new Error("Timed out waiting for condition");
}
describe("AgentSession retry fallback", () => {
let tempDir: TempDir;
let authStorage: AuthStorage;
let modelRegistry: ModelRegistry;
let session: AgentSession | undefined;
beforeEach(async () => {
tempDir = TempDir.createSync("@pi-retry-fallback-");
authStorage = await AuthStorage.create(path.join(tempDir.path(), "testauth.db"));
authStorage.setRuntimeApiKey("anthropic", "anthropic-test-key");
authStorage.setRuntimeApiKey("openai", "openai-test-key");
modelRegistry = new ModelRegistry(authStorage);
});
afterEach(async () => {
if (session) {
await session.dispose();
session = undefined;
}
authStorage.close();
tempDir.removeSync();
});
it("advances through a role-keyed fallback chain across retries", async () => {
const primaryModel = getBundledModel("anthropic", "claude-sonnet-4-5");
const firstFallback = getBundledModel("openai", "gpt-4o-mini");
const secondFallback = getBundledModel("openai", "gpt-4o");
if (!primaryModel || !firstFallback || !secondFallback) {
throw new Error("Expected bundled test models to exist");
}
const requestedModels: string[] = [];
const retryStartEvents: Array<Extract<AgentSessionEvent, { type: "auto_retry_start" }>> = [];
const retryEndEvents: Array<Extract<AgentSessionEvent, { type: "auto_retry_end" }>> = [];
const fallbackAppliedEvents: Array<Extract<AgentSessionEvent, { type: "retry_fallback_applied" }>> = [];
const fallbackSucceededEvents: Array<Extract<AgentSessionEvent, { type: "retry_fallback_succeeded" }>> = [];
const agent = new Agent({
getApiKey: provider => `${provider}-test-key`,
initialState: {
model: primaryModel,
systemPrompt: "Test",
tools: [],
messages: [],
},
streamFn: model => {
requestedModels.push(`${model.provider}/${model.id}`);
const stream = new MockAssistantStream();
queueMicrotask(() => {
if (model.provider === primaryModel.provider && model.id === primaryModel.id) {
const message = createAssistantMessage(model, {
stopReason: "error",
errorMessage: "overloaded_error: provider returned error 503",
});
stream.push({ type: "start", partial: message });
stream.push({ type: "error", reason: "error", error: message });
return;
}
if (model.provider === firstFallback.provider && model.id === firstFallback.id) {
const message = createAssistantMessage(model, {
stopReason: "error",
errorMessage: "service unavailable: 503 overloaded",
});
stream.push({ type: "start", partial: message });
stream.push({ type: "error", reason: "error", error: message });
return;
}
if (model.provider === secondFallback.provider && model.id === secondFallback.id) {
const message = createAssistantMessage(model, {
text: "Recovered on second fallback",
stopReason: "stop",
});
stream.push({
type: "start",
partial: createAssistantMessage(model, { text: "", stopReason: "stop" }),
});
stream.push({ type: "done", reason: "stop", message });
return;
}
throw new Error(`Unexpected model requested during retry fallback test: ${model.provider}/${model.id}`);
});
return stream;
},
});
const settings = Settings.isolated({
"compaction.enabled": false,
"retry.baseDelayMs": 5,
"retry.fallbackChains": {
default: [
`${firstFallback.provider}/${firstFallback.id}`,
`${secondFallback.provider}/${secondFallback.id}`,
],
},
});
settings.setModelRole("default", `${primaryModel.provider}/${primaryModel.id}`);
session = new AgentSession({
agent,
sessionManager: SessionManager.inMemory(),
settings,
modelRegistry,
});
session.subscribe(event => {
if (event.type === "auto_retry_start") {
retryStartEvents.push(event);
}
if (event.type === "auto_retry_end") {
retryEndEvents.push(event);
}
if (event.type === "retry_fallback_applied") {
fallbackAppliedEvents.push(event);
}
if (event.type === "retry_fallback_succeeded") {
fallbackSucceededEvents.push(event);
}
});
await session.prompt("Recover from rate limits");
await session.waitForIdle();
expect(requestedModels).toEqual([
`${primaryModel.provider}/${primaryModel.id}`,
`${firstFallback.provider}/${firstFallback.id}`,
`${secondFallback.provider}/${secondFallback.id}`,
]);
expect(session.model?.provider).toBe(secondFallback.provider);
expect(session.model?.id).toBe(secondFallback.id);
expect(retryStartEvents.map(event => event.delayMs)).toEqual([0, 0]);
expect(fallbackAppliedEvents).toEqual([
{
type: "retry_fallback_applied",
from: `${primaryModel.provider}/${primaryModel.id}`,
to: `${firstFallback.provider}/${firstFallback.id}`,
role: "default",
},
{
type: "retry_fallback_applied",
from: `${firstFallback.provider}/${firstFallback.id}`,
to: `${secondFallback.provider}/${secondFallback.id}`,
role: "default",
},
]);
expect(retryEndEvents).toHaveLength(1);
expect(retryEndEvents[0]).toMatchObject({ success: true, attempt: 2 });
expect(fallbackSucceededEvents).toEqual([
{
type: "retry_fallback_succeeded",
model: `${secondFallback.provider}/${secondFallback.id}`,
role: "default",
},
]);
});
it("suppresses cooled selectors and lazily reverts to the role primary after cooldown expiry", async () => {
const primaryModel = getBundledModel("anthropic", "claude-sonnet-4-5");
const fallbackModel = getBundledModel("openai", "gpt-4o-mini");
if (!primaryModel || !fallbackModel) {
throw new Error("Expected bundled test models to exist");
}
const requestedModels: string[] = [];
let primaryAttempts = 0;
const agent = new Agent({
getApiKey: provider => `${provider}-test-key`,
initialState: {
model: primaryModel,
systemPrompt: "Test",
tools: [],
messages: [],
},
streamFn: model => {
requestedModels.push(`${model.provider}/${model.id}`);
const stream = new MockAssistantStream();
queueMicrotask(() => {
if (model.provider === primaryModel.provider && model.id === primaryModel.id && primaryAttempts === 0) {
primaryAttempts += 1;
const message = createAssistantMessage(model, {
stopReason: "error",
errorMessage: "rate limit exceeded retry-after-ms=200",
});
stream.push({ type: "start", partial: message });
stream.push({ type: "error", reason: "error", error: message });
return;
}
const message = createAssistantMessage(model, {
text: `ok:${model.provider}/${model.id}`,
stopReason: "stop",
});
stream.push({ type: "start", partial: createAssistantMessage(model, { text: "", stopReason: "stop" }) });
stream.push({ type: "done", reason: "stop", message });
});
return stream;
},
});
const settings = Settings.isolated({
"compaction.enabled": false,
"retry.baseDelayMs": 5,
"retry.fallbackChains": {
default: [`${fallbackModel.provider}/${fallbackModel.id}`],
},
"retry.fallbackRevertPolicy": "cooldown-expiry",
});
settings.setModelRole("default", `${primaryModel.provider}/${primaryModel.id}`);
session = new AgentSession({
agent,
sessionManager: SessionManager.inMemory(),
settings,
modelRegistry,
});
await session.prompt("First prompt triggers fallback");
await session.waitForIdle();
expect(requestedModels).toEqual([
`${primaryModel.provider}/${primaryModel.id}`,
`${fallbackModel.provider}/${fallbackModel.id}`,
]);
expect(session.model?.provider).toBe(fallbackModel.provider);
expect(session.model?.id).toBe(fallbackModel.id);
await session.prompt("Immediate second prompt should stay on fallback");
await session.waitForIdle();
expect(requestedModels).toEqual([
`${primaryModel.provider}/${primaryModel.id}`,
`${fallbackModel.provider}/${fallbackModel.id}`,
`${fallbackModel.provider}/${fallbackModel.id}`,
]);
expect(session.model?.provider).toBe(fallbackModel.provider);
expect(session.model?.id).toBe(fallbackModel.id);
await Bun.sleep(240);
await session.prompt("Third prompt should lazily revert to primary");
await session.waitForIdle();
expect(requestedModels).toEqual([
`${primaryModel.provider}/${primaryModel.id}`,
`${fallbackModel.provider}/${fallbackModel.id}`,
`${fallbackModel.provider}/${fallbackModel.id}`,
`${primaryModel.provider}/${primaryModel.id}`,
]);
expect(session.model?.provider).toBe(primaryModel.provider);
expect(session.model?.id).toBe(primaryModel.id);
});
it("preserves thinking on bare fallback selectors and does not overwrite user thinking on restore", async () => {
const primaryModel = getBundledModel("anthropic", "claude-sonnet-4-5");
const fallbackModel = getBundledModel("openai", "gpt-4o-mini");
if (!primaryModel || !fallbackModel) {
throw new Error("Expected bundled test models to exist");
}
const requestedModels: string[] = [];
let primaryAttempts = 0;
const agent = new Agent({
getApiKey: provider => `${provider}-test-key`,
initialState: {
model: primaryModel,
systemPrompt: "Test",
tools: [],
messages: [],
},
streamFn: model => {
requestedModels.push(`${model.provider}/${model.id}`);
const stream = new MockAssistantStream();
queueMicrotask(() => {
if (model.provider === primaryModel.provider && model.id === primaryModel.id && primaryAttempts === 0) {
primaryAttempts += 1;
const message = createAssistantMessage(model, {
stopReason: "error",
errorMessage: "rate limit exceeded retry-after-ms=200",
});
stream.push({ type: "start", partial: message });
stream.push({ type: "error", reason: "error", error: message });
return;
}
const message = createAssistantMessage(model, {
text: `ok:${model.provider}/${model.id}`,
stopReason: "stop",
});
stream.push({ type: "start", partial: createAssistantMessage(model, { text: "", stopReason: "stop" }) });
stream.push({ type: "done", reason: "stop", message });
});
return stream;
},
});
const settings = Settings.isolated({
"compaction.enabled": false,
"retry.baseDelayMs": 5,
"retry.fallbackChains": {
default: [`${fallbackModel.provider}/${fallbackModel.id}`],
},
"retry.fallbackRevertPolicy": "cooldown-expiry",
});
settings.setModelRole("default", `${primaryModel.provider}/${primaryModel.id}:high`);
session = new AgentSession({
agent,
sessionManager: SessionManager.inMemory(),
settings,
modelRegistry,
thinkingLevel: Effort.High,
});
await session.prompt("First prompt triggers bare-selector fallback");
await session.waitForIdle();
expect(requestedModels).toEqual([
`${primaryModel.provider}/${primaryModel.id}`,
`${fallbackModel.provider}/${fallbackModel.id}`,
]);
expect(session.model?.provider).toBe(fallbackModel.provider);
expect(session.model?.id).toBe(fallbackModel.id);
expect(session.thinkingLevel).toBeUndefined();
session.setThinkingLevel(Effort.Low);
await Bun.sleep(240);
await session.prompt("Second prompt should restore model but preserve user thinking change");
await session.waitForIdle();
expect(requestedModels).toEqual([
`${primaryModel.provider}/${primaryModel.id}`,
`${fallbackModel.provider}/${fallbackModel.id}`,
`${primaryModel.provider}/${primaryModel.id}`,
]);
expect(session.model?.provider).toBe(primaryModel.provider);
expect(session.model?.id).toBe(primaryModel.id);
expect(session.thinkingLevel).toBeUndefined();
});
it("normalizes suppression by base selector and clears it on model refresh", async () => {
const future = Date.now() + 60_000;
modelRegistry.suppressSelector("openai/gpt-4o:high", future);
expect(modelRegistry.isSelectorSuppressed("openai/gpt-4o")).toBe(true);
expect(modelRegistry.isSelectorSuppressed("openai/gpt-4o:low")).toBe(true);
await modelRegistry.refresh("offline");
expect(modelRegistry.isSelectorSuppressed("openai/gpt-4o")).toBe(false);
});
});