- Normalized completed image_generation_call results into assistant image blocks.
- Persisted image bytes through the session blob store and rendered them in live, replay, ACP, proxy, telemetry, and HTML paths.
- Added response normalization, persistence, and TUI rendering regressions.
Fixes#4768
- Removed the `computeTokensPerSecond` helper function in favor of a direct calculation.
- Standardized throughput to use total request duration instead of post-TTFT decode time.
- Updated UI usage display to calculate tokens per second against total duration.
- Prevented over-inflation of throughput metrics caused by hidden reasoning tokens.
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
The model selector's persistence path dropped the `:auto` selector when parsing role values, producing a warning ('Invalid thinking level "auto"') and rendering the badge as `inherit` instead of `auto`. Reload of the default role also lost the auto state whenever the role value carried an explicit `:auto` suffix instead of relying on `defaultThinkingLevel`.
Widen the resolver chain (`parseThinkingSuffix`, `splitThinkingSuffix`, `parseModelString`, `parseModelPattern*`, `ResolvedModelRoleValue`, `ResolvedRoleModel`, `ResolveCliModelResult`) to carry the `AUTO_THINKING` sentinel end to end, and coerce it back to `undefined` at concrete-only boundaries (glob scope patterns, retry fallback, advisor, commit pipeline, guided-goal, bench).
Regression tests cover:
- `resolveModelRoleValue("provider/model:auto")` returns explicit auto without a warning.
- `ModelSelector` renders `DEFAULT (auto)` and `SMOL (auto)` when the role value has `:auto`.
- `cycleRoleModels` activates auto thinking on entering a `:auto` role.
- Startup resume activates auto thinking when `modelRoles.default` carries `:auto`.
Fixes#4128
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
- Added management of provider session states during benchmark execution.
- Implemented a teardown process to close and clear session states after request completion.
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
- Added --par flag to execute benchmark runs concurrently with a default degree of 4.
- Added --service-tier flag to allow overriding the provider service tier per benchmark.
- Increased default benchmark run count from 1 to 10 to provide more robust averaging.
- Updated benchmarking logic to process requests in a concurrency-limited pool while preserving output order.
- Implemented pre-flight credential checks to prevent unnecessary worker spawning when authentication is missing.
- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.