- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
Required CLI model registries to expose getAvailable() and used that authenticated
set whenever callers omit availableModels. Deferred SDK and bench/dry-balance
resolution now lets configured roles beat unauthenticated catalog id collisions.
Updated resolver test registries and made the #6508 regression omit the explicit
availableModels option, covering the deferred-caller path from the review.
Fixes#6508
- Normalized completed image_generation_call results into assistant image blocks.
- Persisted image bytes through the session blob store and rendered them in live, replay, ACP, proxy, telemetry, and HTML paths.
- Added response normalization, persistence, and TUI rendering regressions.
Fixes#4768
- Removed the `computeTokensPerSecond` helper function in favor of a direct calculation.
- Standardized throughput to use total request duration instead of post-TTFT decode time.
- Updated UI usage display to calculate tokens per second against total duration.
- Prevented over-inflation of throughput metrics caused by hidden reasoning tokens.
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
The model selector's persistence path dropped the `:auto` selector when parsing role values, producing a warning ('Invalid thinking level "auto"') and rendering the badge as `inherit` instead of `auto`. Reload of the default role also lost the auto state whenever the role value carried an explicit `:auto` suffix instead of relying on `defaultThinkingLevel`.
Widen the resolver chain (`parseThinkingSuffix`, `splitThinkingSuffix`, `parseModelString`, `parseModelPattern*`, `ResolvedModelRoleValue`, `ResolvedRoleModel`, `ResolveCliModelResult`) to carry the `AUTO_THINKING` sentinel end to end, and coerce it back to `undefined` at concrete-only boundaries (glob scope patterns, retry fallback, advisor, commit pipeline, guided-goal, bench).
Regression tests cover:
- `resolveModelRoleValue("provider/model:auto")` returns explicit auto without a warning.
- `ModelSelector` renders `DEFAULT (auto)` and `SMOL (auto)` when the role value has `:auto`.
- `cycleRoleModels` activates auto thinking on entering a `:auto` role.
- Startup resume activates auto thinking when `modelRoles.default` carries `:auto`.
Fixes#4128
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
- Added management of provider session states during benchmark execution.
- Implemented a teardown process to close and clear session states after request completion.
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
- Added --par flag to execute benchmark runs concurrently with a default degree of 4.
- Added --service-tier flag to allow overriding the provider service tier per benchmark.
- Increased default benchmark run count from 1 to 10 to provide more robust averaging.
- Updated benchmarking logic to process requests in a concurrency-limited pool while preserving output order.
- Implemented pre-flight credential checks to prevent unnecessary worker spawning when authentication is missing.
- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.