Required CLI model registries to expose getAvailable() and used that authenticated
set whenever callers omit availableModels. Deferred SDK and bench/dry-balance
resolution now lets configured roles beat unauthenticated catalog id collisions.
Updated resolver test registries and made the #6508 regression omit the explicit
availableModels option, covering the deferred-caller path from the review.
Fixes#6508
`resolveCliModel` ran findExactCliModel's unauthenticated catalog fallback
before configured-role resolution, so the bundled `cursor/default` model (bare
id `default`) shadowed a configured `modelRoles.default`. On machines without
Cursor credentials `--model default` failed with `No API key found for cursor`
instead of resolving the configured, authenticated default role.
Defer the catalog fallback: explicit provider/id references and authenticated
bare ids still win over roles, a configured role now beats an unauthenticated
catalog-only id, and the catalog id is still recovered via the trailing fuzzy
fallback when no role matches.
Fixes#6508
Flat aggregator ids whose prefix collides with a provider slug (e.g.
openai/gpt-oss-120b hosted on OpenRouter) bypassed the authenticated-model
preference: the explicit-provider branch searched the full catalog only.
Keep provider/id exact references authoritative, but let the flat-id
fallback prefer authenticated providers before catalog order.
- Carried configured role identity through deferred CLI model resolution.
- Consulted ordered authenticated role fallbacks after unavailable primaries.
- Added startup regression coverage for missing primary and fallback entries.
Fixes#6283
Made startup resolution rank configured providers before catalog-only matches while preserving explicit provider pins and catalog fallback behavior.
Added regression coverage for shared bare model IDs across authenticated and unauthenticated providers.
Fixes#6150
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
parseModelPatternWithContext ran matchModel on the whole pattern
(including a trailing :level thinking suffix) and only stripped the
suffix if that first pass missed. matchModel's provider-scoped fuzzy
match normalizes colons away and does subsequence matching, so
kimi-for-coding:high matched the longer sibling kimi-for-coding-highspeed
before the suffix was recognized as a thinking level, silently switching
model and billing tier.
Match the full pattern exactly first (new exactOnly mode skips the
fuzzy/substring fallbacks), then strip a valid :level suffix and recurse
before any fuzzy match; fuzzy-match the whole pattern only as a last
resort. Literal ids ending in :max still win via the exact pass.
Fixes#5151
The pre-flight auth check in resolveModelOverrideWithAuthFallback called
getApiKey without a session id. For providers with session-sticky OAuth
credentials, this returned undefined even though the credential was
usable once the subagent session started, causing the auth fallback to
silently replace the configured model with the parent's (#5325).
The subagent's id is now forwarded as the session id so session-sticky
credentials resolve during the pre-flight check. Genuinely broken auth
(stale OAuth, revoked tokens) still falls back as before.
Also propagate model resolution warnings through resolveModelOverride
and log them in the executor so users see why a pattern didn't match.
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Enabled comma-separated string splitting within array-based model patterns.
- Expanded deferred model patterns before registration to align with immediate resolution behavior.
- Unified resolution logic to ensure deferred patterns support the same role aliases and chaining as the standard path.
main independently absorbed batch-1 (incremental grapheme slice 718c7cea2, markdown
stream-prefix cache 705426548 + 3822a83b4) and the pathTo/patch.ts items; restore
main's refined versions wholesale. Port the model-resolver optimization onto main's
resolver shape: hoist per-candidate case folds in matchModel and build the
preference context once per role resolution (matchPatternWithContext) instead of
per fallback pattern. Rewrite changelog entries to the surviving items only.
Five hot-path performance fixes + two low-risk allocation reductions.
No behavior change; all derived counts/orderings are identical.
- session-manager pathTo: leaf->root walk used branch.unshift() per node
(O(n^2) over branch length); now push + single reverse. Backs
getBranch(), hit at ~17 sites per turn.
- edit/modes/patch: collapseConsecutiveSharedLines / collapseRepeatedBlocks
/ trimCommonContext built shared-line sets via
new Set(oldLines.filter(l => newLines.includes(l))) -> O(old*new) per
hunk. Precompute new Set(newLines) and use .has() -> O(old+new).
- task/executor appendRecentOutputTail: re-split + filter + slice + reverse
of the full (up to 8KB) recentOutputTail on every text_delta token. Fast
path extends the current last line in place; full recompute only when a
newline boundary or truncation changes the window. tailLastLineRepresentable
flag guards the trailing-whitespace-only-line edge case.
- task/render renderResult: header booleans (3x .some) + footer counts
(3x .filter) + request total (.reduce) re-scanned details.results ~30x/sec
via the spinner. Single pass derives aborted/failed/mergeFailed/success
counts + requestTotal; booleans derived from counts.
- task/render extractIncrementalReviewResult: re-called normalizeYieldData
internally though both callers had already normalized the same yield data.
Signature now takes pre-normalized RenderYieldItem[].
Honorable mentions (allocation reduction, no algorithmic change):
- config/model-resolver: hoist case-folded pattern out of matchModel filter
passes; build the O(n) preference context once per role in
resolveModelRoleValue and reuse across fallback patterns.
- tools/read countTextLines: count newlines directly instead of allocating
via split("\n"); hashline formatter reuses the line count instead of
recomputing.
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
- Added `resolveAdvisorRoleSelection` to handle the advisor role's distinct configuration logic.
- Implemented `rolePriorityDefaults` to allow the advisor role to alias the `slow` model priority chain.
- Updated `AgentSession` to utilize the new advisor-specific resolution logic during session operations.
getOpenRouterRouteSuffix() used strict parseThinkingLevel(), which never
recognizes the max->xhigh alias, so openrouter/<id>:max was consumed as an
OpenRouter route suffix and cloned into a literal <id>:max model id with the
reasoning level dropped. This hit every exact-selector funnel
(parseModelPattern -> resolveCliModel/--model, resolveModelRoleValue/modelRoles
+ model picker, SDK default role) for the dominant aggregator provider.
Exclude max via parseThinkingSuffix(.., MAX_THINKING_SUFFIX_OPTIONS) so the
pattern falls through to the existing max-aware selector split. Literal :max
ids stay safe (none exist under openrouter; nanogpt literals win via exact
lookup before this path). Adds an openrouter/<id>:max regression test.
Kept synthetic Bedrock inference profile ARN models in enabledModels and SDK scope filtering instead of dropping them when returning available models.
Fixes#3004
Resolved Amazon Bedrock application inference profile ARNs through the provider-specific model resolver and routed Bedrock requests to the ARN region.
Fixes#3004
- Removed the Codex-preferred canonical remapping path for exact `provider/id` model matches.
- Resolved explicit `openai/gpt-5.5` references as `openai` provider without redirecting to `openai-codex`.
- Added regression tests to confirm explicit provider/id and enabled-model patterns are not coalesced to Codex.
Limit provider-priority ranking to provider defaults that share the first fallback default id, preserving mixed-provider startup precedence while keeping the OpenAI/Codex tie fix.\n\nFixes #2807
Route stale OpenAI GPT default roles through canonical Codex selection when the catalog prefers the Codex OAuth transport, and rank shared provider defaults by canonical provider priority.\n\nFixes #2807
Install a subagent's ordered model candidates as child-session retry fallback chains so a retryable provider failure advances to the next candidate instead of killing the worker (issue #2750).
The routing bypass in matchModel() returned early for any provider whose
post-slash id contained a valid @slug, so fuzzy provider-qualified patterns
over non-aggregator ids that legitimately end in @ (e.g. google-vertex/opus@default
-> claude-opus-4-8@default) resolved to nothing. Gate the bypass on
providerModels.some(supportsUpstreamRouting) so only OpenRouter / Vercel Gateway
short-circuit to the routing fallback. Addresses Codex review feedback on #2710.
Kept OpenRouter and Vercel upstream routing suffixes in subagent retry fallback selectors so same-base routed candidates stay distinct.
Resolved retry fallback candidates from the raw selector before model switching so routed fallback models keep their requested upstream route.
Avoided carrying max as an explicit thinking selector from literal provider/model role values while preserving max suffixes on pi role aliases and non-literal selectors.\n\nFixes #2727
Recognized max in provider/model selector parsing where a concrete model lookup can preserve literal :max IDs before falling back to the xhigh alias.\n\nFixes #2727
Recognized max in role aliases, canonical scope expansion, and glob selectors while keeping literal :max model IDs matched before the alias path.\n\nFixes #2727
Prevented provider-scoped fuzzy matching from consuming OpenRouter @upstream selectors whose slug also appears in the model id.
Added resolver coverage for openrouter/deepseek/deepseek-v4-pro@deepseek:high so the upstream routing block and thinking level are both preserved.
Fixes#2708
- Added `omp models` command with `ls`, `find`, `canonical`, and `refresh` actions.
- Removed top-level `--list-models` parsing from CLI args, launch, and main command flow.
- Implemented action-driven model listing with provider filtering, extension loading, and `--json` output.
- Updated unknown provider/model errors and tests to direct users to `omp models` guidance.
- Added `thinking.requiresEffort` to `ThinkingConfig` and baked it in `deriveThinking`/`fillThinkingWireDefaults` via `impliesMandatoryReasoning` for reasoning-only upstreams: Gemini 3.x, Gemini 2.5 Pro, the OpenAI o-series, MiniMax M2, and thinking-only `-reasoner`/`-reasoning` SKUs.
- Added `minimumSupportedEffort()` to `model-thinking.ts` as the clamp target for thinking-off requests on flagged models.
- Moved `stripThinkingVariantToken`/`findThinkingVariantToken` into `identity/family.ts`, taught them the `-reasoning`/`-reasoner` spellings, and re-pointed the `variant-collapse` and coding-agent `model-resolver` imports.
- Dropped `requiresEffort` (with `effortRouting`/`suppressWhenOff`) from collapsed-pair thinking surfaces in `derivePairThinkingSurface`, since the collapsed pair routes off to the bare backing id.
- Regenerated `models.json` and covered derivation, backfill, and reasoning-token pairing in `model-thinking.test.ts` and `variant-collapse.test.ts`.
- Resolved retired effort-tier variant ids in `model-resolver.ts` through the hand-table aliases (`resolveVariantAlias`, `resolveBareVariantAlias`) plus the `X-thinking` → `X` grammar (`stripThinkingVariantToken`), with exact matches always winning while a raw id is live and explicit `:effort` suffixes transferring unchanged.
- Re-keyed models.yml `modelOverrides` and rate-limit selector suppressions from raw member ids onto the collapsed model in `model-registry.ts` (`normalizeSuppressedSelector`, lazy `hasLiveModel` checks so live raw ids keep their own overrides).
- Collapsed custom/config provider model lists at registry rebuild via `collapseBuiltModelVariants`, folding config-defined `X`/`X-thinking` twins into one entry.
- Extended `model-registry.test.ts` and `model-resolver.test.ts` with effort-tier variant collapsing and alias-resolution coverage.
Unconfigured pi/designer now follows modelRoles.default before consulting the Gemini priority chain, matching the smol and slow fallback behavior while preserving priority defaults when no default is configured.
Added regression coverage and updated the coding-agent changelog.
Fixes#2336
When modelRoles.default points at another role alias (e.g. "pi/slow") and the inheriting role is unset, the resolver now recurses through resolveConfiguredRolePattern with a visited-role set so downstream one-layer expanders like completion-bridge's resolveTierModel see concrete model patterns instead of pi/<role>. Self-aliased defaults still collapse to the role's built-in priority chain.
Added regression coverage for both cross-role expansion paths.
Fixes#2336
When modelRoles.default points at pi/smol or pi/slow, unset matching roles now expand back to their built-in priority chain instead of returning the alias as a literal model pattern.
Added regression coverage for self-aliased smol and slow defaults.
Fixes#2336
Unconfigured pi/smol and pi/slow role aliases now inherit modelRoles.default before consulting the cloud-priority list, avoiding silent paid-provider routing for local-default setups.
Added resolver coverage for both roles and updated the coding-agent changelog.
Fixes#2336