Two defects in the ceiling work, both found in review.
The local backend shared the online ceiling, so with `autoThinkingMaxEffort:
max` and a sparse ladder the clamp could snap a `hard` bucket up to `max` —
a tier the 3-bucket on-device classifier can never select. The local branch
now pins `xhigh`.
Applying the ceiling before the Low floor also broke the floor's contract:
on `["minimal", "max"]` under an `xhigh` ceiling the intersection hid `max`,
the code concluded the model "maxes out below Low", and it fell through to
`minimal`. The floor is now resolved against the model's own ladder first
and the ceiling filters that pool, so an excluded top tier yields no level
instead of a sub-Low one.
Docs and changelog now scope the guarantee to what `auto` resolves: a
`thinking.requiresEffort` model whose ladder holds nothing under the ceiling
still receives its lowest supported effort from the transport, because it
accepts nothing else. The test that claimed to prove billing is renamed to
say what it checks.
Prompt assertions now cover the `max` criteria and the tie-break exception,
not just the label, since the label alone is inert. Drops the duplicated
pool-level assertions in favour of the contract-level sparse-ladder case.
Capping the classifier result before `clampAutoThinkingEffort` was not
enough. The clamp seeds `chosen` with `pool[0]`, so a sparse ladder whose
tiers all sit above the request snaps upward instead of down: on
`thinking.efforts: ["max"]` an `xhigh` request returned `max`, letting the
default setting bill the top tier with no opt-in. The same upward snap made
`resolveProvisionalAutoLevel` hand back `max`, breaking the invariant its
doc comment had just claimed.
`clampAutoThinkingEffort` now takes the ceiling and intersects it with the
model's supported tiers, returning `undefined` when nothing is eligible so
auto leaves the current level alone instead of billing an excluded tier.
The classifier passes its configured ceiling and the provisional level
passes XHigh.
AGENTS.md requires catalog values to come from `@oh-my-pi/pi-catalog/<module>`
rather than the pi-ai barrel, which re-exports only the types its own
signatures need.
`max` became a first-class effort tier in d435385a, but the `auto`
classifier prompt still offers only `low|medium|high|xhigh`. On a model
that exposes the tier, `auto` can therefore never reach it — only the
`ultrathink` keyword can, because it bypasses the classifier entirely.
`providers.autoThinkingMaxEffort` (`xhigh` | `max`, default `xhigh`) lifts
that ceiling. Opting in adds `max` to the classifier vocabulary, gated on
the target model actually supporting the tier, and scopes the tie-break
exception to that prompt variant so the default renders byte-for-byte as
before. A classification above the configured ceiling is clamped before
the model clamp, so a hallucinated `max` cannot cross a ceiling the user
did not opt into. The on-device 3-bucket classifier stays capped at
`xhigh`, and the provisional/fallback level still never provisions `max`.
Also corrects the two `Auto-detect per prompt (low-xhigh)` labels and the
stale `xhigh auto ceiling` comment, which the new setting makes wrong.
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.
Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.
Sites fixed:
- utils/title-generator.ts (30 -> 1024)
- utils/commit-message-generator (60 -> 1024)
- tts/speech-enhancer (512 -> 1536)
- auto-thinking/classifier online path (8 -> 1024); classifyLocal
keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
classifyLocal keeps ANSWER_MAX_TOKENS
Fixes#4355
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.
Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.
Fixes#3356
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.
Fixes#2808
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.
Fixes#2198
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.