- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
Two defects in the ceiling work, both found in review.
The local backend shared the online ceiling, so with `autoThinkingMaxEffort:
max` and a sparse ladder the clamp could snap a `hard` bucket up to `max` —
a tier the 3-bucket on-device classifier can never select. The local branch
now pins `xhigh`.
Applying the ceiling before the Low floor also broke the floor's contract:
on `["minimal", "max"]` under an `xhigh` ceiling the intersection hid `max`,
the code concluded the model "maxes out below Low", and it fell through to
`minimal`. The floor is now resolved against the model's own ladder first
and the ceiling filters that pool, so an excluded top tier yields no level
instead of a sub-Low one.
Docs and changelog now scope the guarantee to what `auto` resolves: a
`thinking.requiresEffort` model whose ladder holds nothing under the ceiling
still receives its lowest supported effort from the transport, because it
accepts nothing else. The test that claimed to prove billing is renamed to
say what it checks.
Prompt assertions now cover the `max` criteria and the tie-break exception,
not just the label, since the label alone is inert. Drops the duplicated
pool-level assertions in favour of the contract-level sparse-ladder case.
Pinning the whole rendered prompt made any harmless rewording a test
failure, which AGENTS.md calls out as prompt boilerplate. The guarantee
worth defending is narrower: a user who has not opted in sees no `max`
label and keeps the unconditional tie-break, which is what actually makes
the top tier unreachable. Assert those two facts and drop the snapshot.
Capping the classifier result before `clampAutoThinkingEffort` was not
enough. The clamp seeds `chosen` with `pool[0]`, so a sparse ladder whose
tiers all sit above the request snaps upward instead of down: on
`thinking.efforts: ["max"]` an `xhigh` request returned `max`, letting the
default setting bill the top tier with no opt-in. The same upward snap made
`resolveProvisionalAutoLevel` hand back `max`, breaking the invariant its
doc comment had just claimed.
`clampAutoThinkingEffort` now takes the ceiling and intersects it with the
model's supported tiers, returning `undefined` when nothing is eligible so
auto leaves the current level alone instead of billing an excluded tier.
The classifier passes its configured ceiling and the provisional level
passes XHigh.
`max` became a first-class effort tier in d435385a, but the `auto`
classifier prompt still offers only `low|medium|high|xhigh`. On a model
that exposes the tier, `auto` can therefore never reach it — only the
`ultrathink` keyword can, because it bypasses the classifier entirely.
`providers.autoThinkingMaxEffort` (`xhigh` | `max`, default `xhigh`) lifts
that ceiling. Opting in adds `max` to the classifier vocabulary, gated on
the target model actually supporting the tier, and scopes the tie-break
exception to that prompt variant so the default renders byte-for-byte as
before. A classification above the configured ceiling is clamped before
the model clamp, so a hallucinated `max` cannot cross a ceiling the user
did not opt into. The on-device 3-bucket classifier stays capped at
`xhigh`, and the provisional/fallback level still never provisions `max`.
Also corrects the two `Auto-detect per prompt (low-xhigh)` labels and the
stale `xhigh auto ceiling` comment, which the new setting makes wrong.
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.
Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.
Fixes#3356
- Added `off` and `auto` as valid inputs for the `--thinking` CLI flag.
- Centralized thinking level definitions in `CLI_THINKING_LEVELS` to keep flag options, shell completions, and validation in sync.
- Configured CLI parsing to reject `inherit` as an explicit input to prevent unintended configuration suppression.
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.
Fixes#2808
Mapped the user-facing max thinking selector to the canonical xhigh effort so DeepSeek V4 Pro selectors and --thinking can request provider maximum reasoning.\n\nFixes #2727
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.