Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.
Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.
Sites fixed:
- utils/title-generator.ts (30 -> 1024)
- utils/commit-message-generator (60 -> 1024)
- tts/speech-enhancer (512 -> 1536)
- auto-thinking/classifier online path (8 -> 1024); classifyLocal
keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
classifyLocal keeps ANSWER_MAX_TOKENS
Fixes#4355
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
The model selector's persistence path dropped the `:auto` selector when parsing role values, producing a warning ('Invalid thinking level "auto"') and rendering the badge as `inherit` instead of `auto`. Reload of the default role also lost the auto state whenever the role value carried an explicit `:auto` suffix instead of relying on `defaultThinkingLevel`.
Widen the resolver chain (`parseThinkingSuffix`, `splitThinkingSuffix`, `parseModelString`, `parseModelPattern*`, `ResolvedModelRoleValue`, `ResolvedRoleModel`, `ResolveCliModelResult`) to carry the `AUTO_THINKING` sentinel end to end, and coerce it back to `undefined` at concrete-only boundaries (glob scope patterns, retry fallback, advisor, commit pipeline, guided-goal, bench).
Regression tests cover:
- `resolveModelRoleValue("provider/model:auto")` returns explicit auto without a warning.
- `ModelSelector` renders `DEFAULT (auto)` and `SMOL (auto)` when the role value has `:auto`.
- `cycleRoleModels` activates auto thinking on entering a `:auto` role.
- Startup resume activates auto thinking when `modelRoles.default` carries `:auto`.
Fixes#4128
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.
Fixes#2198
- Added first-party-first provider priority defaults for model ranking.
- Consolidated model resolution to use getModelMatchPreferences from session settings.
- Prioritized providerPriorityRank ahead of usage rank when picking preferred models.
- Added second-pass fallback to default-model or API-key-valid matching order.
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
- Updated OpenAI completions parameterization to honor `disableReasoning` on effort-based compatible models by sending the minimum supported effort.
- Expanded commit and title generation token budgets so reasoning models can return output after internal thinking while non-reasoning calls keep existing limits.
- Switched title generation to request a `set_title` tool call, added extraction from tool-call arguments, and updated tests for the new behavior.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Added canonical model equivalence types, cache helpers, and registry APIs for provider variant lookup.
- Changed model resolution to apply canonical ID overrides/excludes with provider order before fallback matching.
- Added canonical and provider model views in list-models and selector UI with canonical sorting/persistence.
- Updated role/model persistence to store selectors while runtime now resolves concrete canonical-backed provider models.
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).