Commit Graph
16 Commits
Author SHA1 Message Date
Éverton Toffanetto 1f4cdcbbbb fix(coding-agent): cap the local classifier and keep the Low floor
Two defects in the ceiling work, both found in review.

The local backend shared the online ceiling, so with `autoThinkingMaxEffort:
max` and a sparse ladder the clamp could snap a `hard` bucket up to `max` —
a tier the 3-bucket on-device classifier can never select. The local branch
now pins `xhigh`.

Applying the ceiling before the Low floor also broke the floor's contract:
on `["minimal", "max"]` under an `xhigh` ceiling the intersection hid `max`,
the code concluded the model "maxes out below Low", and it fell through to
`minimal`. The floor is now resolved against the model's own ladder first
and the ceiling filters that pool, so an excluded top tier yields no level
instead of a sub-Low one.

Docs and changelog now scope the guarantee to what `auto` resolves: a
`thinking.requiresEffort` model whose ladder holds nothing under the ceiling
still receives its lowest supported effort from the transport, because it
accepts nothing else. The test that claimed to prove billing is renamed to
say what it checks.

Prompt assertions now cover the `max` criteria and the tie-break exception,
not just the label, since the label alone is inert. Drops the duplicated
pool-level assertions in favour of the contract-level sparse-ladder case.
2026-07-26 05:11:53 -03:00
Éverton Toffanetto 3933bbd0f4 fix(coding-agent): keep the auto ceiling out of the model clamp
Capping the classifier result before `clampAutoThinkingEffort` was not
enough. The clamp seeds `chosen` with `pool[0]`, so a sparse ladder whose
tiers all sit above the request snaps upward instead of down: on
`thinking.efforts: ["max"]` an `xhigh` request returned `max`, letting the
default setting bill the top tier with no opt-in. The same upward snap made
`resolveProvisionalAutoLevel` hand back `max`, breaking the invariant its
doc comment had just claimed.

`clampAutoThinkingEffort` now takes the ceiling and intersects it with the
model's supported tiers, returning `undefined` when nothing is eligible so
auto leaves the current level alone instead of billing an excluded tier.
The classifier passes its configured ceiling and the provisional level
passes XHigh.
2026-07-26 04:04:07 -03:00
Éverton Toffanetto e3e4fd3a31 refactor(coding-agent): import THINKING_EFFORTS from the catalog module
AGENTS.md requires catalog values to come from `@oh-my-pi/pi-catalog/<module>`
rather than the pi-ai barrel, which re-exports only the types its own
signatures need.
2026-07-26 03:45:44 -03:00
Éverton Toffanetto fbd3ebffc6 feat(coding-agent): add opt-in max ceiling for auto thinking
`max` became a first-class effort tier in d435385a, but the `auto`
classifier prompt still offers only `low|medium|high|xhigh`. On a model
that exposes the tier, `auto` can therefore never reach it — only the
`ultrathink` keyword can, because it bypasses the classifier entirely.

`providers.autoThinkingMaxEffort` (`xhigh` | `max`, default `xhigh`) lifts
that ceiling. Opting in adds `max` to the classifier vocabulary, gated on
the target model actually supporting the tier, and scopes the tie-break
exception to that prompt variant so the default renders byte-for-byte as
before. A classification above the configured ceiling is clamped before
the model clamp, so a hallucinated `max` cannot cross a ceiling the user
did not opt into. The on-device 3-bucket classifier stays capped at
`xhigh`, and the provisional/fallback level still never provisions `max`.

Also corrects the two `Auto-detect per prompt (low-xhigh)` labels and the
stale `xhigh auto ceiling` comment, which the new setting makes wrong.
2026-07-26 03:16:56 -03:00
can1357 93635e7b6a feat(coding-agent): centralized preprocessing and guidance for small models
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
2026-07-11 12:28:18 +02:00
roboomp 8886a528dc fix(coding-agent): size title/commit/speech/classifier budgets for backends that ignore disableReasoning
Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.

Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.

Sites fixed:
- utils/title-generator.ts       (30   -> 1024)
- utils/commit-message-generator (60   -> 1024)
- tts/speech-enhancer            (512  -> 1536)
- auto-thinking/classifier       online path (8  -> 1024); classifyLocal
                                  keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
                                  classifyLocal keeps ANSWER_MAX_TOKENS

Fixes #4355
2026-07-03 00:47:20 +00:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 f0f7a5ba89 feat(coding-agent): introduced tiny model role for background tasks
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
2026-06-27 07:56:27 +02:00
roboomp 3e502a38bc style: bun run fix 2026-06-24 14:12:42 +00:00
roboomp b0e07f52d2 fix(coding-agent): clamped auto thinking to undefined for models without controllable effort
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.

Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.

Fixes #3356
2026-06-24 14:12:18 +00:00
roboomp 648f41c710 fix(coding-agent): expanded local auto-thinking budget
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.

Fixes #2808
2026-06-16 23:36:29 +00:00
can1357 9629842f33 feat(coding-agent): accepted a model directly in ModelRegistry.resolver 2026-06-12 02:33:46 +02:00
roboompandcan1357 8c3149e5a9 fix(ai): scoped antigravity quota blocks by model family
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.

Fixes #2198
2026-06-10 08:26:00 +02:00
can1357 c10eb5e50e feat(coding-agent): added resolver-based auth retries to image and search tools
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
2026-06-07 04:49:19 +02:00
can1357 cba6641299 feat: enabled resolver-based API key retries with refresh and rotation
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
2026-06-07 04:20:48 +02:00
can1357 7f866a48a8 feat(coding-agent): added per-turn AUTO_THINKING in coding-agent session
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.
2026-05-31 03:32:21 +02:00