Commit Graph

23 Commits

Author SHA1 Message Date
roboomp 41a9afabf1 fix(auto-thinking): clear proxy thinking budget in online classifiers
The online difficulty and unexpected-stop classifiers call the tiny/smol model with disableReasoning plus maxTokens=1024. On the openai-completions transport (LiteLLM), disableReasoning on a reasoning model is downgraded to the lowest reasoning effort, so omp still emits reasoning_effort. LiteLLM/Vertex translates that to an Anthropic thinking.budget_tokens of at least 1024, and max_tokens=1024 is not greater than the budget, so every classifier call 400s.

Give the online classifiers 4096 output tokens so the request clears a proxy-injected minimum thinking budget (and leaves room for the keyword). Local reasoning budgets are unchanged.

Fixes #8610
2026-08-15 05:32:22 +00:00
can1357 b279db1790 test: refactored test suites to eliminate time-based sleeps and polling loops
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
2026-08-13 19:32:22 +02:00
can1357 a872d77068 chore: cleanup dumb tests 2026-08-02 20:39:23 +02:00
can1357 70e6d2c7dc Merge PR #6680: feat(coding-agent): add opt-in max ceiling for auto thinking (@everton-dgn) 2026-07-30 01:48:51 +02:00
Wolfgang Schoenberger e12e575da3 fix(task): reject effort above configured ceiling 2026-07-27 04:13:56 -07:00
Éverton Toffanetto 1f4cdcbbbb fix(coding-agent): cap the local classifier and keep the Low floor
Two defects in the ceiling work, both found in review.

The local backend shared the online ceiling, so with `autoThinkingMaxEffort:
max` and a sparse ladder the clamp could snap a `hard` bucket up to `max` —
a tier the 3-bucket on-device classifier can never select. The local branch
now pins `xhigh`.

Applying the ceiling before the Low floor also broke the floor's contract:
on `["minimal", "max"]` under an `xhigh` ceiling the intersection hid `max`,
the code concluded the model "maxes out below Low", and it fell through to
`minimal`. The floor is now resolved against the model's own ladder first
and the ceiling filters that pool, so an excluded top tier yields no level
instead of a sub-Low one.

Docs and changelog now scope the guarantee to what `auto` resolves: a
`thinking.requiresEffort` model whose ladder holds nothing under the ceiling
still receives its lowest supported effort from the transport, because it
accepts nothing else. The test that claimed to prove billing is renamed to
say what it checks.

Prompt assertions now cover the `max` criteria and the tie-break exception,
not just the label, since the label alone is inert. Drops the duplicated
pool-level assertions in favour of the contract-level sparse-ladder case.
2026-07-26 05:11:53 -03:00
Éverton Toffanetto a94d04cfca test(coding-agent): assert the default classifier prompt semantically
Pinning the whole rendered prompt made any harmless rewording a test
failure, which AGENTS.md calls out as prompt boilerplate. The guarantee
worth defending is narrower: a user who has not opted in sees no `max`
label and keeps the unconditional tie-break, which is what actually makes
the top tier unreachable. Assert those two facts and drop the snapshot.
2026-07-26 04:13:07 -03:00
Éverton Toffanetto 3933bbd0f4 fix(coding-agent): keep the auto ceiling out of the model clamp
Capping the classifier result before `clampAutoThinkingEffort` was not
enough. The clamp seeds `chosen` with `pool[0]`, so a sparse ladder whose
tiers all sit above the request snaps upward instead of down: on
`thinking.efforts: ["max"]` an `xhigh` request returned `max`, letting the
default setting bill the top tier with no opt-in. The same upward snap made
`resolveProvisionalAutoLevel` hand back `max`, breaking the invariant its
doc comment had just claimed.

`clampAutoThinkingEffort` now takes the ceiling and intersects it with the
model's supported tiers, returning `undefined` when nothing is eligible so
auto leaves the current level alone instead of billing an excluded tier.
The classifier passes its configured ceiling and the provisional level
passes XHigh.
2026-07-26 04:04:07 -03:00
Éverton Toffanetto fbd3ebffc6 feat(coding-agent): add opt-in max ceiling for auto thinking
`max` became a first-class effort tier in d435385a, but the `auto`
classifier prompt still offers only `low|medium|high|xhigh`. On a model
that exposes the tier, `auto` can therefore never reach it — only the
`ultrathink` keyword can, because it bypasses the classifier entirely.

`providers.autoThinkingMaxEffort` (`xhigh` | `max`, default `xhigh`) lifts
that ceiling. Opting in adds `max` to the classifier vocabulary, gated on
the target model actually supporting the tier, and scopes the tie-break
exception to that prompt variant so the default renders byte-for-byte as
before. A classification above the configured ceiling is clamped before
the model clamp, so a hallucinated `max` cannot cross a ceiling the user
did not opt into. The on-device 3-bucket classifier stays capped at
`xhigh`, and the provisional/fallback level still never provisions `max`.

Also corrects the two `Auto-detect per prompt (low-xhigh)` labels and the
stale `xhigh auto ceiling` comment, which the new setting makes wrong.
2026-07-26 03:16:56 -03:00
can1357 db937bd149 feat(coding-agent): added coarse effort parameter to task tool
- Add `effort` (`lo`/`med`/`hi`) parameter to task spawn parameters and prompts.
- Implement `resolveTaskEffortLevel` to map coarse task effort onto model-supported thinking ranges.
- Pass effort configuration through executor options and structured subagent requests.
2026-07-24 15:51:34 +02:00
can1357 93635e7b6a feat(coding-agent): centralized preprocessing and guidance for small models
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
2026-07-11 12:28:18 +02:00
can1357 d435385ab1 feat: introduced max reasoning effort tier across model and rpc systems
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
2026-07-10 13:39:42 +02:00
can1357 007a0108b9 test(coding-agent): cover reasoning-safe helper budgets 2026-07-05 13:25:25 +02:00
roboomp b0e07f52d2 fix(coding-agent): clamped auto thinking to undefined for models without controllable effort
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.

Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.

Fixes #3356
2026-06-24 14:12:18 +00:00
can1357 6ff37e346a feat(coding-agent): extended --thinking CLI flag options
- Added `off` and `auto` as valid inputs for the `--thinking` CLI flag.
- Centralized thinking level definitions in `CLI_THINKING_LEVELS` to keep flag options, shell completions, and validation in sync.
- Configured CLI parsing to reject `inherit` as an explicit input to prevent unintended configuration suppression.
2026-06-23 01:46:36 +02:00
can1357 5d72ce1237 Merge PR #2729: fix(coding-agent): accept max thinking alias (@roboomp)
# Conflicts:
#	packages/coding-agent/test/model-resolver.test.ts
#	packages/coding-agent/test/sdk-model-selection.test.ts
2026-06-19 00:58:52 +02:00
roboomp 648f41c710 fix(coding-agent): expanded local auto-thinking budget
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.

Fixes #2808
2026-06-16 23:36:29 +00:00
roboomp 07958045c3 fix(coding-agent): guarded thinking selector maps
Checked thinking selector maps with own-property lookup so inherited Object keys cannot parse as valid efforts or thinking levels.\n\nFixes #2727
2026-06-16 03:35:22 +00:00
roboomp 46a9867773 fix(coding-agent): preserved max model globs
Kept max as a selector alias only after literal model lookup misses and left scoped globs matching literal :max model ids.\n\nFixes #2727
2026-06-16 01:32:49 +00:00
roboomp 83a5260957 fix(coding-agent): accepted max thinking alias
Mapped the user-facing max thinking selector to the canonical xhigh effort so DeepSeek V4 Pro selectors and --thinking can request provider maximum reasoning.\n\nFixes #2727
2026-06-16 01:09:37 +00:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 9d457f73d9 test: migrated test imports to package subpath exports
- Replaced relative `../src` imports with `@oh-my-pi/pi-ai` and `@oh-my-pi/pi-agent-core` subpaths.
2026-06-08 19:03:55 +02:00
can1357 7f866a48a8 feat(coding-agent): added per-turn AUTO_THINKING in coding-agent session
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.
2026-05-31 03:32:21 +02:00