Commit Graph
12 Commits
Author SHA1 Message Date
can1357 93635e7b6a feat(coding-agent): centralized preprocessing and guidance for small models
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
2026-07-11 12:28:18 +02:00
roboomp 8886a528dc fix(coding-agent): size title/commit/speech/classifier budgets for backends that ignore disableReasoning
Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.

Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.

Sites fixed:
- utils/title-generator.ts       (30   -> 1024)
- utils/commit-message-generator (60   -> 1024)
- tts/speech-enhancer            (512  -> 1536)
- auto-thinking/classifier       online path (8  -> 1024); classifyLocal
                                  keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
                                  classifyLocal keeps ANSWER_MAX_TOKENS

Fixes #4355
2026-07-03 00:47:20 +00:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 f0f7a5ba89 feat(coding-agent): introduced tiny model role for background tasks
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
2026-06-27 07:56:27 +02:00
roboomp 3e502a38bc style: bun run fix 2026-06-24 14:12:42 +00:00
roboomp b0e07f52d2 fix(coding-agent): clamped auto thinking to undefined for models without controllable effort
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.

Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.

Fixes #3356
2026-06-24 14:12:18 +00:00
roboomp 648f41c710 fix(coding-agent): expanded local auto-thinking budget
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.

Fixes #2808
2026-06-16 23:36:29 +00:00
can1357 9629842f33 feat(coding-agent): accepted a model directly in ModelRegistry.resolver 2026-06-12 02:33:46 +02:00
roboompandcan1357 8c3149e5a9 fix(ai): scoped antigravity quota blocks by model family
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.

Fixes #2198
2026-06-10 08:26:00 +02:00
can1357 c10eb5e50e feat(coding-agent): added resolver-based auth retries to image and search tools
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
2026-06-07 04:49:19 +02:00
can1357 cba6641299 feat: enabled resolver-based API key retries with refresh and rotation
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
2026-06-07 04:20:48 +02:00
can1357 7f866a48a8 feat(coding-agent): added per-turn AUTO_THINKING in coding-agent session
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.
2026-05-31 03:32:21 +02:00