fix(coding-agent): size title/commit/speech/classifier budgets for backends that ignore disableReasoning
Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.
Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.
Sites fixed:
- utils/title-generator.ts (30 -> 1024)
- utils/commit-message-generator (60 -> 1024)
- tts/speech-enhancer (512 -> 1536)
- auto-thinking/classifier online path (8 -> 1024); classifyLocal
keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
classifyLocal keeps ANSWER_MAX_TOKENS
Fixes #4355
This commit is contained in:
@@ -16,8 +16,12 @@ import { concreteThinkingLevel, toReasoningEffort } from "../thinking";
|
||||
|
||||
const COMMIT_SYSTEM_PROMPT = prompt.render(commitSystemPrompt);
|
||||
const MAX_DIFF_CHARS = 4000;
|
||||
const COMMIT_MAX_TOKENS = 60;
|
||||
const REASONING_SAFE_MAX_TOKENS = 1024;
|
||||
// Cover the "backend ignores `disableReasoning`" case unconditionally: the
|
||||
// static `model.reasoning` catalog flag can't distinguish a thinking model
|
||||
// declared `reasoning: false` (e.g. Qwen3 served locally via llama.cpp) from
|
||||
// one that never emits thinking. `maxTokens` is a hard cap — non-thinking
|
||||
// completions still return in a handful of tokens (issue #4355).
|
||||
const COMMIT_MAX_TOKENS = 1024;
|
||||
|
||||
/** File patterns that should be excluded from commit message generation diffs. */
|
||||
const NOISE_SUFFIXES = [".lock", ".lockb", "-lock.json", "-lock.yaml"];
|
||||
@@ -101,9 +105,7 @@ export async function generateCommitMessage(
|
||||
if (!apiKey) continue;
|
||||
|
||||
try {
|
||||
const maxTokens = candidate.model.reasoning
|
||||
? Math.max(COMMIT_MAX_TOKENS, REASONING_SAFE_MAX_TOKENS)
|
||||
: COMMIT_MAX_TOKENS;
|
||||
const maxTokens = COMMIT_MAX_TOKENS;
|
||||
const response = await completeSimple(
|
||||
candidate.model,
|
||||
{
|
||||
|
||||
Reference in New Issue
Block a user