The online difficulty and unexpected-stop classifiers call the tiny/smol model with disableReasoning plus maxTokens=1024. On the openai-completions transport (LiteLLM), disableReasoning on a reasoning model is downgraded to the lowest reasoning effort, so omp still emits reasoning_effort. LiteLLM/Vertex translates that to an Anthropic thinking.budget_tokens of at least 1024, and max_tokens=1024 is not greater than the budget, so every classifier call 400s.
Give the online classifiers 4096 output tokens so the request clears a proxy-injected minimum thinking budget (and leaves room for the keyword). Local reasoning budgets are unchanged.
Fixes#8610
Align isUnexpectedStopCandidate with #isEmptyAssistantStop: only a signed (non-whitespace thinkingSignature) thinking-only stop is a candidate. Unsigned thinking-only stops stay with the empty-stop retry path and its cap, so unexpected-stop classification no longer re-handles a turn the empty-stop cap already gave up on.
Fixes#7499
isUnexpectedStopCandidate only counted non-whitespace text blocks, so a stopReason:"stop" turn whose sole content was a signed thinking block bypassed the unexpected-stop guard entirely. #isEmptyAssistantStop treats signed thinking as terminal (not empty), so such stops fell through both recovery handlers and the agent silently stopped mid-task.
Count non-whitespace thinking blocks as candidates and feed the thinking text to the classifier when no text block is present.
Fixes#7499
- Added `features.unexpectedStopDetection` and `unexpectedStopModel` settings for opt-in behavior.
- Added assistant-stop handling to classify stop reasons and resume generation with retry prompts.
- Added unexpected-stop classifier logic with candidate checks, model fallback, and YES/NO parsing.
- Added retry tracking that caps auto-continues at three attempts and logs a warning when exceeded.