Files
oh-my-pi/packages/coding-agent/test
roboomp 41a9afabf1 fix(auto-thinking): clear proxy thinking budget in online classifiers
The online difficulty and unexpected-stop classifiers call the tiny/smol model with disableReasoning plus maxTokens=1024. On the openai-completions transport (LiteLLM), disableReasoning on a reasoning model is downgraded to the lowest reasoning effort, so omp still emits reasoning_effort. LiteLLM/Vertex translates that to an Anthropic thinking.budget_tokens of at least 1024, and max_tokens=1024 is not greater than the budget, so every classifier call 400s.

Give the online classifiers 4096 output tokens so the request clears a proxy-injected minimum thinking budget (and leaves room for the keyword). Local reasoning budgets are unchanged.

Fixes #8610
2026-08-15 05:32:22 +00:00
..
2026-07-27 16:43:53 +02:00
2026-07-27 16:43:53 +02:00
2026-07-27 16:43:53 +02:00
2026-07-27 16:43:53 +02:00
2026-07-01 00:07:30 +00:00
2026-08-05 03:07:16 +02:00
2026-05-30 18:08:51 +02:00
2026-07-27 16:43:53 +02:00
2026-06-24 13:44:47 +00:00