c00790fa3a
Ollama Cloud's deepseek-v4-pro and deepseek-v4-flash deployments reject any output budget above 65536 with HTTP 400, despite advertising a 1M context / 384K output (ollama/ollama#16890). Ollama's /api/show never reports this cap, so the catalog left the base models at the full context window and the dated tag deepseek-v4-flash:0731 at a stale 8192 fallback. Pin these ids (base plus tag variants) to min(contextWindow, 65536) at both runtime discovery and generation; other cloud models keep their discovered limits. Fixes #7266