fix(coding-agent): preserve vLLM discovered context windows

Read vLLM max_model_len and OpenAI-compatible context_length metadata during model discovery, route providers.vllm.baseUrl into built-in discovery before cached models exist, and avoid sending local placeholder bearer tokens.

Scope the vLLM model cache to the discovery base URL so endpoint changes refetch immediately, and add focused regression coverage for configured and built-in vLLM discovery.
This commit is contained in:
Gerben Meijer
2026-06-13 18:46:28 +02:00
committed by can1357
parent 8e7a2fac97
commit 4a6eec624b
14 changed files with 429 additions and 27 deletions
@@ -271,6 +271,7 @@ describe("createAgentSession deferred model pattern resolution", () => {
slashCommands: [],
enableMCP: false,
enableLsp: false,
skipPythonPreflight: true,
});
try {