977059bb8e
Copilot's /models endpoint exposes token limits under capabilities.limits. max_prompt_tokens is the actual prompt capacity (what OMP calls contextWindow), while max_context_window_tokens is the total window (prompt + output budget). Using the latter inflates contextWindow, which breaks compaction thresholds, overflow detection, and context promotion. Three changes: 1. mapModel reference selection: always prefer the Copilot-specific bundled reference over the global cross-provider reference. Copilot imposes its own limits that are strictly lower than native provider limits. 2. mapModel contextWindow chain: remove max_context_window_tokens from the fallback. New chain: context_length -> max_prompt_tokens -> reference. 3. generate-models: stop overwriting contextWindow/maxTokens in applyGlobalModelsDevFallback. These are provider-specific and should not be replaced with cross-provider models.dev global references. Also fixes bundled values: github-copilot/gpt-5.4 (400k -> 272k) and github-copilot/gpt-5.2 (264k -> 128k) to match live API max_prompt_tokens. Refs: #225, #226