- Classified model_not_available_for_integrator as a permanent entitlement denial instead of transient fleet skew.
- Preserved the provider response and Available models list while retaining model_not_supported fleet retries.
Fixes#7819
Any Copilot model in the middle of a rollout (claude-sonnet-4.6,
claude-opus-4.6, gpt-5.4, gpt-5.3-codex, ...) returned a raw HTTP 400 on
roughly half of all turns. GET /models on api.githubcopilot.com returns
two different catalogs across repeated calls: part of the fleet serves
those ids, part rejects them with
400 {"error":{"message":"The requested model is not available for
integrator \"copilot-language-server\". ...",
"code":"model_not_available_for_integrator", ...}}
The absorb machinery already existed and was correct
(isCopilotTransientModelError -> isProviderRetryableError -> the
Anthropic transport's PROVIDER_MAX_RETRIES). Only the classifier missed:
it matched the older model_not_supported code and probed err.code /
err.error.code, while the real code is model_not_available_for_integrator
sitting at err.error.error.code (the SDK stores the parsed body on
.error, and Copilot's body is itself {error:{code}}). So
isProviderRetryableError fell through to "4xx => terminal".
Fix the classifier: providerErrorCode() walks the error envelope up to
depth 3 instead of hardcoding a shape, both Copilot model-availability
codes are accepted, and a wire-body text match backs it up because SDK
envelope shapes drift between provider families while the stringified
message does not. This alone restores the retry path, because
isProviderRetryableError consults the provider hook before its
4xx short-circuit.
Retry shape, since a rejection is a per-request replica reroll rather
than upstream backpressure:
- both transports wait a flat delay between model-flap attempts instead
of the growing backoff; generic retryable failures (429/5xx/transport)
keep their linear ramp and Retry-After handling
- the OpenAI-transport budget goes 3 -> 8 attempts, because a measured
~70% flap window produced a turn that needed 6 wire attempts and
exhausting the budget escalates to the agent-level retry, which
restarts the whole turn
Absorbed attempts cost no tokens: rejections are gateway-side, carry no
usage block, and return in ~208ms median versus ~1884ms for a served
request.
Also refresh the exhausted-retry guidance text, which cited a
nonexistent model id and described the cause as a per-client rollout gap
rather than fleet skew.
GitHub Copilot intermittently returns `HTTP 400 model_not_supported`
for preview models (gpt-5.3-codex, gpt-5.4, gpt-5.4-mini, ...) on OAuth
clients other than VS Code, even when `/models` reports the model as
enabled. Root cause is a per-OAuth-client rollout gap across Copilot's
responses backend; repeating the identical request typically lands on
a backend that has the model. See opencode#13313.
- Add `isCopilotTransientModelError` and `callWithCopilotModelRetry`
in `utils/retry` (3 attempts, linear backoff, abort-aware, no-op
for non-Copilot providers).
- Wrap `client.responses.create` in `openai-responses` and the
initial completions stream in `openai-completions` with the retry.
- Extend Anthropic `isProviderRetryableError` to treat Copilot
transient model errors as provider-retryable.
- Rename `rewriteCopilotAuthError` to `rewriteCopilotError` and add
a 400 `model_not_supported` rewrite surfacing actionable guidance
(retry, switch to a GA model, or run from VS Code) after retries
exhaust.
- Rename test file accordingly and add a dedicated retry unit test.