The previous #runSerialized waited on the shared #dispatchTail, then ran
unconditionally: when two or more events queued behind an in-flight run,
each resumed from the same settled await and started its own run in
parallel, defeating the ordering guarantee for a burst landing in one
coalescing window (message_end + agent_end behind a suspended flush).
Each waiter now chains its own link onto the current tail
(tail.then(run, run)), so queued runs start strictly one after another;
the idle path still runs synchronously, preserving the flush timing the
coalescing tests assert on. The in-flight flag clears only when the
settling link is still the tail, so a later chained link's settle does
not clear it early.
Regression test: two message_end events queued behind a suspended window
flush stay serialized (init call count steps 1 -> 2 -> 3 as each gate
opens); fails on the previous implementation.
The Claude/Cursor/Gemini/Windsurf importers appended user entries before
project entries, so a project `enabled: false` could not claim its dedupe key
ahead of a same-named user server and the disable was silently ignored. Load
project entries first, matching the native/Codex loaders, so a project disable
suppresses a same-named user server.
Updated docs/mcp-config.md to reflect the project-first precedence and added
compound regression coverage.
Fixes#7652
Set HOME and os.homedir() to a dedicated temporary directory for each translated-provider fixture, then restore both after every test. This prevents real user MCP configs from shadowing the project fixture through capability deduplication.
Python/Ruby raise on the removed keyword via their keyword-only
signatures; the Julia helper's kwargs... catch-all silently swallowed
agent("..."; model="default") and forwarded it to the bridge, where
the "+": "delete" strip discarded it. A caller could keep thinking a
model override was honored. Reject model loudly, matching the strict
removal decision on #7621.
The Claude Code, Cursor, Gemini CLI, Windsurf, and VS Code importers built
their canonical MCPServer object without mapping serverConfig.enabled, so a
server declared with "enabled": false stayed undefined and the central
suppressServer filter never fired. Only disabledServers masked the gap.
Map the field in each importer, mirroring opencode.ts and codex.ts, and add
a table-driven regression test across all five importers.
Fixes#7652
The 3 -> 8 bump in callWithCopilotModelRetry was shared by the generic
retryable branch, so a persistent status-less transport blip on Copilot
would ramp across 8 attempts (~11.2s of dead time) instead of the 3 it
took before, and a repeated Retry-After 429 could stretch the same way
on top of the transport's own fetchWithRetry budget.
Derive the budget from the failure kind: model-availability 400s keep the
8-attempt reroll, everything else caps at the previous 3.
Also read COPILOT_TRANSIENT_MODEL_CODES with Object.hasOwn — `code` is
provider-controlled, so a 400 body whose code was `__proto__` or
`toString` classified as transient through the prototype chain.
Any Copilot model in the middle of a rollout (claude-sonnet-4.6,
claude-opus-4.6, gpt-5.4, gpt-5.3-codex, ...) returned a raw HTTP 400 on
roughly half of all turns. GET /models on api.githubcopilot.com returns
two different catalogs across repeated calls: part of the fleet serves
those ids, part rejects them with
400 {"error":{"message":"The requested model is not available for
integrator \"copilot-language-server\". ...",
"code":"model_not_available_for_integrator", ...}}
The absorb machinery already existed and was correct
(isCopilotTransientModelError -> isProviderRetryableError -> the
Anthropic transport's PROVIDER_MAX_RETRIES). Only the classifier missed:
it matched the older model_not_supported code and probed err.code /
err.error.code, while the real code is model_not_available_for_integrator
sitting at err.error.error.code (the SDK stores the parsed body on
.error, and Copilot's body is itself {error:{code}}). So
isProviderRetryableError fell through to "4xx => terminal".
Fix the classifier: providerErrorCode() walks the error envelope up to
depth 3 instead of hardcoding a shape, both Copilot model-availability
codes are accepted, and a wire-body text match backs it up because SDK
envelope shapes drift between provider families while the stringified
message does not. This alone restores the retry path, because
isProviderRetryableError consults the provider hook before its
4xx short-circuit.
Retry shape, since a rejection is a per-request replica reroll rather
than upstream backpressure:
- both transports wait a flat delay between model-flap attempts instead
of the growing backoff; generic retryable failures (429/5xx/transport)
keep their linear ramp and Retry-After handling
- the OpenAI-transport budget goes 3 -> 8 attempts, because a measured
~70% flap window produced a turn that needed 6 wire attempts and
exhausting the budget escalates to the agent-level retry, which
restarts the whole turn
Absorbed attempts cost no tokens: rejections are gateway-side, carry no
usage block, and return in ~208ms median versus ~1884ms for a served
request.
Also refresh the exhausted-retry guidance text, which cited a
nonexistent model id and described the cause as a per-client rollout gap
rather than fleet skew.