Files
oh-my-pi/packages/coding-agent/test
Miroslav Drbal [ApoC] 8b028224e3 feat: auto-reconnect MCP servers on connection loss (#482)
* feat: auto-reconnect MCP servers on connection loss

When an HTTP SSE stream drops (server restart, network interruption),
the transport fires onClose, and the manager proactively reconnects
with retry backoff (500ms, 1s, 2s, 4s). Tools are kept in the registry
during reconnection so they remain selected and available to the agent.

If proactive reconnection fails, stale tools stay registered. When the
agent calls one, the tool bridge detects the retriable connection error
(ECONNREFUSED, ECONNRESET, stale session 404/502/503, etc.), triggers
reconnectServer on the manager, and retries the call once on the fresh
connection. Concurrent reconnect attempts for the same server are deduped.

connectToServer now always installs a default onRequest handler for ping
and roots/list (using getProjectDir()), so all connections -- including
short-lived test/probe ones -- properly respond to server-initiated
requests during initialization.

Post-connection setup (resources, prompts, subscriptions) is extracted
into a shared #loadServerResourcesAndPrompts method used by both initial
connection and reconnection paths.

Add /mcp reconnect <name> command for manual recovery after extended
outages where both proactive and reactive reconnection have failed.

* docs: add changelog entry for MCP auto-reconnect

* fix: address P1 review findings in MCP reconnection

- Save server configs before connection attempt so deferred tools can
  reconnect even when the initial connection timed out (P1-1)
- Make waitForConnection() and getConnectionStatus() aware of in-flight
  reconnections so callers wait instead of failing immediately (P1-2)
- Add epoch counter incremented on disconnectAll() and checked in
  connectAndWireServer() to invalidate stale reconnect attempts that
  outlive a manager reset/reload (P1-3)
- Skip servers with pending reconnections in connectServers() to prevent
  parallel connection attempts for the same server

* fix: deferred tool reconnect and non-blocking transport teardown

- DeferredMCPTool.execute now reconnects when getConnection() fails
  ("MCP server not connected"), not only on network errors from
  callTool. Servers that missed the startup window can now be woken
  by the first tool call against their cached tools. (P1-4)
- #doReconnect fire-and-forgets the old transport close instead of
  awaiting it. HttpTransport.close() sends a DELETE with 30s timeout;
  blocking here delayed the first reconnect attempt by that amount
  on every server restart. (P1-5)

* fix: abort-aware reconnect waits and preserve tool selection on reconnect

- Wrap all reconnect() awaits with withAbort(signal) so user
  cancellation (Esc) interrupts the reconnect backoff loop instead
  of blocking for up to 7.5s. Applies to MCPTool (1 site) and
  DeferredMCPTool (2 sites). (P2-1)
- Remove activateDiscoveredMCPTools call from /mcp reconnect handler.
  refreshMCPTools already preserves the user's prior MCP tool
  selection; the extra activation was silently opting into all
  server tools including ones the user had not enabled. (P2-2)

* fix: rebind MCPTool connection after reconnect, add stdio retriable error

- MCPTool.connection is now mutable; after a successful reconnect retry,
  this.connection is rebound to the fresh connection so subsequent calls
  on the same instance (e.g. batched tool calls) use it instead of
  triggering another reconnect cycle. (P2-3)
- Add "Transport closed" to RETRIABLE_PATTERNS. StdioTransport rejects
  pending requests with this message when the subprocess dies, which
  should trigger the reconnect path just like HTTP transport errors. (P2-4)

---------

Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
2026-03-20 23:27:30 +01:00
..
2026-03-14 10:46:26 +01:00