Commit Graph
95 Commits
Author SHA1 Message Date
roboomp 95f64b6142 fix(mcp): clear stale OAuth credential on definitive refresh failure
When an HTTP MCP server returns invalid_grant (or invalid_token / revoked /
plain 401 from the token endpoint) during OAuth refresh, MCPManager
previously logged "MCP OAuth refresh failed, using existing token" and
re-attached the stale access token as Authorization: Bearer on every
subsequent request. The next tool-load 401'd with invalid_token, future
sessions repeated the loop, and the only recovery was to hand-clear the
credential row in agent.db. Reported with Logfire as the trigger; any
remote HTTP MCP that rotates / revokes refresh tokens is affected.

#resolveAuthConfig now reuses pi-ai's isDefinitiveOAuthFailure classifier
(same one auth-broker and AuthStorage use for first-party providers): on
a definitive failure it calls AuthStorage.remove(credentialId), drops the
Bearer entirely, and the next request surfaces a clean auth error so the
user can /mcp reauth <server> (or /mcp unauth) to recover. Transient
failures (network/fetch failed/ECONNREFUSED) still fall back to the
existing token to ride out blips.

Verified with new mcp-manager-oauth-refresh.test.ts (invalid_grant, 401,
transient fallback, happy-path rotation). The full mcp-* test set
(45 tests across 5 files) still passes.

Fixes #1908
2026-06-05 05:47:09 +00:00
can1357 7a7479f7e3 feat(coding-agent/modes): added arrow-key tab switching to dashboards
- Mapped left/right arrows to switch tabs in the extension dashboard.
- Updated agent and extension dashboard footers to show arrow controls.
- Reformatted HTTP transport assignment and reordered an import.
2026-06-04 16:37:50 +02:00
can1357 51865fc1ea test(coding-agent): updated assertions for condensed prompt wording
- Aligned handoff, reminder, and system-prompt expectations with shortened copy.
- Added HTTP transport test for required initialize failures.
- Guarded SSE startup timeout against stale connection races.
2026-06-04 16:35:45 +02:00
roboomp cb2d5b859a fix(coding-agent): restored exa mcp fallback
Restored Exa's unauthenticated MCP fallback when neither auth storage nor EXA_API_KEY provides credentials. Preserved API-key search ordering and bounded the MCP request with the web-search hard timeout. Updated provider copy and regression coverage for the no-key path.\n\nFixes #1860
2026-06-04 13:18:08 +00:00
roboomp cf33b64033 fix(mcp): updated completed tool status icons
MCP tool renderers now merge call and result components so completed blocks replace the pending call header with success or error status.

MCP protocol errors now preserve top-level isError so downstream render state can use the error status consistently.

Added renderer coverage for completed MCP success and error blocks.

Fixes #1855
2026-06-04 12:14:57 +00:00
VoidChecksum bab53df38a fix(coding-agent/mcp): handle async broken-pipe rejections in stdio transport
On Windows, Bun's FileSink surfaces a broken pipe as a *rejected Promise*
(the EPIPE arrives via a processTicksAndRejections tick), not a synchronous
throw. The sync-only `writeFrame()` from #1711 returns `true` for such a
write and lets the rejected promise float; `request()` likewise wrote stdin
without awaiting. Either path escapes as an unhandled rejection, which the
postmortem handler turns into a fatal process.exit(1) when an MCP server
exits mid-handshake.

- writeFrame: also neutralize a rejected Promise returned by write()/flush()
  (covers the notify + #sendResponse paths).
- request(): await write()/flush() so the rejection lands in the existing
  try/catch and rejects the request instead of floating.
- tests: cover the async-rejection path for writeFrame.

Follow-up to #1710 / #1711, which addressed only the synchronous throw. The
existing notify mid-handshake test fails on Windows without this change
because the real FileSink rejects asynchronously.
2026-06-03 10:51:42 +02:00
roboomp 6598c69a2f fix(mcp-stdio): made close() unconditional in resource teardown
Per second review on #1711: when notify()'s write failed, #handleClose()
flipped #connected=false before the throw. connectToServer()'s catch
then called transport.close(), but close() early-returned because
#connected was already false — so #process.kill() and the readLoop
await never ran. A subprocess that closed its stdin without exiting
(parent EPIPE, transport dead, subprocess alive) would leak.

close() no longer guards on #connected for the resource phase. It still
calls #handleClose() once (idempotent — only if #connected is true), then
unconditionally walks the cleanup chain (kill process, null #process,
await + null #readLoop), each step individually guarded so repeat calls
remain no-ops. onClose still fires exactly once per transport lifetime.

Two new tests in StdioTransport.close: (1) close() called after the
read-loop has already EOF'd and torn down still completes cleanup
without throwing and without re-firing onClose; (2) repeated close()
calls fire onClose exactly once and leave #connected=false.
2026-06-02 12:35:35 +00:00
roboomp b870732258 fix(mcp-stdio): surfaced notify() write failures so handshake fails loudly
Per review on #1711: when the write inside notify() fails (FileSink EPIPE
on Windows during the initialize/notifications-initialized race), silently
closing the transport while still resolving the notify promise let
initializeConnection() return a 'connected' handle wrapping a dead
transport. The manager only wires its reconnect onClose handler after
connectToServer() resolves, so a swallowed handshake failure would
neither reconnect nor fail the connection — it would just leak.

notify() now still calls #handleClose() on write failure (so any wired
onClose runs) but additionally throws `Transport closed while sending
notification "<method>"`. The single in-tree notify caller is
initializeConnection() at client.ts:122; the rejection propagates into
connectToServer()'s catch (which closes the transport and rethrows) and
on into the manager's pending-connection error path. #sendResponse() is
unchanged — silent on failure, since a dead subprocess has no use for
the response.

Test coverage updated to document the surfaced-rejection contract and
also assert transport.connected flips to false.
2026-06-02 12:28:07 +00:00
roboomp 3dacaf263f fix(mcp-stdio): swallowed EPIPE from notify and sendResponse stdin writes
StdioTransport.notify() and #sendResponse() wrote to the subprocess's
stdin without try/catch, while the sibling request() already wrapped the
same write/flush. The #connected guard cannot close the race: connect()
sets it synchronously; #handleClose() only clears it from the read loop's
finally after EOF on stdout — strictly later than the parent's next
stdin write. When an MCP server exits between the initialize response and
the notifications/initialized notification, Bun's FileSink throws EPIPE
synchronously on Windows and the async notify wrapper surfaces it as an
unhandled rejection.

Route both call sites through a new writeFrame(stdin, frame) helper that
catches synchronous sink failures and returns a boolean. notify() now
tears the transport down via #handleClose() on failure so the reconnect
machinery engages; #sendResponse() stays silent because a dead subprocess
has no use for the response. The redundant inner try/catch around
#sendResponse in #handleServerRequest() is removed.

Fixes #1710
2026-06-02 12:22:23 +00:00
roboomp 6881c5cc41 fix(mcp): drop stale connection when reconnect breaker trips
Leaving the dead connection in `#connections` made `getConnectionStatus` report `connected` and `waitForConnection` hand a closed transport to callers after the breaker had explicitly suspended the server. Mirror `#doReconnect`'s teardown: detach `onClose`, fire-and-forget `transport.close()`, and drop the entry from `#connections` (plus its in-flight slots in `#pendingConnections`/`#pendingToolLoads`). Tools stay registered in `#tools` so the user can recover with `/mcp reconnect`.

Test asserts `getConnectionStatus("crashy") === "disconnected"` after the burst.

Refs #1592
2026-05-31 15:38:21 +00:00
roboomp 230514d937 fix(mcp): cap automatic reconnect bursts to prevent fork-bomb
A stdio MCP server that completes the initialize + tools/list handshake and then exits cleanly will fire `transport.onClose` on every clean exit, and the old `MCPManager.reconnectServer` path spawned again unconditionally. A misconfigured PHP-shebang MCP (e.g. Laravel Boost in a non-Laravel project) hit this loop and forked 66 487 `php84` processes parented directly to the agent's `bun` PID until macOS force-rebooted.

Add a per-server sliding-window circuit breaker: at most 5 reconnect attempts per 30 s window. The transport `onClose` callback and the per-tool-call retry in `tool-bridge` are subject to the breaker; `/mcp reconnect` passes `{ manual: true }` to reset the window so users can recover after fixing the underlying misconfiguration. Stale `onClose` is detached when the breaker trips so a late EOF event cannot re-arm the loop.

Defended by `mcp-reconnect-storm.test.ts`: a Bun stdio fixture answers the handshake and exits, then asserts the spawn count stays at ≤ 10 (was 127 without the fix).

Fixes #1592
2026-05-31 15:34:29 +00:00
can1357 e831c2c758 chore: reformat 2026-05-30 18:08:51 +02:00
roboomp 2266fdae82 fix(mcp): honor disabled mcp timeouts for sse startup
When the operator disables MCP client-side timeouts via timeout: 0 or OMP_MCP_TIMEOUT_MS=0, do not impose a 1s startup deadline on the optional HTTP GET SSE listener — let the listener wait as long as the server takes so server-to-client messages are not lost.

Refs #1460
2026-05-27 23:37:10 +00:00
roboomp c0c9049cca fix(mcp): bounded optional http sse startup
Abort the optional Streamable HTTP GET SSE listener attempt after a short bounded startup window so POST-only request/response servers can finish initialization.

Fixes #1460
2026-05-27 23:32:38 +00:00
can1357 bc9833d422 feat(mcp): added validation and warning for invalid OMP_MCP_TIMEOUT_MS
- Rejected negative values in addition to non-numeric ones, falling back to per-server config or default 30s.
- Emitted a logger warning when an invalid env value is ignored.
- Added tests covering negative and non-numeric rejection cases.
2026-05-26 20:30:02 +02:00
Can BölükandGitHub a2f4c70299 Merge pull request #1415 from omsrisrieternalradhakrsna/mcp-timeout-env-override
Allow disabling MCP client timeouts
2026-05-26 21:29:19 +03:00
can1357 7f27627a90 feat(mcp/oauth): added RFC 8414 path-ful issuer form to OAuth discovery
- Extended `buildWellKnownUrls` and `#resolveRegistrationEndpoint` to try `/.well-known//` as a third candidate after origin-root and path-prefixed forms.
- Fixed single-segment path handling so `/my-service` is treated as the gateway prefix rather than dropped.
- Fixed missing `await` on `#tryWellKnownForRegistration` that caused path-prefixed fallback to return an unresolved Promise.
- Added tests for single-segment prefix discovery and RFC 8414 path-ful issuer fallback.
2026-05-26 20:22:25 +02:00
SUPREME e415adecd5 Allow disabling MCP client timeouts 2026-05-26 22:05:41 +05:30
Mohd Faiz Hasim 5be0f01a91 feat(coding-agent): support OAuth discovery for path-prefixed auth servers behind gateways
- Extract resource_metadata URL from WWW-Authenticate and follow RFC 9728 chain
- Add buildWellKnownUrls with path-prefixed well-known fallback for gateways
- Fix resolveRegistrationEndpoint to try path-prefixed well-known (was missing await)
- Support relative Mcp-Auth-Server URL resolution against server URL
- Pass resourceMetadataUrl through all discoverOAuthEndpoints call sites
- Add comprehensive tests for path-prefixed, resource_metadata, and relative URL flows
2026-05-26 22:50:32 +08:00
can1357 5364a9bfd0 refactor(tool-discovery): simplified tool discovery API by removing MCP-specific shims
- Removed deprecated MCP-specific type aliases and functions from tool-discovery module, consolidating to unified generic tool discovery API.
- Migrated session and SDK code to use generic filterBySource() and collectDiscoverableTools() instead of MCP-specific variants.
- Removed deprecated interface members including hasQueuedMessages(), FocusPane, AcpBuiltinCommandRuntime, and legacy settings methods.
- Updated test suites to use renamed generic discovery methods and removed back-compat test coverage for legacy MCP shapes.
2026-05-26 15:27:05 +02:00
can1357 e7200c2e4a feat(ai): added unified normalize flow for Google/CCA schema handling
- Implemented a unified normalization flow by switching Google/CCA handling to normalizeSchemaForGoogle/CCA.
- Added normalize.ts with recursive node normalization, nullable-union checks, and combiner collapsing.
- Removed sanitize-google.ts and normalize-cca.ts, replacing them with normalize exports in schema indexes.
- Added spill-to-description utilities with spill/paren modes and `$defs` exclusion for unsupported fields.
- Updated MCP bridge and schema tests to use normalizeSchemaFor* APIs with expanded compatibility checks.
- Documented normalization behavior changes and breaking rename in constraints and package changelog files.
2026-05-16 19:04:09 +02:00
can1357 2867e1f4e3 feat(deps): added pi.zod exports and removed TypeBox package exports
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
2026-05-15 14:46:54 +02:00
Vilmos Nebehaj a7f73ee645 fix(coding-agent): persist dynamically registered MCP OAuth client_id
When an MCP server uses OAuth Dynamic Client Registration (RFC 7591) and
no client_id is pre-configured, MCPOAuthFlow registers a fresh public
PKCE client on each authorize, captures the issued client_id into a
private field, then discards it once the flow object goes out of scope.
At refresh time, MCPManager#resolveAuthConfig calls refreshMCPOAuthToken
with auth.clientId from mcp.json — which is empty for these servers —
so providers that require client_id on the refresh grant (e.g. Linear at
mcp.linear.app/token) reject with HTTP 401 invalid_client. The user is
forced to /mcp reauth manually every time the access token expires.

This change threads the resolved/registered client credentials back out
of the OAuth flow and persists them into mcp.json so refresh has what
it needs indefinitely:

- MCPOAuthFlow exposes resolvedClientId / registeredClientSecret getters.
- MCPCommandController#handleOAuthFlow returns OAuthFlowResult with
  credentialId + clientId + clientSecret, populated from the flow's
  post-login state.
- The initial-connect non-wizard path and /mcp reauth path persist the
  returned client credentials into both auth.{clientId,clientSecret}
  (used at refresh) and oauth.{clientId,clientSecret} (used by future
  /mcp reauth to skip re-registration).
- The wizard's onOAuth callback signature now returns the same shape;
  #launchOAuthFlow folds the registered credentials into wizard state so
  the final mcp.json entry built by #buildServerConfigWithAuth includes
  them under auth.{clientId,clientSecret}.

Servers that configure a static oauth.clientId in mcp.json (Notion,
Slack, Datadog) are unaffected: #tryRegisterClient short-circuits, the
returned clientId equals the configured one, and the write-back is a
no-op.

Adds two MCPOAuthFlow unit tests covering both paths.
2026-05-13 15:39:04 -07:00
can1357 1bde755933 feat(coding-agent): added global singletons for URL protocol handlers
- Added process-wide singleton instances for InternalUrlRouter, AsyncJobManager, and MCPManager.
- Changed internal URL protocols to resolve through registered sessions and scan all active roots/datasets for matches.
- Refactored agent, artifact, memory, rule, skill, jobs, and mcp handlers to use shared manager and rule/skill state.
- Removed per-session protocol/tool wiring and switched tests to initialize and reset global singleton state.
2026-05-12 05:07:52 +02:00
can1357 4a6c56772f fix(coding-agent/mcp): fixed SSE parsing by draining the stream in a single pass
- Refactored SSE response parsing to read the stream with a single in-progress drain loop.
- Resolved the matching request promise when the expected JSON-RPC result or error arrives and continued dispatching remaining SSE messages on the same stream.
- Adjusted error handling to reject on abort/timeout or missing response and clear the timeout when stream parsing completes.
2026-05-10 06:47:19 +02:00
HabibPro1999andcan1357 941683cd66 refactor(coding-agent): extracted generic tool discovery module
Moved the BM25 discovery index out of mcp/discoverable-tool-metadata.ts
into tool-discovery/tool-index.ts and generalized it to work over any
AgentTool source (built-in, MCP, extension, custom) instead of only MCP.

Added DiscoverableTool with explicit source and summary fields. Kept
the MCP-typed helpers (collectDiscoverableMCPTools, searchDiscoverable
MCPTools, etc.) as thin re-exports from the original path so existing
callers continue to work unchanged.

Tests cover the generic index, source filtering, summary fallback, and
back-compat with the legacy MCP-only API.
2026-05-06 18:08:48 +02:00
Can BölükandGitHub 333510aced Merge pull request #890 from apoc/fix/mcp-tool-cache-stability
fix(coding-agent/mcp): stabilize tool ordering and skip redundant prompt rebuilds
2026-05-01 19:09:23 +02:00
can1357 3a46b748a3 perf(ai): dynamic oauth imports per provider
- Documented and removed `utils/oauth` from the `ai` package entrypoint, noting it as a breaking change.
- Refactored `cli`, `auth-storage`, and `utils/oauth` to load provider modules via scoped dynamic `import()` calls.
- Removed top-level provider imports and barrel exports from `utils/oauth/index.ts`, streamlining oauth module loading.
- Consolidated OAuth symbol, type, and provider imports in coding-agent and tests to `@oh-my-pi/pi-ai/utils/oauth` modules.
- Defined `DEFAULT_LOCAL_TOKEN` locally in model-registry and removed its cross-package OAuth import usage.
2026-04-30 15:14:07 +02:00
Miroslav Drbal b5ca55e79f fix(coding-agent/mcp): stabilize tool ordering and skip redundant prompt rebuilds
Two cache-stability fixes for Anthropic prompt caching during MCP server
reconnects, which happen routinely (~5 min per server) in long sessions
due to SSE transport keepalive timeouts.

1) MCPManager: deterministic tool ordering

   `#tools` is now sorted by name after every mutation. The previous
   filter-out + push-to-end pattern in `#replaceServerTools` moved the
   reconnecting server's tools to the end of the array, producing a new
   byte order whenever the reconnect sequence differed from the initial
   discovery sequence. With multiple healthy servers, each reconnect of
   the non-last server flipped the order and invalidated the tools
   cache breakpoint sent to Anthropic.

   Sort applies in `discoverAndConnect` (initial population) and
   `#replaceServerTools` (used by `reconnectServer` and
   `refreshServerTools`). The comparator is character-code based,
   locale-independent and deterministic. `sortMCPToolsByName` is
   exported as a small generic helper and unit-tested.

2) AgentSession: skip system-prompt rebuild when inputs are unchanged

   `#applyActiveToolsByName` (called from `refreshMCPTools` after every
   reconnect) used to unconditionally call `rebuildSystemPrompt` and
   `setSystemPrompt` even when the resulting prompt was byte-identical.
   This wasted CPU on every flap and risked silent cache invalidation
   if the rebuild path ever became non-deterministic.

   Now `#applyActiveToolsByName` computes a stable signature of the
   inputs `rebuildSystemPrompt` reads and skips the rebuild when the
   signature matches the last successful one. The signature covers:
     - active tool names in render order
     - active tool labels and descriptions (rendered as `{{label}}:
       \`{{name}}\`` in the prompt body)
     - when MCP discovery is on, every registry tool's name + label +
       description (the prompt summarizes discoverable-but-inactive
       MCP tools)
     - per-server MCP `instructions` text (embedded under "## MCP
       Server Instructions" in the appended prompt; can change on
       server upgrade while tool list stays identical)

   Server instructions are read via a new optional
   `getMcpServerInstructions` callback on `AgentSessionConfig`, wired
   from the SDK as `() => mcpManager.getServerInstructions()`.

   `refreshBaseSystemPrompt()` continues to rebuild unconditionally and
   refreshes the cached signature, so explicit refreshes still pick up
   ambient changes (edit-mode toggles, memory writes, etc.) that the
   signature does not cover.

Signature inputs deliberately NOT covered: tool input schemas, memory
instructions read from disk, and other ambient state. Callers that
mutate those must call `refreshBaseSystemPrompt()` explicitly; existing
hooks (`#syncEditToolModeAfterModelChange`, memory hooks, `/clear`)
already do.
2026-04-30 15:04:33 +02:00
can1357 88a1072cc5 feat(coding-agent): implemented mcp__-prefixed MCP tool IDs for parsing
- Renamed MCP tool IDs from `mcp_<server>_<tool>` to `mcp__<server>_<tool>`, and changed built-in `grep` to `search`.
- Updated `parseMCPToolName()` and bridge helpers to require and trim the `mcp__` prefix.
- Updated cursor, manager, and session discovery flows to require `mcp__`-prefixed tool names.
- Updated MCP tests and assertion fixtures to use `mcp__`-prefixed tool IDs and expected system prompts.
2026-04-28 01:15:35 +02:00
can1357 52da3674a4 feat(coding-agent): added auto-rebase for stale atom/hashline anchors
- Added auto-rebasing for stale atom and hashline anchors within ±2 lines, with warning diagnostics.
- Removed atom range-locator support and dropped `between` ops, updating docs/tests for `sub` over `set` nudges.
- Changed no-op handling to track unchanged hashline edits and emit contextual hints for unchanged ranges.
- Expanded Anthropic strict error handling to retry on schema-too-complex and compiled-grammar-too-large errors.
- Added anchor-retargeting warning checks and updated edit tests to expect locator and range-locator rejections.
- Updated python tool-call fixtures by adding `title` to executed-cell payloads and adjusting related test expectations.
2026-04-26 05:22:15 +02:00
can1357 44ea247cf8 fix(coding-agent/tools): corrected json-tree inline arg truncation logic
- Updated inline argument formatting in formatArgsInline to truncate key/value pairs by width and show ellipsis when values overflow.
- Fixed JSON-tree output to hide harness-internal INTENT_FIELD and __partialJson keys from top-level tool argument rendering.
- Fixed multiline scalar rendering and empty-structure display paths in renderJsonTreeLines after tree-prefix traversal changes.
2026-04-26 04:32:27 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 e31f2ef6f0 refactor: simplified null checks using optional chaining across TypeScript and Rust modules
- Simplified null/empty checks across TypeScript codebase using optional chaining operator (?.) for improved readability.
- Replaced explicit null checks in validation logic with optional chaining in oauth-discovery, gemini-cli, claude, zai, and lsp modules.
- Updated error handling in Rust command invocation to use double question mark operator (??) for cmd_result.
- Consolidated null validation patterns across tools (bash-skill-urls, browser, gemini-image, resolve) and keybindings using optional chaining.
2026-03-26 14:09:14 +01:00
can1357 4f7b61ab62 feat(coding-agent): added MCP server auto-reconnect with SSE monitoring
- Added auto-reconnect capability for MCP servers with SSE stream monitoring and exponential retry backoff.
- Added tool-level reconnect handling for retriable connection errors (ECONNREFUSED, ECONNRESET, 404/502/503).
- Added `/mcp reconnect <name>` command for manual MCP server recovery.
- Improved reconnect robustness by aborting retries when MCP configuration changes via epoch checking.
- Extended transport reconnect handling to all transport types (stdio, HTTP/SSE) with unified onClose logic.
- Added comprehensive test coverage for MCPManager reconnect behavior and tool-level abort propagation.
2026-03-20 23:48:38 +01:00
can1357 9a37cc8096 refactor(coding-agent/mcp): restructured tool-bridge abort handling and argument normalization
- Replaced custom withAbort() implementation with untilAborted() utility from pi-utils for consistent abort handling.
- Extracted normalizeToolArgs() helper to consolidate argument normalization logic across renderCall(), renderResult(), and execute() methods.
- Added MCPToolCallParams type import to improve type safety for tool parameter handling.
2026-03-20 23:42:02 +01:00
can1357 51228e9811 refactor(coding-agent): restructured MCP reconnection tracking and error handling
- Refactored MCP manager reconnection logic to track configuration changes via reconnectEpoch parameter.
- Consolidated error pattern matching in tool-bridge to use lowercase normalization for consistent comparison.
- Extracted retry provider variables to eliminate repeated ternary expressions in error handling.
- Reformatted code across multiple files for improved readability with consistent multi-line formatting.
- Added test coverage for transport reconnection scenarios including connection reuse after reconnection.
2026-03-20 23:32:07 +01:00
8b028224e3 feat: auto-reconnect MCP servers on connection loss (#482)
* feat: auto-reconnect MCP servers on connection loss

When an HTTP SSE stream drops (server restart, network interruption),
the transport fires onClose, and the manager proactively reconnects
with retry backoff (500ms, 1s, 2s, 4s). Tools are kept in the registry
during reconnection so they remain selected and available to the agent.

If proactive reconnection fails, stale tools stay registered. When the
agent calls one, the tool bridge detects the retriable connection error
(ECONNREFUSED, ECONNRESET, stale session 404/502/503, etc.), triggers
reconnectServer on the manager, and retries the call once on the fresh
connection. Concurrent reconnect attempts for the same server are deduped.

connectToServer now always installs a default onRequest handler for ping
and roots/list (using getProjectDir()), so all connections -- including
short-lived test/probe ones -- properly respond to server-initiated
requests during initialization.

Post-connection setup (resources, prompts, subscriptions) is extracted
into a shared #loadServerResourcesAndPrompts method used by both initial
connection and reconnection paths.

Add /mcp reconnect <name> command for manual recovery after extended
outages where both proactive and reactive reconnection have failed.

* docs: add changelog entry for MCP auto-reconnect

* fix: address P1 review findings in MCP reconnection

- Save server configs before connection attempt so deferred tools can
  reconnect even when the initial connection timed out (P1-1)
- Make waitForConnection() and getConnectionStatus() aware of in-flight
  reconnections so callers wait instead of failing immediately (P1-2)
- Add epoch counter incremented on disconnectAll() and checked in
  connectAndWireServer() to invalidate stale reconnect attempts that
  outlive a manager reset/reload (P1-3)
- Skip servers with pending reconnections in connectServers() to prevent
  parallel connection attempts for the same server

* fix: deferred tool reconnect and non-blocking transport teardown

- DeferredMCPTool.execute now reconnects when getConnection() fails
  ("MCP server not connected"), not only on network errors from
  callTool. Servers that missed the startup window can now be woken
  by the first tool call against their cached tools. (P1-4)
- #doReconnect fire-and-forgets the old transport close instead of
  awaiting it. HttpTransport.close() sends a DELETE with 30s timeout;
  blocking here delayed the first reconnect attempt by that amount
  on every server restart. (P1-5)

* fix: abort-aware reconnect waits and preserve tool selection on reconnect

- Wrap all reconnect() awaits with withAbort(signal) so user
  cancellation (Esc) interrupts the reconnect backoff loop instead
  of blocking for up to 7.5s. Applies to MCPTool (1 site) and
  DeferredMCPTool (2 sites). (P2-1)
- Remove activateDiscoveredMCPTools call from /mcp reconnect handler.
  refreshMCPTools already preserves the user's prior MCP tool
  selection; the extra activation was silently opting into all
  server tools including ones the user had not enabled. (P2-2)

* fix: rebind MCPTool connection after reconnect, add stdio retriable error

- MCPTool.connection is now mutable; after a successful reconnect retry,
  this.connection is rebound to the fresh connection so subsequent calls
  on the same instance (e.g. batched tool calls) use it instead of
  triggering another reconnect cycle. (P2-3)
- Add "Transport closed" to RETRIABLE_PATTERNS. StdioTransport rejects
  pending requests with this message when the subprocess dies, which
  should trigger the reconnect path just like HTTP transport errors. (P2-4)

---------

Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
2026-03-20 23:27:30 +01:00
can1357 51fb0ade87 feat: add discoveryDefaultServers config for MCP discovery mode
fixes #470
2026-03-18 23:05:33 +01:00
21a86a693d feat(mcp): implement roots/list and server-to-client request handling (#474)
* feat(mcp): implement roots/list and server-to-client request handling

Add support for MCP server-to-client JSON-RPC requests across both
stdio and HTTP transports, enabling servers to query client capabilities
such as roots/list during initialization.

Transport layer (types.ts, stdio.ts, http.ts):
- Add onRequest callback to MCPTransport interface for server-initiated
  requests; add toJsonRpcError helper for error code propagation
- Classify incoming messages by checking method+id (request), id-only
  (response), method-only (notification); guard against id:null per
  JSON-RPC 2.0 spec
- StdioTransport: detect server requests in #handleMessage, respond via
  #sendResponse writing JSON-RPC response to subprocess stdin
- HttpTransport: detect server requests via #dispatchSSEMessage across
  all SSE streams (dedicated listener, POST response drain, notify
  piggybacking), respond via #sendServerResponse POST with proper
  Accept header and session ID
- Refactor startSSEListener to resolve once SSE GET connects (not when
  stream ends), enabling await before notifications/initialized; reset
  #sseConnection via .finally() for reconnection after transient failure
- #parseSSEResponse continues reading after capturing the primary
  response to drain piggybacked server requests/notifications; clears
  timeout after capture so drain phase is unbounded
- notify() reads text/event-stream response bodies for piggybacked
  messages; cancels non-SSE response bodies to release connections
- #sendServerResponse includes AbortSignal.timeout and cancels response
  body; fire-and-forget handlers wrapped in try/catch to prevent
  unhandled rejections

Client wiring (client.ts):
- Add onRequest to connectToServer options, wire to transport before
  initialization
- Add awaitable onInitialized hook in initializeConnection, called
  between initialize response (which sets session ID) and initialized
  notification, so SSE stream is open when server sends roots/list
- Pass only signal to transport.request (not full options object)
- Hoist transport ref to outer scope; close on timeout/abort to prevent
  orphaned transports when SSE GET hangs

Manager (manager.ts):
- Wire onRequest handler in connectServers for all MCP connections
- Handle roots/list by returning project CWD as file:// URI via
  pathToFileURL; return -32601 for unsupported methods

Tests (mcp-roots-list.test.ts):
- toJsonRpcError: code extraction, defaults, non-Error values
- Message classification spec tests: request/response/notification/
  unknown dispatch, id:null and id:0 edge cases
- Roots response shape: file:// URI generation, Windows paths, spaces

* fix(mcp): return SSE response immediately instead of blocking on stream drain

The #parseSSEResponse loop continued iterating the SSE stream after
capturing the response for the expected request ID. Since clearTimeout
was called after capture, a server that holds the SSE stream open for
follow-up events (permitted by Streamable HTTP) would block the
request() call indefinitely.

Return the result as soon as it's captured and drain remaining
messages in a detached background task via #readSSEStream, which
already handles dispatch and error swallowing.

* fix(mcp): handle batched JSON-RPC messages in both transports

JSON-RPC 2.0 section 6 allows sending an array of request/notification
objects as a batch. If a server sent a batch, the message classifier
in both transports would fail the 'method in message' check on the
array object and silently drop all contained messages.

Add an Array.isArray guard at the top of #handleMessage (stdio) and
#dispatchSSEMessage (http) that recurses into each element. Defensive
measure — no known MCP server sends batches today, but the guard is
cheap and correct per the JSON-RPC spec.

* fix(mcp): address second Codex review round

- http: break from SSE loop before starting background drain to avoid
  ReadableStream locked error (the for-await iterator still holds the
  reader when #drainSSEBackground was called inline)
- types: toJsonRpcError now accepts plain { code, message } objects,
  not just Error instances, so onRequest handlers can throw structured
  JSON-RPC errors without wrapping in Error
- test: relax Windows path name assertion to toBeTruthy since
  path.basename is platform-dependent for backslash paths; add tests
  for plain-object toJsonRpcError

* fix(mcp): address third Codex review round

- parseSSEResponse: flatten JSON-RPC batch arrays before checking for
  the expected response, so a server that batches the primary response
  with piggybacked requests/notifications in a single SSE event still
  has the response extracted correctly
- sendServerResponse: retry once on 401/403 via onAuthError, matching
  the auth-refresh logic in #executeRequest; prevents server-initiated
  request replies from failing after token expiry on long-lived SSE
  sessions

---------

Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
2026-03-18 23:02:22 +01:00
can1357 66c80fd613 feat: add MCP JSON schema and documentation
fixes #462
2026-03-18 22:43:59 +01:00
can1357 afc2276a12 refactor(ai): extracted OpenAI compat logic into dedicated module
- Extracted OpenAI compatibility detection and resolution logic into dedicated `openai-completions-compat` module.
- Refactored `detectCompat()` and `getCompat()` to delegate to new compat module functions with simplified conditional logic.
- Fixed OAuth redirect URI validation to preserve exact configured values without trailing slash normalization.
- Improved session deletion to return boolean status and display error messages in UI instead of silently failing.
- Added `/session delete` command with Delete key support and confirmation dialogs for session management.
2026-03-17 16:07:26 +01:00
571510bdf3 fix: MCP OAuth exact redirect URIs for Slack-style providers (#454)
* Add exact MCP OAuth redirect URI support

* Allow proxied HTTPS loopback MCP redirects

* Fix MCP OAuth busy-port test determinism

* Make MCP OAuth redirect test deterministic

* fix(coding-agent): honor exact loopback redirect ports

* fix(coding-agent): expand env vars in standalone MCP oauth config

---------

Co-authored-by: can1357 <me@can.ac>
2026-03-17 14:49:23 +01:00
94ad651d17 feat: add MCP tool discovery search (#352)
* Add MCP tool discovery search and live refresh

* Fix MCP discovery review feedback

* Address remaining MCP discovery review comments

* feat: compact MCP discovery search results

* fix: align MCP discovery search contract

* feat: add MCP server tool counts to discovery hints

* fix(agent): corrected stale toolChoice validation against active tools

- Fixed stale forced toolChoice passed to provider after mid-turn tool refresh by validating against active tools.
- Added refreshToolChoiceForActiveTools() to filter invalid tool choices when available tools change.
- Changed getToolChoice config to use computed function instead of static property for dynamic validation.
- Fixed MCP tool selection tracking in coding-agent to distinguish between discovery-enabled and non-discovery sessions.
- Updated search_tool_bm25 to filter already-selected tools before applying limit parameter.

---------

Co-authored-by: can1357 <me@can.ac>
2026-03-16 13:43:43 +01:00
can1357 7818c5c316 feat(coding-agent): enabled interactive input submission and fixed continue path state handling
- Exported `submitInteractiveInput()` function for programmatic submission of user input in interactive mode.
- Fixed continue special path to skip optimistic submission state check for already-started prompts.
- Added 2 test cases covering continue submission behavior and optimistic state cancellation.
2026-03-10 20:15:03 +01:00
can1357 bcfbc710de fix(mcp): force OAuth token refresh on 401/403 auth errors
#resolveAuthConfig only refreshed tokens within the 5-minute pre-expiry
window, so the onAuthError retry path reused stale credentials when
tokens were revoked, clocks skewed, or expires was missing. Add a
forceRefresh parameter and pass true from the auth error handler so
401/403 always triggers an unconditional refresh attempt.
2026-03-10 18:40:18 +01:00
Kevin LoftisandGitHub 6f02d42eb4 feat: add MCP OAuth token refresh (#359)
* Add OAuth token refresh for MCP connections

Proactive refresh with 5-minute buffer before token expiry, plus
retry on 401/403 with automatic token refresh for HTTP transports.
Persist tokenUrl, clientId, and clientSecret in auth config so
refresh can happen without re-prompting the user.

* docs(coding-agent): updated CHANGELOG for MCP OAuth token refresh
2026-03-10 18:31:01 +01:00
can1357 6357b245a5 fix: corrected resource tracking and context cleanup across MCP and virtualization layers
- Fixed resource refresh tracking by storing connection references alongside promises to prevent stale deduplication.
- Fixed update target resolution to explicitly handle missing ompPath and use path.resolve() for consistent normalization.
- Added error handling and logging in Smithery registry detail fetching to gracefully handle failures and track issues.
- Fixed virtualization context cleanup in error paths to prevent partially-started instances from remaining active.
- Fixed API key retrieval to use dynamic provider configuration instead of hardcoded provider string.
- Enhanced test utilities to capture and verify request parameters for improved test coverage and debugging.
2026-03-03 06:07:22 +01:00
c22ba6b49c feat: implemented Smithery Registry for MCP search (#96)
* Implemented Smithery MCP Searchable Registry

* refactor(coding-agent): consolidated MCP registry search and improve error handling

- Extracted `parseCommandArgs` and `stripControlChars` utilities to reduce duplication in MCP command controller. Replaced custom URL opening logic with centralized `openPath` utility across OAuth and registry flows.
- Enhanced Smithery auth error handling to gracefully degrade on file read failures with logging instead of throwing, and improved chmod error reporting. Normalized Smithery API base URL to strip trailing slashes.
- Improved registry search pagination to fetch multiple pages until sufficient results are found, with semantic mode support to preserve API relevance ranking. Deduplicated entries by identity key and applied local sorting only in non-semantic mode.

---------

Co-authored-by: can1357 <me@can.ac>
2026-03-03 04:48:23 +01:00
can1357 f6553a00ff fix(coding-agent): guard stale MCP subscription post-actions 2026-03-03 04:00:30 +01:00