Commit Graph

234 Commits

Author SHA1 Message Date
can1357 b94bfba025 feat(mcp): enforced header precedence and origin policy on remote transports
- Client-generated HTTP/MCP/authorization headers win over configured
  headers case-insensitively (Agent Plugins §7.2.1) via the new
  header-policy fetch wrapper used by the HTTP and legacy SSE transports.
- headerPolicy: "origin-locked" pins configured headers to the configured
  URL's origin: never forwarded across cross-origin redirects, and
  method-changing redirects of JSON-RPC POSTs are refused.
- envPolicy: "literal" exempts stdio env values (and origin-locked
  headers) from config-value resolution: no ambient env-name lookup, no
  __omp_shell("command execution, empty values preserved.")
2026-08-07 05:59:36 +02:00
can1357 39bc9de52f style: applied biome import order and formatting 2026-08-03 15:26:20 +02:00
can1357 9fb082bf14 refactor(utils): consolidated file locking into pi-utils file-lock
- Moved the coding-agent lock-directory primitive to @oh-my-pi/pi-utils/file-lock
  and migrated settings, MCP config-writer, and security store imports.
- Replaced the stats aggregator's parallel ~200-line token/breaker lock protocol
  with the shared primitive: dead owners reclaimed immediately, live-but-wedged
  owners after STATS_SYNC_LOCK_STALE_MS, unstamped acquisitions after the new
  acquireStaleMs grace (10s).
- Shared primitive now treats EPERM kill probes as live owners.
- Rewrote the stats lock-reclamation regressions against the shared protocol
  and moved the file-lock contract test into pi-utils.
2026-08-03 15:25:03 +02:00
roboomp c6a057073b fix(mcp): expand env vars in reauth oauth credentials and reject empty tokens
/mcp reauth read OAuth clientId/clientSecret from the raw, unexpanded config
while URL and resource used expandEnvVarsDeep, so `${VAR}` placeholders were
sent literally to the token exchange. MCPOAuthFlow.exchangeToken() also accepted
any HTTP-success body, storing an empty access token when a provider signals
failure with HTTP 200 (e.g. Slack `{ ok: false, error }`), surfacing only later
as invalid_token.

- Select flow client credentials from runtimeBaseConfig / expanded auth block;
  keep the raw placeholder for the persisted config file.
- Reject token responses without a non-empty access_token, including the
  sanitized provider error when present.
- Add regression tests for env-expanded reauth credentials and HTTP-200 token
  error bodies.

Fixes #7440
2026-08-03 00:33:36 +00:00
can1357 13a36f7c83 refactor(coding-agent/mcp): changed default MCP request ID format to sequential integers
- Change the default MCP JSON-RPC request ID format from snowflake strings to sequential integers.
- Update server configuration schema, connection equivalence checks, and tests to reflect the new integer default.
2026-08-01 20:33:11 +02:00
Jérémy Marchand 4476940a4a docs(mcp): document requestIdFormat as OMP-specific
The option only exists in OMP's own config format, so the OMP-owned discovery
providers are the only ones that parse it. Say so in the schema description, the
MCPServerConfigBase doc, and the changelog, and name the config paths where
setting it actually takes effect, so nobody expects a server imported from
another tool's config to honor it.
2026-07-31 21:04:02 +02:00
Jérémy Marchand 3cb6e5df9c fix(mcp): carry requestIdFormat through MCP config discovery
The option was only present on the transport-facing `MCPServerConfig`, so a
value written in `.omp/mcp.json` or a standalone `.mcp.json` never reached the
transports: discovery normalizes config into the canonical `MCPServer` shape and
`convertToLegacyConfig()` rebuilds the transport config from it, and neither step
knew about the field. Setting `"number"` in the documented config path silently
kept the snowflake-string default, which is the hang the option exists to avoid.

Wire it through the same four places `timeout` already uses: the canonical
`MCPServer` shape, the two OMP-native loaders (with validate-and-warn on an
unrecognized value), and the legacy conversion. Foreign-format providers are
untouched, since the key is OMP-specific.

Also point the `MCPServerConfigBase` doc comment at `RequestIdAllocator` rather
than a helper name that never existed.
2026-07-31 21:03:06 +02:00
Jérémy Marchand 8560188314 feat(mcp): let a server opt into integer JSON-RPC request ids
Apple's `xcrun mcpbridge` decodes JSON-RPC `id` as an integer only. OMP
mints collision-resistant snowflake strings, so the bridge logs
`mcpbridge.DecodeError Code=1`, never replies, and every request hangs
until it times out (#7053). JSON-RPC 2.0 permits String and Number ids
equally, so both shapes are legal and the string default stays.

Add `requestIdFormat: "string" | "number"` to the shared server config
and honor it in all three transports through one allocator. The string
default is unchanged, so this is inert unless a server opts in.

Verified against Xcode 26.3's bridge: with `"number"`, `initialize`
succeeds and `tools/list` returns all 21 tools; with the default, the
same request times out.
2026-07-31 21:03:05 +02:00
can1357 857b70fe99 Merge remote-tracking branch 'refs/remotes/pr/6535' into prep/6535
# Conflicts:
#	packages/coding-agent/src/extensibility/extensions/runner.ts
2026-07-30 01:42:24 +02:00
can1357 d16a251777 chore: reorg tests 2026-07-27 16:43:53 +02:00
can1357 137bad2c8b style: applied biome formatting to review follow-up changes 2026-07-27 16:17:04 +02:00
can1357 41ce810cff fix(mcp): preserved native resource URIs and opaque scheme routing
- Native (non-mcp://) resource URIs now pass through byte-for-byte via
  rawHref; slash elision applies only to the legacy mcp:// wrapper, so
  catalog://root/ style URIs match exact-equality server lookups.
- resources/templates/list failure no longer discards a successful
  resources/list (Promise.allSettled; templates retried later).
- Opaque RFC 3986 URIs (urn:doc, custom:item) are recognized by both
  the router and read-cli discovery gates, with drive-path and
  read-selector false positives guarded.
- Review follow-up for PR #6790.
2026-07-27 16:15:02 +02:00
can1357 9448aac3a6 fix(mcp): kept disable precedence and stable tool collision winners
- Added a suppress load option so disabled servers still claim their
  capability key: a project foo with enabled:false shadows a same-named
  enabled user foo again, while scope-removed entries drop fully.
- Tool-name collisions now resolve by stable server+tool origin key
  instead of manager array order, so reconnect re-appends cannot flip
  the routed implementation.
- Review follow-up for PR #6787.
2026-07-27 16:07:23 +02:00
can1357 b0063dd180 Merge PR #6787: fix(mcp): deduplicate aliased server connections (@roboomp) 2026-07-27 15:57:44 +02:00
Dongmen Laohu 2c77c8535a fix(mcp): resolve native resource URIs 2026-07-27 18:59:12 +08:00
roboomp da11d906ff fix(mcp): unified tool collision handling
Moved first-wins MCP tool-name deduplication and origin-aware warnings into one shared helper used by startup extension registration, SDK custom-tool assembly, and deferred refreshes.

Added an SDK startup regression proving colliding MCP proxy tools keep the first origin instead of silently overwriting it.

Fixes #6786
2026-07-27 10:49:59 +00:00
roboomp 6c96a5ee9f fix(mcp): filter disabled servers before dedup
Applied the denylist and per-server enabled:false exclusions before connection-equivalence deduplication, alongside project scope, so a disabled higher-priority server can no longer shadow a differently-named equivalent enabled server and leave no connection. Parameterized LoadOptions<T> so the pre-dedup filter sees the typed item.

Fixes #6786
2026-07-27 10:37:52 +00:00
roboomp 394eaaeae2 fix(mcp): filter project scope before dedup
Applied the project-scope filter before connection-equivalence deduplication so a project server can no longer shadow a differently-named but equivalent user server and then be dropped, leaving none.

Fixes #6786
2026-07-27 10:31:18 +00:00
Anthony "Asterisk" Ambuehl 29625f08c2 feat(mcp): add mcp_notification extension event + multi-listener API
Convert MCPManager's dangling single-slot setOnNotification callback into
a multi-listener API and expose server-initiated MCP notifications as an
extension event so extensions can bridge push-capable MCP servers (e.g.
peer messaging, ticket nudges) into session behavior.

API changes:
- Removed: MCPManager.setOnNotification(handler) — single-slot, zero callers
- Added:   MCPManager.addNotificationListener(listener): () => void
           Multi-listener with per-listener error isolation, returns unsub.
- Added:   'mcp_notification' extension event
           Payload: { server: string; method: string; params: unknown }

Wired in sdk.ts: one listener bridges to extensionRunner.emitMcpNotification,
captured under postmortem for teardown.

Tests: 3 new (multi-listener fanout, error isolation, unsubscribe),
fixture pattern matches neighboring mcp tests. bun check passes (biome +
tsgo).

Docs: extensions.md (new MCP notifications subsection with bridging
example), mcp-runtime-lifecycle.md (Server-initiated notifications
section), CHANGELOG.
2026-07-24 13:36:03 -07:00
can1357 e8502806ff fix(mcp): retry smithery poll timeouts until the authorization deadline
A single hung/slow poll now aborts with TimeoutError after 30s; without
this, that one timeout escaped #waitForSmitheryCliApiKey and dropped the
browser login to the manual API-key fallback instead of retrying until
the 5-minute deadline. Catch isTimeoutError in the poll loop and
continue; export SmitheryCliPollResponse to type the retried response.
2026-07-23 22:15:21 +02:00
can1357 db2601eaff Merge PR #6425: fix(mcp): bound oauth discovery and smithery poll fetches with abort timeouts (@roboomp) 2026-07-23 22:15:21 +02:00
roboomp af3883c052 fix(mcp): bound oauth discovery and smithery poll fetches with abort timeouts
MCP OAuth endpoint discovery ran metadata, well-known, and recursive
authorization-server fetches with no AbortSignal, and the Smithery
browser-login poll received a never-aborting signal. An endpoint that
accepts the TCP connection but never responds stalled /mcp add,
/mcp reauth, the add wizard, or /mcp smithery-login indefinitely; the
5-minute login/poll deadlines run only after discovery resolves or
between polls, so a hung fetch never reached them.

- discoverOAuthEndpoints and fetchResourceMetadataScopes gain an optional
  signal and wrap every fetch in withTimeoutSignal(DISCOVERY_FETCH_TIMEOUT_MS,
  opts?.signal), threaded through the recursive authorization_servers call.
- pollSmitheryCliAuthSession wraps its fetch in
  withTimeoutSignal(SMITHERY_POLL_TIMEOUT_MS, signal) so a hung poll aborts
  and the loop reaches its 5-minute deadline.

Fixes #4103
2026-07-23 19:45:41 +00:00
roboomp 15577406f8 fix(mcp): serialize mcp.json config writes and use unique temp path
Every exported read-modify-write on mcp.json (add/update/remove server, disabled/force-enabled lists) now runs under a per-file withFileLock, so overlapping in-process or cross-process mutations no longer lose updates. writeMCPConfigFile writes to a pid+uuid temp file instead of a shared ${filePath}.tmp, so concurrent writers cannot rename each other's temp out from under them (ENOENT / clobber).

Fixes #4104
2026-07-23 19:40:13 +00:00
can1357 80ec1cb6ae Merge PR #6347: feat: optionally render MCP results as Markdown (@zeroknots) 2026-07-23 11:53:05 +02:00
zeroknots dd510f2ccc feat(coding-agent): render MCP Markdown results 2026-07-23 13:26:00 +07:00
slee1996 2145ab8f8e fix(coding-agent): retry MCP tool auth challenges 2026-07-22 23:44:38 -06:00
roboomp 04179bf0b8 fix(mcp): route task proxies through source tool and mark tools non-strict
MCP-backed tools never declared an explicit strict value, so OpenAI-family
serializers (post-#4336/#4340) had no false to preserve and models over-filled
mutually exclusive optional fields. Task/subagent proxies also rebuilt a raw
tools/call instead of executing through the source MCPTool, bypassing intent
stripping, placeholder pruning, local-URL resolution, reconnect, abort, and
result metadata; strict servers rejected proxied calls with
unrecognized_keys ["i"].

- MCPTool/DeferredMCPTool now declare `readonly strict = false as const`.
- createMCPProxyTools delegates to the current source tool, re-resolved by raw
  MCP server/tool metadata so reconnect replacements are honored, and keeps the
  Task 60s timeout by combining its abort signal with the caller's.
- Regression coverage: strict flags, proxy parity for i/placeholder shaping,
  declared-i passthrough, and reconnect re-resolution.

Fixes #6208
2026-07-21 22:37:12 +00:00
Mathews-Tom f0d31bcf56 chore: merge upstream/main (mechanical, merge-tree verified clean) 2026-07-18 04:29:23 +05:30
can1357 357d683007 Merge PR #5894: fix(mcp): process-group kill and SIGKILL escalate on stdio transport close (@Mathews-Tom) 2026-07-17 21:22:06 +02:00
Mathews-Tom 22f3701d92 test(mcp): distinguish zombie grandchildren from live ones in kill checks
kill(pid, 0) succeeds for a zombie too: a grandchild whose parent (the
killed leader) is gone sits as <defunct> until whatever reaps orphans
gets around to it, which can lag on some hosts. processExists() now
reads the process's ps state and treats a zombie as already reaped
instead of still alive, so the group-kill regression tests assert what
they actually claim to test.
2026-07-18 00:47:14 +05:30
Mathews-Tom 224796b13a fix(mcp): sweep group SIGKILL after a cooperative detached leader exit
terminateStdioProcess() treated a detached leader's cooperative SIGTERM
exit as proof the whole process group was gone, so close() skipped the
group SIGKILL and left a SIGTERM-trapping/ignoring grandchild running as
an orphan — exactly the process tree this change set out to reap. A
detached transport now always sweeps the group SIGKILL after SIGTERM,
even when the leader itself already exited.

waitForProcessExit() also left its losing Bun.sleep() timer running
after Promise.race settled from the other side, holding the event loop
open for up to the full grace window on every close(). It now uses a
cancellable setTimeout cleared in a finally block.

Adds a regression test spawning a non-trapping leader with a
SIGTERM-trapping grandchild to cover the gap the first fix closes.
2026-07-18 00:02:21 +05:30
Mathews-Tom c30c3ab0fb fix(mcp): process-group kill and SIGKILL escalate on stdio transport close
Before: StdioTransport.close() did a bare `this.#process.kill()` — a single
direct SIGTERM to the immediate child, with no wait and no escalation. On
Linux (and other non-Windows/non-macOS POSIX hosts), the MCP server is
spawned detached (setsid, its own session leader) so terminal job-control
signals can't stop it. A detached server that traps/ignores SIGTERM — or a
grandchild it spawns inside that session — survived omp process exit and
was orphaned, re-parented to PID 1.

After: close() runs a bounded, idempotent teardown:
  1. End stdin first (cooperative EOF) so a well-behaved server can exit on
     its own before any signal is sent.
  2. Send SIGTERM: to the whole process group (negative-pid `process.kill`)
     when this transport actually spawned detached on a POSIX host, else to
     the direct child only. A negative-pid signal is never attempted for a
     non-detached transport, since it could hit an unrelated group. ESRCH
     from the group signal means the group is already gone (treated as
     success); any other group-signal failure falls back to a direct-child
     signal.
  3. Wait up to ~1s for the direct child to exit; if it hasn't, escalate to
     SIGKILL (group-or-direct, same rule as step 2) and wait a further
     bounded ~0.5s before returning. Total worst case (~1.5s) stays well
     inside the ~3s MCP disconnect-all budget in agent-session.ts dispose().

`#process` is captured into a local and nulled before the first `await`, so
repeat/concurrent close() calls see it already cleared and skip re-signaling
— idempotent per the existing contract documented above close().

Extracted the signal/escalate logic into an exported `terminateStdioProcess`
(plus a `KillableSubprocess` structural type, decoupled from the stdio pipe
generics) so tests can drive group-signal escalation with an explicit
`detached` flag — `StdioTransport.connect()` ties `detached` to the host's
real `process.platform` via `resolveStdioSpawnCommand()`, so a POSIX
detached session can't be reproduced end-to-end through `connect()` on a
non-Linux dev/CI host, but a real detached process group can still be
spawned directly on any POSIX host to exercise it.

Tests added to stdio.test.ts: detached child trapping SIGTERM escalates to
SIGKILL; a detached parent's SIGTERM-trapping grandchild is only reached by
the group SIGKILL (proves group, not direct-child-only, signaling); a
well-behaved child closes promptly without escalating; a non-detached
transport never attempts a group signal. Extended
test/mcp-stdio-transport.test.ts's existing close() idempotency coverage
with a case where the first close() had to run the full escalation path.

Fixes #5578.
2026-07-17 23:22:06 +05:30
roboomp 189ac462cc fix(mcp): narrowed failed DCR block to unapproved clients
Only a 403 response that identifies unapproved_client now blocks generateAuthUrl before its clientless probe. Invalid metadata, invalid redirect, generic forbidden, retryable, and server failures preserve the probe path.

Consolidated fallback coverage across 400, 403, 429, and 503 responses.

Fixes #5852
2026-07-17 14:56:56 +00:00
roboomp 490686602c fix(mcp): excluded retryable 4xx DCR statuses from reauth block
Added #isDefinitiveRegistrationRejection so only non-retryable 4xx DCR errors block generateAuthUrl; 408/425/429 fall through to the clientless authorization probe alongside transport and 5xx failures.

Fixes #5852
2026-07-17 14:49:51 +00:00
roboomp 7faa6eb2dd fix(mcp): scoped reauth block to definitive DCR rejections
Only a 4xx DCR client error blocks generateAuthUrl; transport (status 0) and 5xx failures fall through to the clientless authorization probe so providers accepting client-less authorization keep working.

Fixes #5852
2026-07-17 14:45:56 +00:00
roboomp 33f2f00d58 fix(mcp): blocked reauth after failed client registration
Surfaced rejected dynamic client registration from generateAuthUrl before probing or returning an authorization URL without client_id.

Added coverage for a Cropwise-style 403 unapproved_client response.

Fixes #5852
2026-07-17 14:39:15 +00:00
can1357 3a4220a16a merge PR #5699 via eval/pr-5699: fix(mcp): escape Windows cmd shim command path too 2026-07-17 04:39:14 +02:00
roboomp e509fc3cbc fix(mcp): escape Windows cmd shim command path too
Applied the same cmd.exe percent/quote neutralization to the resolved command token so a % in the shim path is not expanded before launch.

Fixes #5696
2026-07-16 13:49:14 +00:00
roboomp 0788933110 fix(mcp): escape Windows cmd shim args against injection
Building the cmd.exe /c command line and spawning with windowsVerbatimArguments so cmd.exe expansion cannot eat or inject on %VAR%, quote, and metacharacter args (BatBadBut / CVE-2024-24576).

Fixes #5696
2026-07-16 13:39:42 +00:00
roboomp 967befdf42 fix(mcp): corrected Windows cmd shim spawning
Passed batch commands and arguments separately to cmd.exe /c so cmd.exe no longer strips a wrapper quote into the command token.

Fixes #5696
2026-07-16 12:58:05 +00:00
Kormákur 4a510a912f fix(mcp): match tool ownership by server name, not tool-name prefix
MCPManager evicted a server's tools by matching the raw mcp__<name>_
prefix against sanitized tool names. One server's sanitized name can
prefix another's (atlassian vs imported atlassian:atlassian), so every
reconnect of the shorter-named server dropped the sibling's tools and
re-announced them moments later, spamming paired xd:// unmount/mount
notices on each transport flap. Names containing sanitized characters
never prefix-matched at all, leaving stale tools registered after
disconnect. Replacement and removal now match mcpServerName.
2026-07-16 10:47:56 +00:00
can1357 ebe79d6f53 fix(mcp): retain DCR metadata fallback 2026-07-14 22:58:47 +02:00
can1357 80f9329e31 Merge PR #5456: fix(mcp): preserve discovered registration endpoint (@roboomp) 2026-07-14 22:58:47 +02:00
roboomp fb98465923 fix(mcp): preserved discovered registration endpoint
Threaded the authorization server's advertised registration endpoint through OAuth discovery, add, reauth, and the client flow instead of deriving metadata from the authorization endpoint.

Added pathful-issuer discovery and end-to-end DCR regression coverage.

Fixes #5267
2026-07-14 17:46:24 +00:00
can1357 4b5c32a092 Merge PR #4950: fix(mcp): resolve local image paths for tool calls (@roboomp) 2026-07-14 18:45:22 +02:00
can1357 4c3df0ee39 Merge remote-tracking branch 'origin/farm/903c642e/fix-stdio-tcc-spawn' 2026-07-11 07:38:17 +02:00
can1357 1758a6c749 Merge remote-tracking branch 'origin/farm/6d5f8413/fix-mcp-oauth-refresh-race' 2026-07-11 00:10:43 +02:00
roboomp b60dc669ea fix(auth): fenced oauth refresh writes
- Fenced final OAuth refresh update and terminal-disable CAS statements by row id, serialized credential data, active lease owner, and unexpired lease time.
- Passed an AbortSignal through MCP OAuth token refresh and bounded owned refresh operations below the lease TTL while awaiting the aborted fetch to settle.
- Added regressions for stolen-lease update/disable attempts and timed-out MCP token fetch abort behavior.

Fixes #5081
2026-07-10 21:49:18 +00:00
roboomp cf021ad393 fix(auth): serialized mcp oauth refreshes
- Added durable SQLite refresh ownership for stored OAuth rows, with canonical re-read before refresh and compare-and-set persistence.
- Routed MCP proactive and forced OAuth refresh through the shared owner so waiters reuse the winner's rotated credential.
- Added MCP regression tests for shared SQLite refresh ownership and stale invalid_grant losers.

Fixes #5081
2026-07-10 20:25:49 +00:00
Victor Araújo bce6aa89a4 fix(mcp): included OAuth scopes in dynamic client registration
Clerk and similar providers bind DCR clients to only the scopes declared at
registration. Authorize then requests scopes_supported (including openid),
which rejects with "client is not allowed to request scope 'openid'". Match
Claude Code by sending config.scopes as RFC 7591 scope on the DCR body.
2026-07-10 15:08:28 -03:00