- Client-generated HTTP/MCP/authorization headers win over configured
headers case-insensitively (Agent Plugins §7.2.1) via the new
header-policy fetch wrapper used by the HTTP and legacy SSE transports.
- headerPolicy: "origin-locked" pins configured headers to the configured
URL's origin: never forwarded across cross-origin redirects, and
method-changing redirects of JSON-RPC POSTs are refused.
- envPolicy: "literal" exempts stdio env values (and origin-locked
headers) from config-value resolution: no ambient env-name lookup, no
__omp_shell("command execution, empty values preserved.")
- Moved the coding-agent lock-directory primitive to @oh-my-pi/pi-utils/file-lock
and migrated settings, MCP config-writer, and security store imports.
- Replaced the stats aggregator's parallel ~200-line token/breaker lock protocol
with the shared primitive: dead owners reclaimed immediately, live-but-wedged
owners after STATS_SYNC_LOCK_STALE_MS, unstamped acquisitions after the new
acquireStaleMs grace (10s).
- Shared primitive now treats EPERM kill probes as live owners.
- Rewrote the stats lock-reclamation regressions against the shared protocol
and moved the file-lock contract test into pi-utils.
/mcp reauth read OAuth clientId/clientSecret from the raw, unexpanded config
while URL and resource used expandEnvVarsDeep, so `${VAR}` placeholders were
sent literally to the token exchange. MCPOAuthFlow.exchangeToken() also accepted
any HTTP-success body, storing an empty access token when a provider signals
failure with HTTP 200 (e.g. Slack `{ ok: false, error }`), surfacing only later
as invalid_token.
- Select flow client credentials from runtimeBaseConfig / expanded auth block;
keep the raw placeholder for the persisted config file.
- Reject token responses without a non-empty access_token, including the
sanitized provider error when present.
- Add regression tests for env-expanded reauth credentials and HTTP-200 token
error bodies.
Fixes#7440
- Change the default MCP JSON-RPC request ID format from snowflake strings to sequential integers.
- Update server configuration schema, connection equivalence checks, and tests to reflect the new integer default.
The option only exists in OMP's own config format, so the OMP-owned discovery
providers are the only ones that parse it. Say so in the schema description, the
MCPServerConfigBase doc, and the changelog, and name the config paths where
setting it actually takes effect, so nobody expects a server imported from
another tool's config to honor it.
The option was only present on the transport-facing `MCPServerConfig`, so a
value written in `.omp/mcp.json` or a standalone `.mcp.json` never reached the
transports: discovery normalizes config into the canonical `MCPServer` shape and
`convertToLegacyConfig()` rebuilds the transport config from it, and neither step
knew about the field. Setting `"number"` in the documented config path silently
kept the snowflake-string default, which is the hang the option exists to avoid.
Wire it through the same four places `timeout` already uses: the canonical
`MCPServer` shape, the two OMP-native loaders (with validate-and-warn on an
unrecognized value), and the legacy conversion. Foreign-format providers are
untouched, since the key is OMP-specific.
Also point the `MCPServerConfigBase` doc comment at `RequestIdAllocator` rather
than a helper name that never existed.
Apple's `xcrun mcpbridge` decodes JSON-RPC `id` as an integer only. OMP
mints collision-resistant snowflake strings, so the bridge logs
`mcpbridge.DecodeError Code=1`, never replies, and every request hangs
until it times out (#7053). JSON-RPC 2.0 permits String and Number ids
equally, so both shapes are legal and the string default stays.
Add `requestIdFormat: "string" | "number"` to the shared server config
and honor it in all three transports through one allocator. The string
default is unchanged, so this is inert unless a server opts in.
Verified against Xcode 26.3's bridge: with `"number"`, `initialize`
succeeds and `tools/list` returns all 21 tools; with the default, the
same request times out.
- Native (non-mcp://) resource URIs now pass through byte-for-byte via
rawHref; slash elision applies only to the legacy mcp:// wrapper, so
catalog://root/ style URIs match exact-equality server lookups.
- resources/templates/list failure no longer discards a successful
resources/list (Promise.allSettled; templates retried later).
- Opaque RFC 3986 URIs (urn:doc, custom:item) are recognized by both
the router and read-cli discovery gates, with drive-path and
read-selector false positives guarded.
- Review follow-up for PR #6790.
- Added a suppress load option so disabled servers still claim their
capability key: a project foo with enabled:false shadows a same-named
enabled user foo again, while scope-removed entries drop fully.
- Tool-name collisions now resolve by stable server+tool origin key
instead of manager array order, so reconnect re-appends cannot flip
the routed implementation.
- Review follow-up for PR #6787.
Moved first-wins MCP tool-name deduplication and origin-aware warnings into one shared helper used by startup extension registration, SDK custom-tool assembly, and deferred refreshes.
Added an SDK startup regression proving colliding MCP proxy tools keep the first origin instead of silently overwriting it.
Fixes#6786
Applied the denylist and per-server enabled:false exclusions before connection-equivalence deduplication, alongside project scope, so a disabled higher-priority server can no longer shadow a differently-named equivalent enabled server and leave no connection. Parameterized LoadOptions<T> so the pre-dedup filter sees the typed item.
Fixes#6786
Applied the project-scope filter before connection-equivalence deduplication so a project server can no longer shadow a differently-named but equivalent user server and then be dropped, leaving none.
Fixes#6786
A single hung/slow poll now aborts with TimeoutError after 30s; without
this, that one timeout escaped #waitForSmitheryCliApiKey and dropped the
browser login to the manual API-key fallback instead of retrying until
the 5-minute deadline. Catch isTimeoutError in the poll loop and
continue; export SmitheryCliPollResponse to type the retried response.
MCP OAuth endpoint discovery ran metadata, well-known, and recursive
authorization-server fetches with no AbortSignal, and the Smithery
browser-login poll received a never-aborting signal. An endpoint that
accepts the TCP connection but never responds stalled /mcp add,
/mcp reauth, the add wizard, or /mcp smithery-login indefinitely; the
5-minute login/poll deadlines run only after discovery resolves or
between polls, so a hung fetch never reached them.
- discoverOAuthEndpoints and fetchResourceMetadataScopes gain an optional
signal and wrap every fetch in withTimeoutSignal(DISCOVERY_FETCH_TIMEOUT_MS,
opts?.signal), threaded through the recursive authorization_servers call.
- pollSmitheryCliAuthSession wraps its fetch in
withTimeoutSignal(SMITHERY_POLL_TIMEOUT_MS, signal) so a hung poll aborts
and the loop reaches its 5-minute deadline.
Fixes#4103
Every exported read-modify-write on mcp.json (add/update/remove server, disabled/force-enabled lists) now runs under a per-file withFileLock, so overlapping in-process or cross-process mutations no longer lose updates. writeMCPConfigFile writes to a pid+uuid temp file instead of a shared ${filePath}.tmp, so concurrent writers cannot rename each other's temp out from under them (ENOENT / clobber).
Fixes#4104
MCP-backed tools never declared an explicit strict value, so OpenAI-family
serializers (post-#4336/#4340) had no false to preserve and models over-filled
mutually exclusive optional fields. Task/subagent proxies also rebuilt a raw
tools/call instead of executing through the source MCPTool, bypassing intent
stripping, placeholder pruning, local-URL resolution, reconnect, abort, and
result metadata; strict servers rejected proxied calls with
unrecognized_keys ["i"].
- MCPTool/DeferredMCPTool now declare `readonly strict = false as const`.
- createMCPProxyTools delegates to the current source tool, re-resolved by raw
MCP server/tool metadata so reconnect replacements are honored, and keeps the
Task 60s timeout by combining its abort signal with the caller's.
- Regression coverage: strict flags, proxy parity for i/placeholder shaping,
declared-i passthrough, and reconnect re-resolution.
Fixes#6208
kill(pid, 0) succeeds for a zombie too: a grandchild whose parent (the
killed leader) is gone sits as <defunct> until whatever reaps orphans
gets around to it, which can lag on some hosts. processExists() now
reads the process's ps state and treats a zombie as already reaped
instead of still alive, so the group-kill regression tests assert what
they actually claim to test.
terminateStdioProcess() treated a detached leader's cooperative SIGTERM
exit as proof the whole process group was gone, so close() skipped the
group SIGKILL and left a SIGTERM-trapping/ignoring grandchild running as
an orphan — exactly the process tree this change set out to reap. A
detached transport now always sweeps the group SIGKILL after SIGTERM,
even when the leader itself already exited.
waitForProcessExit() also left its losing Bun.sleep() timer running
after Promise.race settled from the other side, holding the event loop
open for up to the full grace window on every close(). It now uses a
cancellable setTimeout cleared in a finally block.
Adds a regression test spawning a non-trapping leader with a
SIGTERM-trapping grandchild to cover the gap the first fix closes.
Before: StdioTransport.close() did a bare `this.#process.kill()` — a single
direct SIGTERM to the immediate child, with no wait and no escalation. On
Linux (and other non-Windows/non-macOS POSIX hosts), the MCP server is
spawned detached (setsid, its own session leader) so terminal job-control
signals can't stop it. A detached server that traps/ignores SIGTERM — or a
grandchild it spawns inside that session — survived omp process exit and
was orphaned, re-parented to PID 1.
After: close() runs a bounded, idempotent teardown:
1. End stdin first (cooperative EOF) so a well-behaved server can exit on
its own before any signal is sent.
2. Send SIGTERM: to the whole process group (negative-pid `process.kill`)
when this transport actually spawned detached on a POSIX host, else to
the direct child only. A negative-pid signal is never attempted for a
non-detached transport, since it could hit an unrelated group. ESRCH
from the group signal means the group is already gone (treated as
success); any other group-signal failure falls back to a direct-child
signal.
3. Wait up to ~1s for the direct child to exit; if it hasn't, escalate to
SIGKILL (group-or-direct, same rule as step 2) and wait a further
bounded ~0.5s before returning. Total worst case (~1.5s) stays well
inside the ~3s MCP disconnect-all budget in agent-session.ts dispose().
`#process` is captured into a local and nulled before the first `await`, so
repeat/concurrent close() calls see it already cleared and skip re-signaling
— idempotent per the existing contract documented above close().
Extracted the signal/escalate logic into an exported `terminateStdioProcess`
(plus a `KillableSubprocess` structural type, decoupled from the stdio pipe
generics) so tests can drive group-signal escalation with an explicit
`detached` flag — `StdioTransport.connect()` ties `detached` to the host's
real `process.platform` via `resolveStdioSpawnCommand()`, so a POSIX
detached session can't be reproduced end-to-end through `connect()` on a
non-Linux dev/CI host, but a real detached process group can still be
spawned directly on any POSIX host to exercise it.
Tests added to stdio.test.ts: detached child trapping SIGTERM escalates to
SIGKILL; a detached parent's SIGTERM-trapping grandchild is only reached by
the group SIGKILL (proves group, not direct-child-only, signaling); a
well-behaved child closes promptly without escalating; a non-detached
transport never attempts a group signal. Extended
test/mcp-stdio-transport.test.ts's existing close() idempotency coverage
with a case where the first close() had to run the full escalation path.
Fixes#5578.
Only a 403 response that identifies unapproved_client now blocks generateAuthUrl before its clientless probe. Invalid metadata, invalid redirect, generic forbidden, retryable, and server failures preserve the probe path.
Consolidated fallback coverage across 400, 403, 429, and 503 responses.
Fixes#5852
Added #isDefinitiveRegistrationRejection so only non-retryable 4xx DCR errors block generateAuthUrl; 408/425/429 fall through to the clientless authorization probe alongside transport and 5xx failures.
Fixes#5852
Only a 4xx DCR client error blocks generateAuthUrl; transport (status 0) and 5xx failures fall through to the clientless authorization probe so providers accepting client-less authorization keep working.
Fixes#5852
Surfaced rejected dynamic client registration from generateAuthUrl before probing or returning an authorization URL without client_id.
Added coverage for a Cropwise-style 403 unapproved_client response.
Fixes#5852
Building the cmd.exe /c command line and spawning with windowsVerbatimArguments so cmd.exe expansion cannot eat or inject on %VAR%, quote, and metacharacter args (BatBadBut / CVE-2024-24576).
Fixes#5696
MCPManager evicted a server's tools by matching the raw mcp__<name>_
prefix against sanitized tool names. One server's sanitized name can
prefix another's (atlassian vs imported atlassian:atlassian), so every
reconnect of the shorter-named server dropped the sibling's tools and
re-announced them moments later, spamming paired xd:// unmount/mount
notices on each transport flap. Names containing sanitized characters
never prefix-matched at all, leaving stale tools registered after
disconnect. Replacement and removal now match mcpServerName.
Threaded the authorization server's advertised registration endpoint through OAuth discovery, add, reauth, and the client flow instead of deriving metadata from the authorization endpoint.
Added pathful-issuer discovery and end-to-end DCR regression coverage.
Fixes#5267
- Fenced final OAuth refresh update and terminal-disable CAS statements by row id, serialized credential data, active lease owner, and unexpired lease time.
- Passed an AbortSignal through MCP OAuth token refresh and bounded owned refresh operations below the lease TTL while awaiting the aborted fetch to settle.
- Added regressions for stolen-lease update/disable attempts and timed-out MCP token fetch abort behavior.
Fixes#5081
- Added durable SQLite refresh ownership for stored OAuth rows, with canonical re-read before refresh and compare-and-set persistence.
- Routed MCP proactive and forced OAuth refresh through the shared owner so waiters reuse the winner's rotated credential.
- Added MCP regression tests for shared SQLite refresh ownership and stale invalid_grant losers.
Fixes#5081
Clerk and similar providers bind DCR clients to only the scopes declared at
registration. Authorize then requests scopes_supported (including openid),
which rejects with "client is not allowed to request scope 'openid'". Match
Claude Code by sending config.scopes as RFC 7591 scope on the DCR body.