Commit Graph

132 Commits

Author SHA1 Message Date
roboomp 0b88ab32f8 fix(mcp): wrap unresolvable Windows commands in cmd.exe for PATHEXT lookup
Bun.spawn -> CreateProcess only appends `.exe` to extensionless names; `.cmd`/`.bat` are never tried. Bare MCP commands like `npx` (which exists only as `npx.cmd` on Windows) crashed the subprocess ~140ms after spawn with ENOENT/EINVAL whenever our own PATH walk couldn't pin the file down (empty `Bun.env.PATH` under a restricted parent process, UNC mounts that reject `fs.access`, locked-down shells).

`resolveStdioSpawnCommand` now routes any unresolved bare command through `cmd.exe /d /s /c` so Windows's PATHEXT search runs Windows-native. Direct-spawn fast path is preserved for resolved `.exe`/`.com` files; the existing `.cmd`/`.bat` wrap is unchanged. The reporter's stated cause ("process.env not merged") was incorrect — the merge has been in place at `transports/stdio.ts:316-319` since well before 16.1.14 — but the symptom they hit is real.

Fixes #3250
2026-06-22 10:49:42 +00:00
roboomp 75b21611f7 fix(mcp): sanitized startup server names
Sanitized MCP server names before status formatting so configured keys cannot leak home paths, tabs, newlines, or oversized text into the TUI.

Fixes #3150
2026-06-20 23:43:16 +02:00
roboomp 8b2012420f fix(mcp): sanitized startup failure status
Sanitized MCP failure text before it reaches the startup status formatter, including tabs, newlines, home paths, and long server output.

Fixes #3150
2026-06-20 21:23:27 +00:00
roboomp 655fed1e48 fix(mcp): updated startup connection status
Emitted MCP connection lifecycle events through the startup event bus so the TUI can replace the initial connecting banner with connected, pending, or failed server state.

Added manager and interactive-mode coverage for mixed success/failure MCP startup updates.

Fixes #3150
2026-06-20 21:14:41 +00:00
can1357 50bc9353aa fix(mcp): keep resources when resources/templates/list unimplemented (#2838) 2026-06-18 02:46:48 +02:00
can1357 adaa032333 style(coding-agent/mcp): reformatted MCP render functions for multiline callback formatting
- Reflowed `renderMCPCall` and `renderMCPResult` `WidthAwareText` calls over multiline formatting only.
2026-06-17 20:42:46 +02:00
can1357 c81adde1f1 fix(coding-agent/mcp): made MCP rendering adapt to available content width
- Wrapped MCP call and result renderers in WidthAwareText so formatting can use runtime content width.
- Computed inline argument preview budget from the available content width instead of a fixed 70-character cap.
- Updated raw text truncation in MCP results to target the measured content width while preserving expand hints and warnings.
2026-06-17 19:16:12 +02:00
jms830 6a0e80ee56 fix(mcp): treat resources/templates/list -32601 as empty, not a resource-load failure 2026-06-16 21:25:03 -04:00
can1357 eaba315cbd security(coding-agent): sanitized artifact names and wrapped extension and MCP tools
- Sanitized artifact filenames by normalizing tool names before composing spill paths.
- Applied `wrapToolWithMetaNotice` to custom tool adapters and RPC-host tools in agent-session setup.
- Wrapped SDK-registered extension/custom tools with the same meta-notice adapter during session creation.
2026-06-17 01:53:24 +02:00
can1357 3f82589ec1 fix: fixed OAuth and profile-boundary regressions across CLI and env handling
- Fixed OAuth credentials to keep unknown fields in schema while preserving existing shape checks.
- Fixed MCP OAuth IDs to be profile-scoped and avoid deleting credentials from non-active profiles.
- Fixed string-flag parsing so PROFILE_BOOTSTRAP_BOUNDARY tokens are not consumed as values.
- Fixed active-profile directory resolution to refresh after env updates so profile .env overrides apply.
2026-06-15 03:20:45 +02:00
Ogrodev f9bc96e96c fix(coding-agent): harden profile auth shipping gaps 2026-06-14 20:30:50 -03:00
Ogrodev 0123a46f83 Merge remote-tracking branch 'upstream/main' into feat/profiles-and-alias
# Conflicts:
#	packages/coding-agent/src/cli/args.ts
2026-06-14 19:10:31 -03:00
metaphorics 8e21f4b41d fix(coding-agent): route MCP connecting banner through the render tree
Deferred MCP discovery wrote 'Connecting to MCP servers: …' straight to process.stderr while the TUI owned the terminal, overdrawing the chat input box border. onMCPConnecting now emits McpConnectingEvent on the mcp:connecting channel; InteractiveMode subscribes and renders it via showStatus (status container), mirroring the LSP-startup pattern. New mcp/startup-events.ts holds the channel, type, and formatMCPConnectingMessage.
2026-06-14 17:24:21 +02:00
usr_bin_roygbiv 15d481c84e fix(mcp): prevent stdio transport close from hanging on blocked read loop 2026-06-14 00:59:20 -05:00
Ogrodev 6450d3b46a Merge remote-tracking branch 'upstream/main' into feat/profiles-and-alias
# Conflicts:
#	packages/coding-agent/src/cli/args.ts
2026-06-13 18:05:16 -03:00
can1357 a2e63c70b0 feat(coding-agent): enabled discovered MCP server auth and namespaced reauth targets
- Allowed colon-separated MCP server names in validation and updated validation errors.
- Added discovery-aware MCP config resolution for /mcp auth, test, and unauth flows.
- Changed streaming behavior so superseded agent_end events don't stop active loaders.
2026-06-13 14:31:33 +02:00
Ogrodev 6c844c86a1 Merge remote-tracking branch 'upstream/main' into feat/profiles-and-alias 2026-06-12 07:50:40 -03:00
can1357 1ceae07764 fix(coding-agent): prevented non-node Windows cmd shims from being launched as node
- Updated the Windows npm shim resolver to resolve shim files against `cwd` and inspect `_prog` before treating a batch wrapper as a node launcher.
- Added a node-only guard so non-node wrappers, such as python shims, fall back to standard cmd.exe execution.
- Adjusted mcp stdio tests with a new non-node shim fixture and a simplified notify race case to validate transport teardown behavior.
2026-06-12 10:01:55 +02:00
roboomp 4938f47254 fix(mcp): preserved cwd precedence for unqualified .cmd commands
Restored cmd.exe's lookup order on Windows so an unqualified MCP command (e.g. server.cmd) checks the configured cwd before iterating PATH, keeping a project-local shim from being shadowed by a same-named global one.
2026-06-12 07:28:13 +00:00
roboomp aa862b8b32 fix(mcp): launched npm cmd shims directly
Resolved Windows npm-generated .cmd MCP shims to their Node entrypoint before spawning so CodeGraph keeps ownership of stdio instead of disconnecting behind a transient cmd.exe wrapper.

Fixes #2367
2026-06-12 07:22:10 +00:00
Ogrodev ef3ae501fb Merge upstream/main into feat/profiles-and-alias 2026-06-11 13:07:52 -03:00
Can Bölük d35efd67e4 Merge branch 'main' into fix-mcp-oauth-resource 2026-06-11 15:53:15 +02:00
roboomp ae49f83913 fix(mcp): hid windows cmd stdio shims
Wrap Windows MCP .cmd/.bat stdio launches through a hidden cmd.exe invocation so PATH shims such as npx keep stdio attached without flashing a console.

Fixes #2287
2026-06-11 04:00:05 +00:00
Ogrodev 2f60eaf938 feat(coding-agent): bind MCP OAuth credentials per profile via url-keyed ids
Store MCP OAuth credentials under deterministic mcp_oauth:<url> ids in each
profile's agent.db with refresh material embedded, so a definition-only entry
in a shared project mcp.json resolves each profile's own credential instead
of profiles clobbering each other's auth.credentialId pointer.

- Refresh material is single-source: embedded credential fields win over the
  config auth block (which may belong to another profile); legacy rows fall
  back to the auth block wholesale
- Wire the 401 refresh hook off the resolvable credential, not the auth
  block, so definition-only bindings refresh mid-session too
- The url-keyed fallback never overrides a pinned Authorization header
- Send prompt=consent by default (oauth.prompt to override, "" to omit) so
  reauth can switch accounts past an active browser session
- /mcp reauth fails fast on stdio transports (with an mcp-remote ~/.mcp-auth
  hint), probes http/sse without OAuth injection, GCs the superseded legacy
  row only after the flow succeeds, and leaves definition-only entries
  untouched on disk
- DCR-issued client secrets stay embedded in the stored credential and are
  never written into config files; user-supplied secrets survive reauth
2026-06-10 18:40:07 -03:00
Clark Tomlinson df97002df9 fix: include MCP OAuth resource indicator 2026-06-10 15:58:25 -04:00
can1357 7b71a6016d fix(coding-agent): routed Windows batch stdio launches through cmd.exe
- Added Windows batch-command detection and COMSPEC-based cmd.exe resolution for MCP stdio spawns.
- Escaped and quoted batch command arguments, then routed .cmd and .bat commands through cmd /d /s /c.
- Updated the stdio transport tests to assert wrapped cmd.exe command arrays and escaping behavior.
2026-06-10 04:40:08 +02:00
roboomp 376675253f fix(mcp): preserved windows cmd argv launch
Resolved the Windows stdio MCP regression by keeping resolved .cmd shims on the direct argv spawn path instead of rewriting them through cmd.exe /c. Added regression coverage for explicit and PATHEXT-resolved codegraph.cmd commands.\n\nFixes #2220
2026-06-10 02:18:35 +00:00
can1357 39c08f5434 fix(coding-agent): stopped web-search query mangling and API-key log leakage
removed the rewrite replacing every 202x with the current year (corrupted CVE ids); MCP request logs redact key/token/secret/auth query params; fetch honors declared charsets, surfaces transport causes, retries 429 once abort-aware, flags mid-stream truncation, stops double-downloading binaries; browser tab reopen/registry/single-flight races fixed, queued opens honor abort, init failures release the temp hold; MCP calls get a default timeout and per-line SSE parse guards.
2026-06-10 01:28:04 +02:00
roboomp db351abe49 fix(mcp): escaped cmd shim quotes
Escaped literal double quotes with cmd caret syntax before invoking Windows .cmd MCP shims so JSON args cannot break out of the quoted argument.

Refs #2174
2026-06-09 08:35:44 +00:00
roboomp 6b4bcb81e4 fix(mcp): preserved percent args in cmd shims
Escaped literal percent signs before joining Windows cmd.exe shim command strings so MCP server args are not consumed by cmd environment expansion.

Refs #2174
2026-06-09 08:30:16 +00:00
roboomp b48960e63e fix(mcp): resolved windows stdio shims
Resolved Windows stdio MCP commands through PATHEXT before spawning so tools installed as .cmd shims launch from bare command configs.

Fixes #2174
2026-06-09 08:16:45 +00:00
can1357 eb1a46baf5 feat: added injectable fetch transport across AI and coding network flows
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
2026-06-09 04:51:17 +02:00
can1357 48e6009b67 Merge remote-tracking branch 'origin/farm/daabf133/mcp-startup-no-block-on-slow-servers' 2026-06-09 00:23:48 +02:00
roboomp 54e7ccb468 fix(coding-agent): unblock omp startup when an MCP server stalls
MCPManager.connectServers used to fall through to an unbounded
Promise.allSettled over every still-pending server without cached tools,
so a single MCP server stuck waiting on the per-request MCP timeout
(OMP_MCP_TIMEOUT_MS, default 30 000 ms) gated the entire UI ready
signal — exactly the 30.282 s stall the reporter observed against
sbox-superdocs in #2100.

Drop the fallback wait. Pending-without-cache servers are left in flight
and their tools surface via the existing background #onToolsChanged ->
refreshMCPTools path the moment the connect completes; failures continue
to log through the background catch handler (gated on
allowBackgroundLogging) so users still see which server failed.

Adds a regression test that spawns an unresponsive stdio MCP fixture
and asserts connectServers returns inside the 250 ms STARTUP_TIMEOUT_MS
window (padded for CI jitter). The same test times out at 15 s without
the patch.

Fixes #2100
2026-06-08 20:41:51 +00:00
can1357 31b6f0bf31 refactor(ai): consolidated provider config into single-source registry
- Derived descriptors, default-model map, env keys, login list, and refresh dispatch from one ProviderDefinition per provider.
- Disabled OpenAI Codex stream obfuscation and interrupted whitespace-only tool-call argument deltas.
- Derived auth-broker callback ports and paste-code login set from the registry.
2026-06-08 18:48:43 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
Guts 83c8105be6 fix(mcp): declare approval tier for MCP tools to prevent hangs in non-yolo mode
MCPTool and DeferredMCPTool now declare approval = 'write' instead of
implicitly defaulting to 'exec'. Without this, the approval system
requires user confirmation for every MCP tool call in non-yolo modes,
but the confirmation prompt never renders in the TUI while streaming,
causing the agent to hang indefinitely.

Also propagate the approval property through customToolToDefinition()
in sdk.ts, which was silently dropping it during CustomTool ->
ToolDefinition conversion.
2026-06-07 16:22:33 +02:00
roboomp 95f64b6142 fix(mcp): clear stale OAuth credential on definitive refresh failure
When an HTTP MCP server returns invalid_grant (or invalid_token / revoked /
plain 401 from the token endpoint) during OAuth refresh, MCPManager
previously logged "MCP OAuth refresh failed, using existing token" and
re-attached the stale access token as Authorization: Bearer on every
subsequent request. The next tool-load 401'd with invalid_token, future
sessions repeated the loop, and the only recovery was to hand-clear the
credential row in agent.db. Reported with Logfire as the trigger; any
remote HTTP MCP that rotates / revokes refresh tokens is affected.

#resolveAuthConfig now reuses pi-ai's isDefinitiveOAuthFailure classifier
(same one auth-broker and AuthStorage use for first-party providers): on
a definitive failure it calls AuthStorage.remove(credentialId), drops the
Bearer entirely, and the next request surfaces a clean auth error so the
user can /mcp reauth <server> (or /mcp unauth) to recover. Transient
failures (network/fetch failed/ECONNREFUSED) still fall back to the
existing token to ride out blips.

Verified with new mcp-manager-oauth-refresh.test.ts (invalid_grant, 401,
transient fallback, happy-path rotation). The full mcp-* test set
(45 tests across 5 files) still passes.

Fixes #1908
2026-06-05 05:47:09 +00:00
can1357 7a7479f7e3 feat(coding-agent/modes): added arrow-key tab switching to dashboards
- Mapped left/right arrows to switch tabs in the extension dashboard.
- Updated agent and extension dashboard footers to show arrow controls.
- Reformatted HTTP transport assignment and reordered an import.
2026-06-04 16:37:50 +02:00
can1357 51865fc1ea test(coding-agent): updated assertions for condensed prompt wording
- Aligned handoff, reminder, and system-prompt expectations with shortened copy.
- Added HTTP transport test for required initialize failures.
- Guarded SSE startup timeout against stale connection races.
2026-06-04 16:35:45 +02:00
roboomp cb2d5b859a fix(coding-agent): restored exa mcp fallback
Restored Exa's unauthenticated MCP fallback when neither auth storage nor EXA_API_KEY provides credentials. Preserved API-key search ordering and bounded the MCP request with the web-search hard timeout. Updated provider copy and regression coverage for the no-key path.\n\nFixes #1860
2026-06-04 13:18:08 +00:00
roboomp cf33b64033 fix(mcp): updated completed tool status icons
MCP tool renderers now merge call and result components so completed blocks replace the pending call header with success or error status.

MCP protocol errors now preserve top-level isError so downstream render state can use the error status consistently.

Added renderer coverage for completed MCP success and error blocks.

Fixes #1855
2026-06-04 12:14:57 +00:00
VoidChecksum bab53df38a fix(coding-agent/mcp): handle async broken-pipe rejections in stdio transport
On Windows, Bun's FileSink surfaces a broken pipe as a *rejected Promise*
(the EPIPE arrives via a processTicksAndRejections tick), not a synchronous
throw. The sync-only `writeFrame()` from #1711 returns `true` for such a
write and lets the rejected promise float; `request()` likewise wrote stdin
without awaiting. Either path escapes as an unhandled rejection, which the
postmortem handler turns into a fatal process.exit(1) when an MCP server
exits mid-handshake.

- writeFrame: also neutralize a rejected Promise returned by write()/flush()
  (covers the notify + #sendResponse paths).
- request(): await write()/flush() so the rejection lands in the existing
  try/catch and rejects the request instead of floating.
- tests: cover the async-rejection path for writeFrame.

Follow-up to #1710 / #1711, which addressed only the synchronous throw. The
existing notify mid-handshake test fails on Windows without this change
because the real FileSink rejects asynchronously.
2026-06-03 10:51:42 +02:00
roboomp 6598c69a2f fix(mcp-stdio): made close() unconditional in resource teardown
Per second review on #1711: when notify()'s write failed, #handleClose()
flipped #connected=false before the throw. connectToServer()'s catch
then called transport.close(), but close() early-returned because
#connected was already false — so #process.kill() and the readLoop
await never ran. A subprocess that closed its stdin without exiting
(parent EPIPE, transport dead, subprocess alive) would leak.

close() no longer guards on #connected for the resource phase. It still
calls #handleClose() once (idempotent — only if #connected is true), then
unconditionally walks the cleanup chain (kill process, null #process,
await + null #readLoop), each step individually guarded so repeat calls
remain no-ops. onClose still fires exactly once per transport lifetime.

Two new tests in StdioTransport.close: (1) close() called after the
read-loop has already EOF'd and torn down still completes cleanup
without throwing and without re-firing onClose; (2) repeated close()
calls fire onClose exactly once and leave #connected=false.
2026-06-02 12:35:35 +00:00
roboomp b870732258 fix(mcp-stdio): surfaced notify() write failures so handshake fails loudly
Per review on #1711: when the write inside notify() fails (FileSink EPIPE
on Windows during the initialize/notifications-initialized race), silently
closing the transport while still resolving the notify promise let
initializeConnection() return a 'connected' handle wrapping a dead
transport. The manager only wires its reconnect onClose handler after
connectToServer() resolves, so a swallowed handshake failure would
neither reconnect nor fail the connection — it would just leak.

notify() now still calls #handleClose() on write failure (so any wired
onClose runs) but additionally throws `Transport closed while sending
notification "<method>"`. The single in-tree notify caller is
initializeConnection() at client.ts:122; the rejection propagates into
connectToServer()'s catch (which closes the transport and rethrows) and
on into the manager's pending-connection error path. #sendResponse() is
unchanged — silent on failure, since a dead subprocess has no use for
the response.

Test coverage updated to document the surfaced-rejection contract and
also assert transport.connected flips to false.
2026-06-02 12:28:07 +00:00
roboomp 3dacaf263f fix(mcp-stdio): swallowed EPIPE from notify and sendResponse stdin writes
StdioTransport.notify() and #sendResponse() wrote to the subprocess's
stdin without try/catch, while the sibling request() already wrapped the
same write/flush. The #connected guard cannot close the race: connect()
sets it synchronously; #handleClose() only clears it from the read loop's
finally after EOF on stdout — strictly later than the parent's next
stdin write. When an MCP server exits between the initialize response and
the notifications/initialized notification, Bun's FileSink throws EPIPE
synchronously on Windows and the async notify wrapper surfaces it as an
unhandled rejection.

Route both call sites through a new writeFrame(stdin, frame) helper that
catches synchronous sink failures and returns a boolean. notify() now
tears the transport down via #handleClose() on failure so the reconnect
machinery engages; #sendResponse() stays silent because a dead subprocess
has no use for the response. The redundant inner try/catch around
#sendResponse in #handleServerRequest() is removed.

Fixes #1710
2026-06-02 12:22:23 +00:00
roboomp 6881c5cc41 fix(mcp): drop stale connection when reconnect breaker trips
Leaving the dead connection in `#connections` made `getConnectionStatus` report `connected` and `waitForConnection` hand a closed transport to callers after the breaker had explicitly suspended the server. Mirror `#doReconnect`'s teardown: detach `onClose`, fire-and-forget `transport.close()`, and drop the entry from `#connections` (plus its in-flight slots in `#pendingConnections`/`#pendingToolLoads`). Tools stay registered in `#tools` so the user can recover with `/mcp reconnect`.

Test asserts `getConnectionStatus("crashy") === "disconnected"` after the burst.

Refs #1592
2026-05-31 15:38:21 +00:00
roboomp 230514d937 fix(mcp): cap automatic reconnect bursts to prevent fork-bomb
A stdio MCP server that completes the initialize + tools/list handshake and then exits cleanly will fire `transport.onClose` on every clean exit, and the old `MCPManager.reconnectServer` path spawned again unconditionally. A misconfigured PHP-shebang MCP (e.g. Laravel Boost in a non-Laravel project) hit this loop and forked 66 487 `php84` processes parented directly to the agent's `bun` PID until macOS force-rebooted.

Add a per-server sliding-window circuit breaker: at most 5 reconnect attempts per 30 s window. The transport `onClose` callback and the per-tool-call retry in `tool-bridge` are subject to the breaker; `/mcp reconnect` passes `{ manual: true }` to reset the window so users can recover after fixing the underlying misconfiguration. Stale `onClose` is detached when the breaker trips so a late EOF event cannot re-arm the loop.

Defended by `mcp-reconnect-storm.test.ts`: a Bun stdio fixture answers the handshake and exits, then asserts the spawn count stays at ≤ 10 (was 127 without the fix).

Fixes #1592
2026-05-31 15:34:29 +00:00
can1357 e831c2c758 chore: reformat 2026-05-30 18:08:51 +02:00
roboomp 2266fdae82 fix(mcp): honor disabled mcp timeouts for sse startup
When the operator disables MCP client-side timeouts via timeout: 0 or OMP_MCP_TIMEOUT_MS=0, do not impose a 1s startup deadline on the optional HTTP GET SSE listener — let the listener wait as long as the server takes so server-to-client messages are not lost.

Refs #1460
2026-05-27 23:37:10 +00:00