Commit Graph
257 Commits
Author SHA1 Message Date
Guts 72e11ca5e9 fix(sdk): bind dynamic approval callbacks to original tool instance 2026-06-07 16:58:25 +02:00
Guts 83c8105be6 fix(mcp): declare approval tier for MCP tools to prevent hangs in non-yolo mode
MCPTool and DeferredMCPTool now declare approval = 'write' instead of
implicitly defaulting to 'exec'. Without this, the approval system
requires user confirmation for every MCP tool call in non-yolo modes,
but the confirmation prompt never renders in the TUI while streaming,
causing the agent to hang indefinitely.

Also propagate the approval property through customToolToDefinition()
in sdk.ts, which was silently dropping it during CustomTool ->
ToolDefinition conversion.
2026-06-07 16:22:33 +02:00
can1357 9bd9e3127e feat(coding-agent): added /tan background forking with prompt cache inheritance
- Added `/tan` slash command registration and interactive handling.
- Added TanCommandController validation and async task scheduling for `/tan` dispatch.
- Added session cloning that suppresses breadcrumbs, copies artifacts, and handles abort cleanup.
- Added `promptCacheKey` support in Agent and inherited `providerPromptCacheKey` in session creation.
2026-06-07 06:52:15 +02:00
can1357 cba6641299 feat: enabled resolver-based API key retries with refresh and rotation
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
2026-06-07 04:20:48 +02:00
can1357 f552ce4e6d perf(coding-agent): deferred heavy module imports to startup paths
- Lazy-loaded OTEL SDK, HTML export, TTSR, and autoresearch modules.
- Made resolveMemoryBackend async to import backends on demand.
- Replaced backend resolution with direct settings reads for rekey checks.
2026-06-06 23:25:18 +02:00
can1357 6efe86c07b fix(coding-agent): honored boolean env flag overrides over settings
- Made PI_INTENT_TRACING, PI_AUTO_QA, PI_PY, and PI_JS take precedence when set.
- Fell back to config when the env flag is unset instead of ORing.
- Surfaced PI_PY=0/PI_JS=0 in the disabled-backend error messages.
2026-06-06 16:30:09 +02:00
can1357 2e425b7b65 docs(coding-agent): documented auto discovery mode behavior
- Described "auto" default gating MCP tools past 40-tool threshold.
- Noted late resolution in createAgentSession after registry exists.
- Updated legacy mcp.discoveryMode mapping to MCP-only.
2026-06-06 01:14:14 +02:00
can1357 f5a938f859 feat(coding-agent): added auto tool discovery mode
- Made "auto" the default, hiding MCP tools past 40-tool threshold.
- Centralized discovery mode resolution in shared mode helper.
- Activated search tool in createAgentSession once full registry exists.
2026-06-06 01:12:26 +02:00
can1357 9114607ce6 feat(coding-agent): surfaced active model in system prompt
- Rendered the active model identifier into the project prompt.
- Rebuilt the cached base prompt on model switch to avoid staleness.
2026-06-05 23:08:44 +02:00
can1357 0340604f64 Merge remote-tracking branch 'origin/farm/9c8b047b/fix-async-job-manager-singleton-overwrite' 2026-06-05 11:41:15 +02:00
roboomp be9e5218ed fix(sdk): cleared async singleton after startup failures
If createAgentSession failed after installing a newly created AsyncJobManager but before AgentSession took ownership, the process-global singleton stayed installed. The new singleton guard then caused the next top-level session to skip constructing a scoped manager, disabling async bash/task support.

The startup-error cleanup now clears the singleton only when it still points at the newly created manager, disposes that manager, and then continues the existing registry/kernel cleanup. The regression test forces a startup failure after singleton installation and verifies the next top-level session can create and use its own async manager.
2026-06-05 09:34:10 +00:00
roboomp ac304eacdc fix(sdk): scoped async job snapshots to sessions
AgentSession now stores the same scoped AsyncJobManager reference that tools receive: owning top-level sessions use their constructed manager, subagents inherit the parent's manager, and secondary in-process top-level sessions get no manager when a singleton is already live.

getAsyncJobSnapshot and ACP delivery drains now use that scoped manager instead of AsyncJobManager.instance(), so secondary sessions cannot report or drain the primary session's background jobs. The regression test covers a secondary session created while the primary has a Main-owned running job.
2026-06-05 09:29:30 +00:00
can1357 2dab082a68 fix(tui): restored live block boundary reporting to TUI
- Appended newly sealed transcript blocks to native scrollback once.
- Deferred only the active live block during ED3-risk streaming.
- Hardened snapshot TTL parsing to handle whitespace-only env values.
2026-06-05 11:29:27 +02:00
roboomp eda5eebe71 fix(sdk): route bash/task/job through ToolSession.asyncJobManager to keep secondary sessions isolated
Per PR review on #1926: a secondary in-process top-level createAgentSession() that exposes bash/task/job tools would still call AsyncJobManager.instance() at execute time, register on the primary's manager, and have the primary's onJobComplete enqueue results into the primary's yieldQueue — corrupting the owning session's conversation.

ToolSession now carries an asyncJobManager reference scoped to its session: the constructed manager for top-level sessions, the inherited singleton for subagents (so their bash/task completions still flow into the spawning conversation as before), and undefined for secondary in-process top-level sessions that found a singleton already installed. bash, task, and job tools resolve the manager through ToolSession instead of the process-global singleton, so a secondary session whose tools attempt async work fails fast with the standard "Async job manager unavailable" error instead of contaminating the primary.
2026-06-05 09:22:18 +00:00
can1357 b8a602ac2a test(auth): switched OAuth ranking tests to weighted selection
- Replaced top-rank assertions with weighted-preference distribution checks.
- Added cases for equal-priority balancing and 2x best-bucket cap.
- Added coding-agent snapshot-cache boot and seed tests.
2026-06-05 11:15:32 +02:00
can1357 5004ba057f feat(auth-broker): added encrypted local snapshot cache
- Added AES-GCM cache for at-rest broker snapshots keyed on token, with URL as additional data.
- Added `onSnapshot` hook to RemoteAuthCredentialStore for persisting applied snapshots.
- Exposed cache read/write and TTL defaults through the coding-agent re-exports.
- Added `getAuthBrokerSnapshotCachePath` with `OMP_AUTH_BROKER_SNAPSHOT_CACHE` override.
2026-06-05 11:09:47 +02:00
roboomp 36db2530a9 fix(sdk): keep primary AsyncJobManager when secondary top-level session disposes
Any in-process secondary createAgentSession() (e.g. the Agent Control Center's create flow in agent-dashboard.ts) was constructing its own AsyncJobManager, overwriting the process-global singleton, and then clearing it on its own dispose. The primary session still held its #ownedAsyncJobManager reference, but AsyncJobManager.instance() was undefined for the rest of the process — the task async path hard-failed with "Async execution is enabled but no async job manager is available" and only a full restart cleared it.

- sdk.ts: skip constructing/installing a second AsyncJobManager when a singleton is already live, so secondary top-level sessions share the owning session's manager instead of clobbering it.\n- agent-session.ts: scope #cancelOwnAsyncJobs so a secondary session inheriting the singleton with the default MAIN_AGENT_ID can no longer cancel the primary session's running bash/task jobs at dispose time. Subagents still reach the inherited singleton via their unique agent ids; the owning session still cancels its own jobs through #ownedAsyncJobManager.

Fixes #1923
2026-06-05 09:06:53 +00:00
can1357 4a9ee14e33 feat(coding-agent): wrapped mid-turn steers in wire-only envelope
- Marked steering user messages and wrapped them pre-LLM so the model sees a `` envelope.
- Kept transcripts and persisted history with the user's original raw text.
- Restored `AssistantMessageComponent` stable-prefix completion API.
2026-06-04 17:41:56 +02:00
roboomp 15d4d7fcf9 fix(providers): enabled append-only auto for xiaomi
Centralized append-only context auto-mode resolution so SDK sessions, interactive sessions, and the status display use the same provider checks. Auto mode now enables for Xiaomi/SGLang hosts and explicit stored-request compat signals while preserving DeepSeek behavior.

Added focused regression coverage for Xiaomi Token Plan SGLang HiCache endpoints and explicit on/off behavior.

Fixes #1851
2026-06-04 11:59:47 +00:00
can1357 3f22f939a2 feat(coding-agent/plan-mode): shared approved plan context with spawned subagents
- Added `loadOverallPlanReference` to resolve a session plan reference from local storage and skip empty or missing files.
- Updated task execution to read the active plan reference (except in plan mode) and pass it into each spawned subagent.
- Extended the subagent system prompt and session SDK/tools plumbing so subagents receive and render the approved plan path and contents.
2026-06-04 04:10:53 +02:00
roboompandcan1357 dbdc77697c fix(coding-agent): retried session resume even when initial restore failed
Post-extension session-model retry now covers the case where the initial restore failed entirely (e.g. saved default unavailable, last active role supplied by an extension) and the settings default filled in the active model. Also recomputes thinking-level from full precedence against the reclaimed model so a fallback model's defaultLevel does not become sticky.\n\nFixes #1649
2026-06-02 09:15:45 +02:00
roboompandcan1357 0494454528 fix(coding-agent): retried role restore after extension providers
Initial startup resume runs before extension providers register, so a role model supplied by an extension fell back to the saved default. Retry the preferred session-model candidates once provider registrations are processed and re-resolve thinking level for the new model.\n\nFixes #1649
2026-06-02 09:15:45 +02:00
roboompandcan1357 00adc9c5d6 fix(coding-agent): fell back to saved default on resume
Try the last active role model first, then the saved default model when the role model cannot be restored during session switching or startup resume.\n\nFixes #1649
2026-06-02 09:15:45 +02:00
roboompandcan1357 bb0991e6f3 fix(coding-agent): restored last active role model
Use the last model_change role when resuming an existing session instead of always restoring models.default. Covered both /resume switchSession and startup continue paths.\n\nFixes #1649
2026-06-02 09:15:44 +02:00
can1357 384a206737 refactor(task): replaced numeric-prefix ids with name-first agent output ids
- Changed `AgentOutputManager` to use requested names verbatim, adding `-2`/`-3` suffixes only on repeats (e.g. `Anna`, `Anna-2`).
- Renamed main agent id from `0-Main` to `Main`; nested ids now use dot notation without numeric prefix (e.g. `Parent.Child`).
- Updated task widget to render dotted hierarchy as `Parent>Child` breadcrumb without leading index.
- Resume scan now tracks seen names instead of a counter to avoid clobbering prior outputs.
2026-06-02 06:50:03 +02:00
can1357 68430dee5c chore: renamed mnemosyne package to mnemopi
- Updated package name, directory, and binary from mnemosyne to mnemopi.
- Updated all lockfile references and workspace paths accordingly.
2026-05-31 08:45:12 +02:00
can1357 2ddc9c5bc9 feat(coding-agent): added turn-budget parsing, multipliers and hard caps
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
2026-05-31 08:03:45 +02:00
can1357 7613c2a913 feat(coding-agent): expanded eval execution and bridge paths with workflow helpers
Expand eval execution/bridge paths with args/log/phase/budget support, add workflow helpers (parallel/pipeline), expose usage statistics, and add eval integration tests and docs.
2026-05-31 07:40:18 +02:00
can1357 d1bd14f020 feat(coding-agent): propagated mcpManager and localProtocolOptions to subagents
- Stored mcpManager and localProtocolOptions on ToolSession so nested subagents inherit them without relying on process-global singletons.
- TaskTool now uses the session's localProtocolOptions and mcpManager when spawning sub-tasks, falling back to defaults if absent.
2026-05-31 06:49:20 +02:00
can1357 b48b825344 fix(coding-agent): process @file before session creation; drop built-in flag-name list
Two review fixes for the extension-flag/initial-prompt work:

1. @file ordering — `processFileArguments` runs `process.exit(1)` on a
   missing/unreadable file. It had been moved after `createSession`, which
   writes the terminal breadcrumb eagerly (SessionManager.create →
   #newSessionSync), so `omp @missing.md "x"` left a junk session/breadcrumb
   behind before exiting.

   Resolve extension-registered CLI flags BEFORE creating the session: load the
   session's extensions up front (new `loadSessionExtensions` helper, the single
   source of createAgentSession's discovery-branch logic), build an
   ExtensionFlagSink straight from the loaded extensions + runtime, re-parse
   argv, then process @file args — all before any session exists. The loaded
   result is handed back to createAgentSession via `preloadedExtensions` (now
   checked before `disableExtensionDiscovery`, so it can't double-load) and the
   same EventBus is shared, so no extra work. This keeps the P1#1 fix
   (`--flag @value` is the flag's value, not a file) while failing fast with no
   session side effects.

2. "Can we avoid the big list of names?" — removed the hand-maintained
   `BUILTIN_FLAG_NAMES` set (and its stale "rejected at registration" doc).
   `applyExtensionFlags` now always falls back to recovering a flag's value from
   argv when parseArgs didn't surface it; the recovery scan mirrors parseArgs's
   consumption rules (flag-looking space-form values stay their own flag) and is
   a no-op for flags that were absent or already surfaced, so no list of
   built-in names is needed.

Adds `ExtensionRunner.aggregateFlags` (static) so getFlags and the CLI's
pre-session sink share one implementation.

Tests: pre-session flag resolution via the exact main.ts sink pattern;
list-free recovery of an arbitrary colliding built-in (`--model`); and the
flag-looking-value rule. Verified typecheck + extension/runner/acp suites.
2026-05-31 06:23:15 +02:00
can1357 5344bcbc69 fix(coding-agent): persisted resolved auto thinking level on session resume
- Auto classification now writes the concrete effort to the session log after the first real user turn.
- Resumed sessions restore the last resolved effort instead of reverting to pending auto.
- Added `dedupeReply` opt-out flag for ephemeral turn reply deduplication.
2026-05-31 04:10:52 +02:00
can1357 7f866a48a8 feat(coding-agent): added per-turn AUTO_THINKING in coding-agent session
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.
2026-05-31 03:32:21 +02:00
can1357 caeaf4e4d2 feat(coding-agent): added builtin default rules and disable controls
- Added 14 bundled TTSR rules (TypeScript and Rust conventions) embedded into the binary via the lowest-priority `builtin-defaults` provider.
- Extracted rule bucketing into `bucketRules` with support for `disabledRules` and `builtinRules` settings.
- Added `ttsr.builtinRules` and `ttsr.disabledRules` settings to control which rules are active per session.
2026-05-31 02:09:55 +02:00
can1357 6d5aba8397 feat: implemented Mnemosyne memory state typing for prompt generation
- Resolved memory backend before instruction assembly and used it to build developer instructions.
- Added a `memoryRootEnabled` prompt option when the memory backend id is `local`.
- Typed Mnemosyne memory state types and updated `registerMnemosyneState()` to call `setMnemosyneSessionState()`.
- Documented that `@oh-my-pi/pi-mnemosyne/diagnose` export was added in the mnemosyne changelog.
2026-05-30 16:55:53 +02:00
can1357 99365385aa feat(mnemosyne): added mnemosyne parent-state sync in delegated sessions
- Added parentMnemosyneSessionState propagation from session state through SDK, executor, and task options into nested sessions.
- Added getMnemosyneSessionState() and rekeying logic to refresh Mnemosyne IDs during session sync, switch, and restore.
- Added Mnemosyne reset and teardown cleanup on unaliasing or restoration to avoid stale state.
2026-05-30 16:22:12 +02:00
can1357 89b33823f3 feat(memory): wired Mnemosyne backend into recall/retain/reflect tools
- Extended tool factories to activate on `memory.backend === "mnemosyne"` in addition to `hindsight`.
- Implemented Mnemosyne execution paths in all three tools using `recallEnhanced`, `remember`, and `beam.formatContext`.
- Exposed `getMnemosyneSessionState` on `ToolSession` and wired it through `createAgentSession`.
- Added usage guidance for `recall`, `retain`, and `reflect` to Mnemosyne static instructions.
- Replaced Hindsight-only contract tests with expanded suite covering both backends.
2026-05-30 14:47:46 +02:00
can1357 b4238b10d3 fix: resolved auth-gateway handling of 429 usage-limit responses
- Classified usage-limit gateway responses as `429 rate_limit_error` in auth handling paths.
- Aligned auth-gateway and pi-native key retrieval with derived `sessionId` for `getApiKey` lookups.
- Handled usage-limit auth failures by rotating credentials with retry hints and returning undefined when none available.
- Replaced stream auth checks with retryable-upstream logic for 401 and usage-limit errors before content.
- Expanded `extractRetryHint` parsing for `~`, `sec`, `ms`, and minute/hour units.
- Added coverage for classifyGatewayError, retry-hint parsing variants, and stream-auth retry edge cases.
2026-05-28 02:56:54 +02:00
can1357 c5055d6623 feat(ai): added OpenRouter routing-variant suffix support
- Added `openrouterVariant` option to `SimpleStreamOptions` and `OpenAICompletionsOptions` to append routing suffixes (`:nitro`, `:floor`, `:online`, `:exacto`) to OpenRouter model IDs at request time.
- Skips appending when the model ID already carries an explicit colon-suffix.
- Exposed `providers.openrouterVariant` setting in the coding-agent UI under Settings → Providers.
- Plumbed through `pi-native-server` forwarder and `AgentSession` options preparation.
2026-05-27 18:41:47 +02:00
can1357 ff94f91104 feat: added extraBody support, xAI fixes, and image provider updates
- Added `extraBody` merging into OpenAI Responses request params.
- Fixed xAI OAuth redirect URI to fail fast on port conflicts.
- Exposed `antigravity` and `xai` as explicit `providers.image` options.
- Added `isImageProviderPreference` guard, replacing inline string checks.
- Fixed TTS tool to resolve output path relative to cwd and require write approval.
2026-05-27 15:20:55 +02:00
can1357 a3b27d0cdb chore(ai): fixup xAI cherrypick 2026-05-27 15:15:53 +02:00
cognitiveandcan1357 bde0114f87 feat(coding-agent): add xAI Grok Voice TTS tool
Adds packages/coding-agent/src/tools/tts.ts: a CustomTool that POSTs to
https://api.x.ai/v1/tts using the shared xAI credentials helper
(supports both SuperGrok OAuth and plain XAI_API_KEY).

Built-in voices: ara, eve (default), leo, rex, sal. xAI also accepts
custom voice IDs (the schema does not enum-restrict voice_id). Output
codec inferred from output_path suffix (.wav → wav, else mp3). Max
15,000 characters per request. Composes the callers abort signal with
a 60s timeout fence.

Wired into sdk.ts immediately after the image-gen tool registration,
matching the await logger.time(...) pattern.

Ported from NousResearch/hermes-agent (MIT) — tools/tts_tool.py
L167-171 (constants) and L896-959 (_generate_xai_tts).

Op: extend
2026-05-27 15:08:13 +02:00
roboomp e2431d59dc fix(coding-agent): kept hidden resolve tool registered when plan mode is enabled
createAgentSession() removed the hidden `resolve` tool from the registry
whenever no active tool advertised `deferrable: true`. Plan mode dispatches
its plan-approval `resolve { action: "apply", extra: { title } }` call
through a standing handler installed by InteractiveMode (no deferrable tool
involved), so read-only plan-mode toolsets (e.g. `read`, `search`, `find`,
`web_search`) silently activated plan mode without `resolve`. The agent had
no callable tool to submit the finalized plan and got stuck on the post-turn
tool-decision reminder.

Keep `resolve` registered whenever `plan.enabled` is true so the standing
handler always has a callable tool. The hidden flag still prevents `resolve`
from appearing in the active tool set until plan mode (or a deferrable tool's
preview action) opts in.

Fixes #1428
2026-05-27 07:00:47 +00:00
can1357 e4a16451ec feat(coding-agent): added coding-agent approval types and mode options
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
2026-05-26 21:52:16 +02:00
can1357 7e7cb39170 chore: drop dead extensionRunner ternaries in sdk.ts 2026-05-26 20:53:35 +02:00
oldschoolaandcan1357 f5273eee6f fix(coding-agent): address PR #1378 review findings
- Decouple the per-tool approval gate from extension presence. ExtensionRunner
  and the ExtensionToolWrapper that hosts the gate are now constructed
  unconditionally in createAgentSession. Previously the runner was only built
  when extensionsResult.extensions.length > 0, so the entire approval system
  silently disappeared for sessions with no extensions loaded — any
  tools.approvalMode: prompt|custom setting was a no-op without feedback.
  Today this hole was masked by createAutoresearchExtension always being
  pushed inline; the unconditional construction makes the safety invariant
  explicit, and a new regression test in approval-mode.test.ts pins it.

- Extend CRITICAL_BASH_PATTERNS to cover remote-fetch-then-execute shapes
  that the original `bash <(curl …)` regex missed:
  - `source <(curl …)` / `. <(curl …)` (anchored at command boundary so
    `find . -name foo` doesn't false-positive)
  - `eval "$(curl …)"` / `eval $(curl …)` / `eval `curl …``
  Also adds `chmod -R` symbolic-mode forms (`u+x`, `u+rwx,o+w …`) targeting
  filesystem root, and `tee` / `tee -a` writes to /etc/{passwd,shadow,sudoers}
  (the standard way to write root-owned files without redirect). Benign
  forms (`source ./local.sh`, `chmod -R u+x ./build`, `tee /var/log/app.log`,
  `eval "$VAR"`) are pinned negative in the test suite.

- Extend formatApprovalPrompt with payload previews for the destructive tools
  that previously rendered as bare `Allow tool: <name>`: eval (language +
  first cell's code), task (agent + first task's id + assignment), ast_edit
  (first op's pattern / replacement / paths), browser (action + tab + url +
  code), and write content (alongside path). For `task` in particular this
  closes the gap that docs/approval-mode.md's "parent's approval covers the
  subagent" claim was waving at — the prompt now actually shows what's being
  delegated.

- Tighten isMcpToolName: drop the fallback `|| toolName.includes("__")` so
  an extension tool legally named `my__feature` or `pkg__util__do` is no
  longer falsely labelled `Origin: MCP server tool` in the approval prompt.
  Strict `mcp__` prefix only.

- Revert the cargo-cult `{ autoApprove: true } as AgentToolContext` insertions
  in agent-session-python-cleanup.test.ts and sdk-move-cwd.test.ts. The tests
  create sessions without passing settings, so the wrapper falls through to
  approvalMode "auto" automatically; the explicit flag was unnecessary and
  the `as AgentToolContext` cast hid that autoApprove lives on
  CustomToolContext, not AgentToolContext.

- Document in commands/launch.ts the dual --auto-approve declaration (oclif
  Flags for --help, manual parseArgs for runtime) so a future rename catches
  both call sites.

- Promote the subagent caveat in docs/approval-mode.md to a callout near the
  top: anything `task` is asked to do runs unattended once the parent task
  call is approved.

Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts → 75 pass / 0 fail
  (was 57; +18 cases covering new remote-exec patterns, chmod symbolic, tee
  /etc, isMcp negative, and eval/task/ast_edit/browser/write payload previews)
- bun test packages/coding-agent/test/tools/approval-mode.test.ts → 7 pass /
  0 fail (was 7; +1 case asserting extensionRunner is always constructed)
- bun tsc --noEmit -p packages/coding-agent → clean
- bun x biome check . → clean
- Windows EBUSY tempdir-cleanup noise in agent-session-python-cleanup and
  sdk-move-cwd is pre-existing on this branch (already documented in the
  PR body) and absent on Linux CI.
2026-05-26 20:53:35 +02:00
oldschoolaandcan1357 4d26453a0b feat(coding-agent): restore per-tool approval policies with safer defaults
Re-introduces the per-tool approval system from luzidd's commit 39124f3 (which
is no longer reachable from main) and improves it before re-landing.

What's restored:
- ApprovalPolicy (allow/deny/prompt) plus DEFAULT_APPROVAL_POLICIES.
- ACTION_EXCEPTIONS registry (LSP read-only, bash critical patterns).
- getApprovalPolicy() six-level resolution order.
- ExtensionToolWrapper.execute() gate before extension handlers.
- --auto-approve / --yolo CLI flag and tools.approval.<tool> user config.
- docs/approval-mode.md user guide.

What's improved over the original:
- Replaced unchecked 'as any' casts with typed unknown narrowing helpers.
- Validate userConfig values: invalid strings, numbers, etc. fall through to
  the built-in default instead of being silently honoured (typo no longer
  locks a tool out or grants implicit approval).
- Expanded CRITICAL_BASH_PATTERNS: chmod -R /, chown -R /, bash <(curl ...),
  writes to /etc/passwd|shadow|sudoers, shutdown/reboot/halt/init 0,
  kill -9 1, nc -e / nc -c reverse shells. Pattern shapes require a
  command-position boundary so 'npm run reboot-tests' and 'echo "shutdown the
  queue"' don't false-positive.
- Added DEBUG_READONLY_ACTIONS exception so DAP inspection actions (threads,
  stack_trace, variables, scopes, read_memory, …) auto-allow while
  execution-side actions (launch, attach, continue, evaluate, write_memory,
  set_breakpoint, …) still prompt.
- formatApprovalPrompt: labels mcp__<server>__<tool> calls as MCP server
  tools, surfaces ssh host + command, recognises the modern § hashline header
  for edit, and truncates >240-char fields so a heredoc-sized body cannot
  blow out the confirmation dialog.
- Test suite grown from 40 to 57 cases — new coverage for invalid user
  config, the extended critical-bash patterns, benign-keyword negatives,
  debug exceptions, MCP/ssh prompt formatting, and command truncation.

Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts -> 57 pass
- bun x biome check . -> clean
- bun run check:ts across all 9 workspaces -> clean
2026-05-26 20:53:33 +02:00
roboomp 5955eb13f7 style: bun run fix 2026-05-26 15:14:48 +00:00
roboomp 7487eb330b fix(task): activate yield tool when subagent has explicit tool list
Plan-mode subagents (and any subagent with an explicit `agent.tools` array)
were given the `yield` tool in the registry but not in
`agent.state.tools`. The session prompts and idle reminders still
demanded a `yield` call to terminate, so the model would reason
"there doesn't seem to be a yield tool available" and the turn went
nowhere.

`createTools` correctly appends `yield` to the registry when
`requireYieldTool: true`, but `createAgentSession` then derived the
active tool list from `options.toolNames` directly, dropping `yield`
again. Mirror the invariant already enforced in
`parseAgentFields` (discovery/helpers.ts): when `requireYieldTool` is
set and the caller passes an explicit list, append `yield` to it
before normalization.

Fixes #1408
2026-05-26 15:14:27 +00:00
can1357 5364a9bfd0 refactor(tool-discovery): simplified tool discovery API by removing MCP-specific shims
- Removed deprecated MCP-specific type aliases and functions from tool-discovery module, consolidating to unified generic tool discovery API.
- Migrated session and SDK code to use generic filterBySource() and collectDiscoverableTools() instead of MCP-specific variants.
- Removed deprecated interface members including hasQueuedMessages(), FocusPane, AcpBuiltinCommandRuntime, and legacy settings methods.
- Updated test suites to use renamed generic discovery methods and removed back-compat test coverage for legacy MCP shapes.
2026-05-26 15:27:05 +02:00
can1357 8a5b3e9552 feat(eval): added shared executor inheritance for subagents with concurrent async cells
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
2026-05-26 14:37:56 +02:00