- Decouple the per-tool approval gate from extension presence. ExtensionRunner
and the ExtensionToolWrapper that hosts the gate are now constructed
unconditionally in createAgentSession. Previously the runner was only built
when extensionsResult.extensions.length > 0, so the entire approval system
silently disappeared for sessions with no extensions loaded — any
tools.approvalMode: prompt|custom setting was a no-op without feedback.
Today this hole was masked by createAutoresearchExtension always being
pushed inline; the unconditional construction makes the safety invariant
explicit, and a new regression test in approval-mode.test.ts pins it.
- Extend CRITICAL_BASH_PATTERNS to cover remote-fetch-then-execute shapes
that the original `bash <(curl …)` regex missed:
- `source <(curl …)` / `. <(curl …)` (anchored at command boundary so
`find . -name foo` doesn't false-positive)
- `eval "$(curl …)"` / `eval $(curl …)` / `eval `curl …``
Also adds `chmod -R` symbolic-mode forms (`u+x`, `u+rwx,o+w …`) targeting
filesystem root, and `tee` / `tee -a` writes to /etc/{passwd,shadow,sudoers}
(the standard way to write root-owned files without redirect). Benign
forms (`source ./local.sh`, `chmod -R u+x ./build`, `tee /var/log/app.log`,
`eval "$VAR"`) are pinned negative in the test suite.
- Extend formatApprovalPrompt with payload previews for the destructive tools
that previously rendered as bare `Allow tool: <name>`: eval (language +
first cell's code), task (agent + first task's id + assignment), ast_edit
(first op's pattern / replacement / paths), browser (action + tab + url +
code), and write content (alongside path). For `task` in particular this
closes the gap that docs/approval-mode.md's "parent's approval covers the
subagent" claim was waving at — the prompt now actually shows what's being
delegated.
- Tighten isMcpToolName: drop the fallback `|| toolName.includes("__")` so
an extension tool legally named `my__feature` or `pkg__util__do` is no
longer falsely labelled `Origin: MCP server tool` in the approval prompt.
Strict `mcp__` prefix only.
- Revert the cargo-cult `{ autoApprove: true } as AgentToolContext` insertions
in agent-session-python-cleanup.test.ts and sdk-move-cwd.test.ts. The tests
create sessions without passing settings, so the wrapper falls through to
approvalMode "auto" automatically; the explicit flag was unnecessary and
the `as AgentToolContext` cast hid that autoApprove lives on
CustomToolContext, not AgentToolContext.
- Document in commands/launch.ts the dual --auto-approve declaration (oclif
Flags for --help, manual parseArgs for runtime) so a future rename catches
both call sites.
- Promote the subagent caveat in docs/approval-mode.md to a callout near the
top: anything `task` is asked to do runs unattended once the parent task
call is approved.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts → 75 pass / 0 fail
(was 57; +18 cases covering new remote-exec patterns, chmod symbolic, tee
/etc, isMcp negative, and eval/task/ast_edit/browser/write payload previews)
- bun test packages/coding-agent/test/tools/approval-mode.test.ts → 7 pass /
0 fail (was 7; +1 case asserting extensionRunner is always constructed)
- bun tsc --noEmit -p packages/coding-agent → clean
- bun x biome check . → clean
- Windows EBUSY tempdir-cleanup noise in agent-session-python-cleanup and
sdk-move-cwd is pre-existing on this branch (already documented in the
PR body) and absent on Linux CI.
The approval gate added in 0efa60b7d requires either a UI runner or an
autoApprove context flag. The python-cleanup tests call EvalTool.execute
directly, bypassing the agent loop that normally supplies context.
Pass { autoApprove: true } as AgentToolContext at three direct call sites.
This matches the in-loop behaviour for tests that opt into approval-free
execution and unblocks CI for #1378.
- Replaced external watchdog timers with per-request SDK timeouts for first-event budget across OpenAI, Anthropic, and Azure providers.
- Keyed Python shared kernels by (sessionId, cwd) to prevent cross-directory state bleed.
- Deduplicated concurrent cold-start session acquisition for JS and Python executors.
- Moved `isOpenAIResponsesProgressEvent` to shared module and scoped display output routing per run for interleaved async cells.
- Added `providerRetryWait` and `retryWait` hooks to stream/usage options so tests bypass real scheduler delays.
- Parameterized GitHub Copilot poll intervals and Copilot model retry base delay for fast test execution.
- Replaced `Bun.sleep`/`setTimeout` polling loops with `AbortSignal` event listeners in agent session tests.
- Consolidated auth-gateway E2E helpers into a shared `test/helpers` module, eliminating duplicated `checkGatewayAvailable` implementations.
- Migrated credential-disabled tests from SQLite-backed stores to an in-memory store, removing temp-dir lifecycle overhead.
- Added inline hashline parse and apply support for `<` prepend and `+` append operations with prefix+suffix edits.
- Added fail-fast behavior to reject inline modify ops combined with delete or replace on same line.
- Renamed HASHLINE_* and mode symbols to HL_* in prompt tooling, read/search checks, and prompt templates.
- Standardized separators to `PI_HL_SEP`/`HL_EDIT_SEP` and fixed `HL_BODY_SEP='|'`, updating parser formatting behavior.
- Updated benchmark subtype constants and python cleanup test setup to use HL_* values and AgentRegistry mock failure injection.
The warmup path no longer produces prelude docs, so the cached-session
warmup it implemented added no value over the create-on-first-execute
path that withKernelSession already covers. Remove warmPythonEnvironment,
the backend warm() hook, the eval-tool warmup loop, the createTools
warmup preflight, and the forcePythonWarmup option. Simplify
ExecutorBackendCallOptions into ExecutorBackendExecOptions since execute
is now the only consumer.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
fixed retained-kernel restart and owner cleanup edge cases during recovery and disposal
tracked async user_python hooks during disposal-sensitive execution paths and hardened startup warmup tracking
strengthened cleanup and kernel lifecycle regressions to remove deadlocks, false positives, and timing flakes
scoped retained-kernel ownership to agent sessions and cleaned it up on session disposal
cleaned up warmed python owners on session startup failure and rejected new direct and tool-based python starts during disposal, including async hook, preflight, and warmup races