- Removed `href`, `hrefr`, and `hline` Handlebars helpers along with shared hashline anchor state; unused by any template.
- Changed blank lines between ops from silent separators to literal payload lines appended to the open op.
- Added overlapping-delete validation to reject before/after-block patch patterns.
- Simplified hashline prompt doc, removing template-helper examples and tightening rules.
- Added a conditional section in the project system prompt that only appears when context files are present.
- Documented that preloaded agent context files should not be searched for, covering AGENTS.md, CLAUDE.md, and similar rules files.
- Reinforced that those agent/context files are already available in context and should be treated as noise otherwise.
- Adjusted handoff test Promise resolver typings to use non-undefined string values.
- Updated the mocked generateHandoff promise to resolve with a string handoff value.
- Resolved the pending handoff promise with "handoff" in test cleanup to satisfy the contract.
- Added optional skipRoleRank controls to model and canonical-model sort helpers to make role ranking conditional.
- Re-sorted fuzzy-filtered model and canonical model results with role rank skipped so stronger query matches are not outranked by weak default-role matches.
- Short-circuited `agent_end` when `#checkCompaction` deferred handoff, skipping rewind/todo passes and `agent.continue()` race.
- Aborted retry/compaction paths in `AgentSession.dispose()` before draining post-prompt tasks so `/exit` and Ctrl+C no longer hang.
- Added handoff-deadlock regression tests in `agent-session-handoff.test.ts` to prevent reordering races.
- Added non-structural single-line prefix and suffix duplicate checks for `A-B:` replacement ranges.
- Gated the new boundary absorber behind `autoDropPureInsertDuplicates` and integrated it with existing structural and multi-line hashline absorption logic.
- Updated hashline schema documentation, prompts, and tests, and recorded the behavior change in the changelog.
- Removed post-filter re-sorting for canonical model matches so fuzzyFilter's relevance order is preserved.
- Stopped reapplying role, MRU, and version sorting after fuzzy filtering canonical and non-canonical models.
- Documented the stable-sort rationale so fuzzy text relevance remains the primary ordering.
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
- Added a new `approvalMode` argument to CLI parsing with validation for `auto`, `prompt`, and `custom` values.
- Registered `--approval-mode` on the launch command so it appears in generated help output.
- Applied the parsed approval mode as a runtime override on `Settings`, ensuring downstream `tools.approvalMode` reads reflect the CLI value.
- Decouple the per-tool approval gate from extension presence. ExtensionRunner
and the ExtensionToolWrapper that hosts the gate are now constructed
unconditionally in createAgentSession. Previously the runner was only built
when extensionsResult.extensions.length > 0, so the entire approval system
silently disappeared for sessions with no extensions loaded — any
tools.approvalMode: prompt|custom setting was a no-op without feedback.
Today this hole was masked by createAutoresearchExtension always being
pushed inline; the unconditional construction makes the safety invariant
explicit, and a new regression test in approval-mode.test.ts pins it.
- Extend CRITICAL_BASH_PATTERNS to cover remote-fetch-then-execute shapes
that the original `bash <(curl …)` regex missed:
- `source <(curl …)` / `. <(curl …)` (anchored at command boundary so
`find . -name foo` doesn't false-positive)
- `eval "$(curl …)"` / `eval $(curl …)` / `eval `curl …``
Also adds `chmod -R` symbolic-mode forms (`u+x`, `u+rwx,o+w …`) targeting
filesystem root, and `tee` / `tee -a` writes to /etc/{passwd,shadow,sudoers}
(the standard way to write root-owned files without redirect). Benign
forms (`source ./local.sh`, `chmod -R u+x ./build`, `tee /var/log/app.log`,
`eval "$VAR"`) are pinned negative in the test suite.
- Extend formatApprovalPrompt with payload previews for the destructive tools
that previously rendered as bare `Allow tool: <name>`: eval (language +
first cell's code), task (agent + first task's id + assignment), ast_edit
(first op's pattern / replacement / paths), browser (action + tab + url +
code), and write content (alongside path). For `task` in particular this
closes the gap that docs/approval-mode.md's "parent's approval covers the
subagent" claim was waving at — the prompt now actually shows what's being
delegated.
- Tighten isMcpToolName: drop the fallback `|| toolName.includes("__")` so
an extension tool legally named `my__feature` or `pkg__util__do` is no
longer falsely labelled `Origin: MCP server tool` in the approval prompt.
Strict `mcp__` prefix only.
- Revert the cargo-cult `{ autoApprove: true } as AgentToolContext` insertions
in agent-session-python-cleanup.test.ts and sdk-move-cwd.test.ts. The tests
create sessions without passing settings, so the wrapper falls through to
approvalMode "auto" automatically; the explicit flag was unnecessary and
the `as AgentToolContext` cast hid that autoApprove lives on
CustomToolContext, not AgentToolContext.
- Document in commands/launch.ts the dual --auto-approve declaration (oclif
Flags for --help, manual parseArgs for runtime) so a future rename catches
both call sites.
- Promote the subagent caveat in docs/approval-mode.md to a callout near the
top: anything `task` is asked to do runs unattended once the parent task
call is approved.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts → 75 pass / 0 fail
(was 57; +18 cases covering new remote-exec patterns, chmod symbolic, tee
/etc, isMcp negative, and eval/task/ast_edit/browser/write payload previews)
- bun test packages/coding-agent/test/tools/approval-mode.test.ts → 7 pass /
0 fail (was 7; +1 case asserting extensionRunner is always constructed)
- bun tsc --noEmit -p packages/coding-agent → clean
- bun x biome check . → clean
- Windows EBUSY tempdir-cleanup noise in agent-session-python-cleanup and
sdk-move-cwd is pre-existing on this branch (already documented in the
PR body) and absent on Linux CI.
- approval: user 'tool: deny' now wins over critical-pattern override
(the override only tightens allow->prompt; it must never re-arm a denied tool).
- approval: rename hindsight policy keys to match registered tool names
(recall/retain/reflect, not hindsight_recall/hindsight_retain).
- approval: head+tail truncation for bash/ssh command prompts so a
destructive suffix buried after a long benign preamble stays visible.
- task/executor: force tools.approvalMode='auto' in createSubagentSettings
so subagents (which have no UI) cannot deadlock on per-tool prompts;
the parent's approval of the task call is the authorization.
- docs/approval-mode: rewrite so every example surfaces tools.approvalMode
and explains that tools.approval is ignored outside 'custom' mode.
Stale comment said 'layers user config on top' which sounded like custom
mode merged defaults with config. The actual behaviour (and what the user
asked for): in custom mode user config wins; built-in defaults only fill
gaps for tools the user hasn't configured. auto and prompt modes ignore
tools.approval entirely.
New global setting under /settings -> Interaction that controls the tool
approval flow:
auto (default) Skip every approval prompt — yolo. Matches --auto-approve.
prompt Built-in per-tool defaults only. Destructive tools (bash,
edit, write, eval, ssh) require confirmation; read-only
tools auto-allow; tools.approval.<tool> overrides ignored.
custom tools.approval.<tool> config wins. Built-in defaults only
fall back for tools the user hasn't configured. Critical
safety patterns (rm -rf /, fork bombs, curl|bash) still
prompt even when the tool is user-allowed.
The CLI --auto-approve / --yolo flag always wins regardless of the setting,
preserving the automation/CI path.
Wires through ExtensionToolWrapper.execute(): the wrapper reads
tools.approvalMode from settings, derives userPolicies only for custom mode,
and feeds the existing requiresApproval() resolver. Resolution order inside
requiresApproval already places user config above built-in defaults, so
'config wins' falls out naturally in custom mode.
Adds test/tools/approval-mode.test.ts covering all three modes, the CLI
override, the built-in fallback in custom mode, and the critical-pattern
override that fires even when bash is user-allowed.
The approval gate added in 0efa60b7d requires either a UI runner or an
autoApprove context flag. The python-cleanup tests call EvalTool.execute
directly, bypassing the agent loop that normally supplies context.
Pass { autoApprove: true } as AgentToolContext at three direct call sites.
This matches the in-loop behaviour for tests that opt into approval-free
execution and unblocks CI for #1378.
Re-introduces the per-tool approval system from luzidd's commit 39124f3 (which
is no longer reachable from main) and improves it before re-landing.
What's restored:
- ApprovalPolicy (allow/deny/prompt) plus DEFAULT_APPROVAL_POLICIES.
- ACTION_EXCEPTIONS registry (LSP read-only, bash critical patterns).
- getApprovalPolicy() six-level resolution order.
- ExtensionToolWrapper.execute() gate before extension handlers.
- --auto-approve / --yolo CLI flag and tools.approval.<tool> user config.
- docs/approval-mode.md user guide.
What's improved over the original:
- Replaced unchecked 'as any' casts with typed unknown narrowing helpers.
- Validate userConfig values: invalid strings, numbers, etc. fall through to
the built-in default instead of being silently honoured (typo no longer
locks a tool out or grants implicit approval).
- Expanded CRITICAL_BASH_PATTERNS: chmod -R /, chown -R /, bash <(curl ...),
writes to /etc/passwd|shadow|sudoers, shutdown/reboot/halt/init 0,
kill -9 1, nc -e / nc -c reverse shells. Pattern shapes require a
command-position boundary so 'npm run reboot-tests' and 'echo "shutdown the
queue"' don't false-positive.
- Added DEBUG_READONLY_ACTIONS exception so DAP inspection actions (threads,
stack_trace, variables, scopes, read_memory, …) auto-allow while
execution-side actions (launch, attach, continue, evaluate, write_memory,
set_breakpoint, …) still prompt.
- formatApprovalPrompt: labels mcp__<server>__<tool> calls as MCP server
tools, surfaces ssh host + command, recognises the modern § hashline header
for edit, and truncates >240-char fields so a heredoc-sized body cannot
blow out the confirmation dialog.
- Test suite grown from 40 to 57 cases — new coverage for invalid user
config, the extended critical-bash patterns, benign-keyword negatives,
debug exceptions, MCP/ssh prompt formatting, and command truncation.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts -> 57 pass
- bun x biome check . -> clean
- bun run check:ts across all 9 workspaces -> clean
- Updated the auth-gateway OpenAI responses caching test to use shared E2E helper utilities and the common gateway URL constant.
- Initialized and reset in-memory settings in the nested live rendering test fixture to isolate test state between runs.
- Replaced Bun.sleep with scheduler.wait for Node-compatible cancellable sleeps.
- Added module-level timestamp gate to skip yields within 50ms of the last one.
- Threaded AbortSignal through ExponentialYield.sleep to cancel losing timers in race.
- Added tests covering gate behaviour and stray-timer cancellation.
- Rejected negative values in addition to non-numeric ones, falling back to per-server config or default 30s.
- Emitted a logger warning when an invalid env value is ignored.
- Added tests covering negative and non-numeric rejection cases.
- Extended `buildWellKnownUrls` and `#resolveRegistrationEndpoint` to try `/.well-known//` as a third candidate after origin-root and path-prefixed forms.
- Fixed single-segment path handling so `/my-service` is treated as the gateway prefix rather than dropped.
- Fixed missing `await` on `#tryWellKnownForRegistration` that caused path-prefixed fallback to return an unresolved Promise.
- Added tests for single-segment prefix discovery and RFC 8414 path-ful issuer fallback.
- Adjusted hashline natural-order preview handling so op-insert and op-replace tokens now emit inline body content as payload when available.
- Added an inline-body presence check so those op tokens are skipped only if path or payload is missing for that line.
- Dropped the `fileType: natives.FileType.File` restriction so glob searches can return directories as well as files.
- Updated the find tool prompt to document directory results and trailing-slash output.
- Added tests verifying directory matches are included and emitted with a trailing `/`.
- Clamped `formatDuration` to return `0ms` for non-positive, NaN, or infinite inputs.
- Updated usage report rendering to suppress reset countdowns when `resetsAt` is absent or no longer in the future.
- Added unit tests for `formatDuration` covering clamped values and standard duration formatting.
brush_core::interp::setup_open_file_with_contents wrote the entire heredoc/here-string body into an anonymous pipe synchronously before handing the reader to the downstream command. Bodies that exceed the OS pipe buffer (~4 KiB on Windows, 16-64 KiB on macOS) deadlocked the writer forever, and the bash tool tripped its 305 s hard timeout without ever launching the consumer. The Linux fast path still uses F_SETPIPE_SZ to grow the pipe inline; every other platform (and Linux bodies that overflow pipe-max-size) now decouples the write onto a fire-and-forget thread that terminates on drain or BrokenPipe.
Adds a 256 KiB regression test that exercises the worst-case shape (: builtin, which never drains stdin), guarded by tokio::time::timeout(10s) so a regression fails CI fast instead of hanging.
- Migrated OAuth provider authentication from standalone `pi-ai` CLI to in-process `AuthStorage.login()` flow in coding-agent.
- Made provider argument optional for `login` and `logout` commands with interactive provider picker when omitted.
- Added `list` command to enumerate registered OAuth providers with optional `--json` output format.
- Removed `pi-ai` CLI binary and `bin` entry from @oh-my-pi/ai package; library API remains unchanged.
- Updated documentation and examples to reflect new `omp auth-broker` command interface and in-process OAuth flow.
Threaded cache freshness/authoritativeness through #loadCachedStandardProviderModels so dropProviderModels only fires when the cached Vertex project-catalog row is both fresh and authoritative. A stale or non-authoritative snapshot (e.g. after ADC discovery failure rewrote the row with authoritative=0) now keeps the bundled Gemini fallback in place, which would otherwise be the last working catalog in API-key-only environments.
Refs #1412
- Updated `TempDirGuard` creation in grep tests to include PID and an atomic sequence, preventing temp path collisions.
- Removed the `branch` filter from GitHub action run queries so results are matched by `head_sha` only.
- Adjusted run-watch calls to the simplified `fetchRunsForCommit` interface without the branch argument.
Added Google Vertex OpenAI-compatible model discovery with ADC auth and treated authoritative Vertex project catalogs as replacements for bundled Gemini fallbacks in the model registry.
Fixes#1412