- Added support for multi-range line selectors on URLs (e.g., `:5-10,20-30`) and combining `:raw` mode with line range selectors.
- Added support for line range selectors on directory listings with offset and limit parameters.
- Fixed `:raw` selector being ignored for JSON and feed URLs and directory listing line selectors dropping offset parameter.
- Added clear error message for line offset beyond directory listing end.
- Refactored URL parsing and directory reading to support multiple comma-separated ranges and improved line-based slicing logic.
- Added comprehensive test coverage for multi-range selectors, raw mode combinations, and directory range operations.
- Replaced `LINE↑`/`LINE↓`/`A-B:` op sigils with unified `A-B:` anchor + `|`/`↑`/`↓` payload sigils.
- Added `mode: "replacement"` tag to insert edits so the applier distinguishes replace-bucket from insert-bucket lines.
- Removed lenient fallbacks (implicit continuation, inline payload acceptance, escaped delimiter stripping).
- Updated grammar, prompt, tokenizer, parser, applier, and messages to match the new format.
- Changed two test expectations from exact string match (toBe) to substring match (toContain) for the '(no output)' text.
- This allows tests to pass when the output contains additional content beyond the expected string.
- Measured bash wall-clock duration for direct, terminal-bridge, and interactive execution paths.
- Recorded wall time in result notices and details, then stripped the duplicated literal notice during shell rendering.
- Updated the renderer to include wall time in the status label and added tests for the new wall-time behavior.
- In `resolveApproval`, yolo mode now returns the user policy directly (`allow`/`prompt`/`deny`) and ignores tool `override` prompts.
- Updated approval-mode and approval unit tests to match the new behavior for critical bash patterns under yolo and auto-approve.
- Updated docs and settings metadata to describe yolo as user-policy-driven rather than override-driven.
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
- Decouple the per-tool approval gate from extension presence. ExtensionRunner
and the ExtensionToolWrapper that hosts the gate are now constructed
unconditionally in createAgentSession. Previously the runner was only built
when extensionsResult.extensions.length > 0, so the entire approval system
silently disappeared for sessions with no extensions loaded — any
tools.approvalMode: prompt|custom setting was a no-op without feedback.
Today this hole was masked by createAutoresearchExtension always being
pushed inline; the unconditional construction makes the safety invariant
explicit, and a new regression test in approval-mode.test.ts pins it.
- Extend CRITICAL_BASH_PATTERNS to cover remote-fetch-then-execute shapes
that the original `bash <(curl …)` regex missed:
- `source <(curl …)` / `. <(curl …)` (anchored at command boundary so
`find . -name foo` doesn't false-positive)
- `eval "$(curl …)"` / `eval $(curl …)` / `eval `curl …``
Also adds `chmod -R` symbolic-mode forms (`u+x`, `u+rwx,o+w …`) targeting
filesystem root, and `tee` / `tee -a` writes to /etc/{passwd,shadow,sudoers}
(the standard way to write root-owned files without redirect). Benign
forms (`source ./local.sh`, `chmod -R u+x ./build`, `tee /var/log/app.log`,
`eval "$VAR"`) are pinned negative in the test suite.
- Extend formatApprovalPrompt with payload previews for the destructive tools
that previously rendered as bare `Allow tool: <name>`: eval (language +
first cell's code), task (agent + first task's id + assignment), ast_edit
(first op's pattern / replacement / paths), browser (action + tab + url +
code), and write content (alongside path). For `task` in particular this
closes the gap that docs/approval-mode.md's "parent's approval covers the
subagent" claim was waving at — the prompt now actually shows what's being
delegated.
- Tighten isMcpToolName: drop the fallback `|| toolName.includes("__")` so
an extension tool legally named `my__feature` or `pkg__util__do` is no
longer falsely labelled `Origin: MCP server tool` in the approval prompt.
Strict `mcp__` prefix only.
- Revert the cargo-cult `{ autoApprove: true } as AgentToolContext` insertions
in agent-session-python-cleanup.test.ts and sdk-move-cwd.test.ts. The tests
create sessions without passing settings, so the wrapper falls through to
approvalMode "auto" automatically; the explicit flag was unnecessary and
the `as AgentToolContext` cast hid that autoApprove lives on
CustomToolContext, not AgentToolContext.
- Document in commands/launch.ts the dual --auto-approve declaration (oclif
Flags for --help, manual parseArgs for runtime) so a future rename catches
both call sites.
- Promote the subagent caveat in docs/approval-mode.md to a callout near the
top: anything `task` is asked to do runs unattended once the parent task
call is approved.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts → 75 pass / 0 fail
(was 57; +18 cases covering new remote-exec patterns, chmod symbolic, tee
/etc, isMcp negative, and eval/task/ast_edit/browser/write payload previews)
- bun test packages/coding-agent/test/tools/approval-mode.test.ts → 7 pass /
0 fail (was 7; +1 case asserting extensionRunner is always constructed)
- bun tsc --noEmit -p packages/coding-agent → clean
- bun x biome check . → clean
- Windows EBUSY tempdir-cleanup noise in agent-session-python-cleanup and
sdk-move-cwd is pre-existing on this branch (already documented in the
PR body) and absent on Linux CI.
- approval: user 'tool: deny' now wins over critical-pattern override
(the override only tightens allow->prompt; it must never re-arm a denied tool).
- approval: rename hindsight policy keys to match registered tool names
(recall/retain/reflect, not hindsight_recall/hindsight_retain).
- approval: head+tail truncation for bash/ssh command prompts so a
destructive suffix buried after a long benign preamble stays visible.
- task/executor: force tools.approvalMode='auto' in createSubagentSettings
so subagents (which have no UI) cannot deadlock on per-tool prompts;
the parent's approval of the task call is the authorization.
- docs/approval-mode: rewrite so every example surfaces tools.approvalMode
and explains that tools.approval is ignored outside 'custom' mode.
New global setting under /settings -> Interaction that controls the tool
approval flow:
auto (default) Skip every approval prompt — yolo. Matches --auto-approve.
prompt Built-in per-tool defaults only. Destructive tools (bash,
edit, write, eval, ssh) require confirmation; read-only
tools auto-allow; tools.approval.<tool> overrides ignored.
custom tools.approval.<tool> config wins. Built-in defaults only
fall back for tools the user hasn't configured. Critical
safety patterns (rm -rf /, fork bombs, curl|bash) still
prompt even when the tool is user-allowed.
The CLI --auto-approve / --yolo flag always wins regardless of the setting,
preserving the automation/CI path.
Wires through ExtensionToolWrapper.execute(): the wrapper reads
tools.approvalMode from settings, derives userPolicies only for custom mode,
and feeds the existing requiresApproval() resolver. Resolution order inside
requiresApproval already places user config above built-in defaults, so
'config wins' falls out naturally in custom mode.
Adds test/tools/approval-mode.test.ts covering all three modes, the CLI
override, the built-in fallback in custom mode, and the critical-pattern
override that fires even when bash is user-allowed.
Re-introduces the per-tool approval system from luzidd's commit 39124f3 (which
is no longer reachable from main) and improves it before re-landing.
What's restored:
- ApprovalPolicy (allow/deny/prompt) plus DEFAULT_APPROVAL_POLICIES.
- ACTION_EXCEPTIONS registry (LSP read-only, bash critical patterns).
- getApprovalPolicy() six-level resolution order.
- ExtensionToolWrapper.execute() gate before extension handlers.
- --auto-approve / --yolo CLI flag and tools.approval.<tool> user config.
- docs/approval-mode.md user guide.
What's improved over the original:
- Replaced unchecked 'as any' casts with typed unknown narrowing helpers.
- Validate userConfig values: invalid strings, numbers, etc. fall through to
the built-in default instead of being silently honoured (typo no longer
locks a tool out or grants implicit approval).
- Expanded CRITICAL_BASH_PATTERNS: chmod -R /, chown -R /, bash <(curl ...),
writes to /etc/passwd|shadow|sudoers, shutdown/reboot/halt/init 0,
kill -9 1, nc -e / nc -c reverse shells. Pattern shapes require a
command-position boundary so 'npm run reboot-tests' and 'echo "shutdown the
queue"' don't false-positive.
- Added DEBUG_READONLY_ACTIONS exception so DAP inspection actions (threads,
stack_trace, variables, scopes, read_memory, …) auto-allow while
execution-side actions (launch, attach, continue, evaluate, write_memory,
set_breakpoint, …) still prompt.
- formatApprovalPrompt: labels mcp__<server>__<tool> calls as MCP server
tools, surfaces ssh host + command, recognises the modern § hashline header
for edit, and truncates >240-char fields so a heredoc-sized body cannot
blow out the confirmation dialog.
- Test suite grown from 40 to 57 cases — new coverage for invalid user
config, the extended critical-bash patterns, benign-keyword negatives,
debug exceptions, MCP/ssh prompt formatting, and command truncation.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts -> 57 pass
- bun x biome check . -> clean
- bun run check:ts across all 9 workspaces -> clean
- Replaced external watchdog timers with per-request SDK timeouts for first-event budget across OpenAI, Anthropic, and Azure providers.
- Keyed Python shared kernels by (sessionId, cwd) to prevent cross-directory state bleed.
- Deduplicated concurrent cold-start session acquisition for JS and Python executors.
- Moved `isOpenAIResponsesProgressEvent` to shared module and scoped display output routing per run for interleaved async cells.
- Removed deprecated MCP-specific type aliases and functions from tool-discovery module, consolidating to unified generic tool discovery API.
- Migrated session and SDK code to use generic filterBySource() and collectDiscoverableTools() instead of MCP-specific variants.
- Removed deprecated interface members including hasQueuedMessages(), FocusPane, AcpBuiltinCommandRuntime, and legacy settings methods.
- Updated test suites to use renamed generic discovery methods and removed back-compat test coverage for legacy MCP shapes.
- Added `cwd` and `env` optional parameters to kernel execution API for runtime working directory and environment variable control.
- Implemented runtime environment setup in Python runner with `_apply_request_runtime()` to apply cwd and env from request before code execution.
- Enhanced SIGINT handler management with `active_executions` counter and `_begin_exec_sigint()` / `_end_exec_sigint()` functions to prevent state mutation during concurrent execution.
- Changed `SearchRenderArgs.paths` parameter type from `string[]` to `string | string[]` to accept single string paths.
- Added comprehensive test coverage for kernel cwd updates, timeout interruption safety, and SystemExit handling in shared executor sessions.
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.
- Removed vim edit mode and automatically map existing vim configurations to hashline mode.
- Deleted VimTool class, VimEngine implementation, and all vim-specific editing logic (2409 lines).
- Removed vim mode from EditMode union type, edit tool strategies, and configuration schemas.
- Deleted vim parser, command handler, buffer manager, and renderer modules.
- Updated documentation and tests to remove vim mode references and add deprecation mapping.
Adds the three regression tests the PR body claimed but #1389 had not actually shipped:
- $-prefixed identifier resolution: asserts resolveSymbolColumn("$store")
on a line with bar$store + $store returns the standalone column (16),
not the substring inside the compound identifier (7).
- create→edit ordering on the same URI: the motivating LSP §3.16.2 case
("Extract to new file" code actions emitting [CreateFile, TextDocumentEdit]
for the same URI). Pre-fix this threw ENOENT.
- Folder-delete subtree flush: mirror of the existing folder-rename test
for the delete arm — child-file edits must land before the parent folder
goes away.
Also adds a CHANGELOG sentence noting that when a WorkspaceEdit supplies
both `changes` and `documentChanges`, the new code uses documentChanges
exclusively per LSP §3.16.2 (previously the two were merged).
- BARE_IDENTIFIER_RE accepts $-prefixed names ($store, $count, RxJS/Svelte/
Angular). Word-boundary check now fires on those names, so searches no
longer return the offset inside a compound identifier like bar$store.
- applyWorkspaceEdit walks documentChanges in declared order, flushing
per-URI text edits immediately before any subsequent resource op for the
same URI. Folder rename/delete ops flush every pending URI under the
affected subtree. Rename ops also flush pending edits queued against
renameOp.newUri (and descendants) before fs.rename runs.
- Legacy changes-map-only WorkspaceEdit payloads are unchanged.
Refactor:
- session.ts: extract mapDebugpyMissingModule helper; replace the duplicated
inline check in launch/attach catch blocks. Add jsdoc on DapStartRequestFailure.settled
documenting per-call ownership and how throwPreferredDapStartError consumes it.
- path-utils.ts: replace the no-op keepOpaqueResourceUri branch with an
OPAQUE_RESOURCE_SCHEMES Set so the structure carries the intent. Functionally
equivalent; new opaque schemes become a one-line Set change.
Tests:
- dap-launch-failures: cover the debugpy stderr -> 'pip install debugpy'
rewrite for launch and attach, plus a negative case (non-debugpy adapter
with the substring in stderr is left untouched).
- dap-launch-failures: model the delayed-launch-failure case the new
settled-race in throwPreferredDapStartError defends against. FakeDapClient
gains optional launchErrorDelayMs/attachErrorDelayMs.
- dap-launch-failures (DebugTool): assert adapter:'debugpy' early-throw
surfaces 'python not found in PATH' on both launch and attach when
selectLaunchAdapter/selectAttachAdapter return null, and the unspecified-adapter
path still falls back to the generic 'No debugger adapter' error.
- find.ts: export validateFindPathInputs and pin the new backslash-escape
semantics (\, no longer trips the comma-joined heuristic) plus the
existing brace-expansion and rejection paths.
- patch.ts: cover the post-write verification error message. The user-facing
ToolError must contain the caller-supplied relative path and not the
absolute resolvedPath (which still lives in the structured context for
log correlation).
- split-internal-url-sel: reword two mcp:// test comments that described a
'peeler refuses' guard that doesn't exist; rename the tests to reflect the
actual opaque-scheme rule.
- browser tool's existing-tab re-nav defaults to waitUntil: 'load' (matching
new-tab path); identical acquireTab() calls no longer hang on dev servers
- patch tool error path uses caller-supplied relative path; absolute
resolvedPath stays in structured context only ($HOME no longer leaks to TUI)
- splitInternalUrlSel keeps mcp:// resource URIs opaque even when they end in
':raw' or '/:1-50' (McpProtocolHandler matches by verbatim URI)
- find tool: timeout signal honored by onMatch; partial results sorted by
mtime desc; backslash-escaped commas skipped in path-list validation
- DAP throwPreferredDapStartError waits up to 50ms for the underlying
launch/attach error instead of one microtask
- debug tool surfaces 'python missing' and 'pip install debugpy' diagnostics
separately when adapter: 'debugpy' is requested
- Updated hashline markers to `¶` headers and `^/v/->` operators across constants, prompts, docs, and tests.
- Reworked hashline grammar and parser to support `ANCHOR<SIGIL>[INLINE_PAYLOAD]` with optional inline bodies and `A-B` ranges.
- Changed range and marker syntax from `..`/`"` to `-` and suffix `^/v/->` forms like `7v` and `A->`.
- Aligned `sameLineRange()` output and BOF/EOF handling so `|TEXT` remains cosmetic and payload now follows the op line.
- Centralized op-line detection by replacing local regex helpers with `isHashlineOpLineText` for payload terminator and bad-op checks.
- Deleted the `ValidationVerdict` type and `evaluateOutputAgainstSchema` API from the output schema validator.
- Updated `yield.ts` to bind `buildOutputValidator`'s error directly to `schemaError` during validator setup.
- Removed the obsolete evaluator tests and adjusted validation success fixture to match the raw summary input shape.
- Unified output schema construction and validation by adding buildOutputValidator and using it in YieldTool and task executor.
- Added MAX_SCHEMA_RETRIES so YieldTool now retries schema failures three times with hints before overriding.
- Updated failure handling to use shared summarizeValidationFailure and formatters for required-field reporting.
- Added tests for output-schema-validator and YieldTool covering malformed schemas and nested-array retry edge cases.
- Centralized OAuth access lifecycle in `AuthStorage`, returning identity metadata and new access-result types.
- Added 60-second skew and strict expiry checks, returning undefined/throws for stale or expired OAuth credentials.
- Removed provider-local token refresh flows from Gemini, Gemini CLI, Antigravity, Kimi, and related OAuth helpers.
- Migrated web-search providers from `AgentStorage` to `AuthStorage` session-aware lookup with `authStorage`/`sessionId`/`signal` flow.
- Replaced `findAnthropicAuth`/DB auth lookup with `buildAnthropicAuthConfig` and explicit base-url override/env fallback ordering.
- Added OpenAI Codex and Gemini web search provider options with updated setup/auth descriptions.
- Updated Codex OAuth flow to refresh near-expiry tokens during web_search and persist the refreshed credentials.
- Plumbed AgentStorage through search orchestrator, scrapers, and fetch paths so providers share session credentials.
- Refactored web provider and credential helpers to accept caller-provided AgentStorage and resolve keys synchronously.
The report_finding tool's priority is exposed as a string enum
("P0"-"P3") for ergonomics, but the reviewer agent and every
custom review agent declare priority as `type: number` in their
JTD output schema. The cast at executor.ts:1473 lied about the
runtime shape, so the auto-injected `findings[].priority` flowed
through as strings and every yield with at least one finding was
rejected with `findings.0.priority: expected number, received string`,
forcing the run into the schema_violation exit path.
Added `toReviewFinding(details)` in tools/review.ts that maps the
priority enum to its numeric ordinal via the existing PRIORITY_INFO
table and use it at the boundary in executor.ts. Render paths still
see the original `ReportFindingDetails` shape (string priority)
through normalizeReportFindings, so display formatting is unaffected.
Fixes#1350
Loaded marketplace lspServers metadata from Claude plugin caches and embedded it for OMP marketplace installs so config-only plugins register without package code.
Fixes#1352
The JTD-to-JSON-Schema converter post-processed convertSchema's
output with normalizeMixedSchemaNode, which walked back into the
emitted JSON Schema looking for nested JTD forms. Inside a
properties block, user-defined property names whose keys happened
to collide with JTD keywords ('ref', 'elements', 'values',
'optionalProperties', 'discriminator') were misclassified as JTD
forms and re-rewritten - corrupting properties like { ref: { type:
'string' } } into { $ref: '#/$defs/[object Object]' } and breaking
the built-in explore agent's output validator with
schema_violation: files.0.ref: must not be present.
convertSchema is already fully recursive and emits pure JSON Schema,
so the post-walk is both unnecessary and unsafe. Drop it.
Fixes#1345
- Changed hashline format to canonical `§` section headers and `"`/`"`/`≔` operations across grammar, parser, and docs.
- Reworked range and op parsing so single anchors are valid, legacy `-`/`-=` ops error, and empty `≔` payloads now delete ranges.
- Removed legacy `HL_EDIT_SEP`, `$HSEP$`, and `hsep` plumbing, adopting raw payload lines.
- Updated execution, input, renderer, and streaming flows to use `HL_FILE_PREFIX` and `HL_OP_CHARS` helpers.
- Updated session-stats parsing to `§`/`≔`/`"`/`"` format, patch envelopes, and bumped parser versions.
- Removed the hashline-separator benchmark script and its PI_HL_SEP job orchestration.
Appended the stealth iframe to documentElement when document.head is not available during new-document evaluation.
Added regression coverage for the null-head bootstrap path.
Fixes#1267
Adds a shared classifyProviderHttpError helper that maps the well-known
failure shapes (status 401/402/403 and bodies matching credits / quota
/ insufficient) into compact SearchProviderError messages so the
orchestrator advances to the next provider instead of bailing.
Wired into every HTTP-talking provider (codex, exa, gemini, anthropic,
brave, jina, kimi, perplexity, searxng, synthetic, tavily, zai) and the
two wrapped-error providers (kagi via KagiApiError, parallel via
ParallelApiError). The orchestrator now collects per-provider failures
and emits a joined summary ("exa: 403 forbidden; codex: credits
exhausted; ...") when the whole chain fails.
Test updates: web-search-{exa,tavily,kagi} message expectations now
assert the compact "<id>: <status> <reason>" form.
- Added splitInternalUrlSel to iteratively peel internal-URL selector chunks while preserving unsupported schemes like mcp://.
- Updated ReadTool to use the internal splitter before routing so selector parsing is handled via parseSel.
- Added unit tests covering malformed selectors, namespaced skill hosts, and unchanged behavior for non-URLs or unsupported schemes.
- Added capParseErrors in shared render utilities and updated ast_grep and ast_edit to return capped parseErrors plus parseErrorsTotal.
- Threaded the preserved totals into parse-error formatting and renderer output so labels and overflow counts report the full number of issues.
- Added an ast_grep test asserting parse errors are capped at PARSE_ERRORS_LIMIT while parseErrorsTotal retains the original count.
- Set GIT_CONFIG_* and GIT_TERMINAL_PROMPT environment variables to prevent test interference from user gitconfig, LFS filters, signing, and credential helpers.
- Enhanced git error messages to include stdout when stderr is empty, providing better diagnostics for test failures.
- Replaced symlink-resolving pathIsWithin with lexical path containment check to prevent test isolation bypass via symlinked extensions.
- Added `providerRetryWait` and `retryWait` hooks to stream/usage options so tests bypass real scheduler delays.
- Parameterized GitHub Copilot poll intervals and Copilot model retry base delay for fast test execution.
- Replaced `Bun.sleep`/`setTimeout` polling loops with `AbortSignal` event listeners in agent session tests.
- Consolidated auth-gateway E2E helpers into a shared `test/helpers` module, eliminating duplicated `checkGatewayAvailable` implementations.
- Migrated credential-disabled tests from SQLite-backed stores to an in-memory store, removing temp-dir lifecycle overhead.
- Added `dev.autoqa.consent` setting and single-flight popup handler wired through `InteractiveMode`.
- Added `flushGrievances` to batch-POST unpushed rows to `dev.autoqaPush.endpoint` with cooldown and single-flight deduplication.
- Added `omp grievances push` subcommand with TTY progress bar for manual draining.
- Migrated shared DB logic to `openAutoQaDb` (with `pushed` column migration) exported from `report-tool-issue`.
- Extracted `description` from type-array and nullable branches so it lives on the wrapper, not duplicated onto each variant.
- Replaced inline enum-type inference with `inferStrictPrimitiveTypeFromEnumOrConst`, covering both `enum` and `const` in sanitize and enforce paths.
- Mixed-primitive enums and non-primitive consts now fall back to non-strict instead of producing a typeless schema that OpenAI rejects on the wire.