- Switched extensibility loader imports from namespace-style `zod` imports to named `z` imports in `zod/v4`.
- Updated extensibility type interfaces to use `typeof z` for injected `zod` modules in hook, extension, tool, and command APIs.
ACP_BUILTIN_SLASH_COMMANDS only carries primary names; the reserved set
passed to getRegisteredCommands was therefore missing aliases like
"models" (/model) and "force:" (/force). An extension registering one
of these aliases would appear in the palette but the builtin would win
at dispatch time (lookupBuiltinSlashCommand searches aliases too).
Export ACP_BUILTIN_RESERVED_NAMES from acp-builtins — the union of all
primary names and aliases for ACP-surfaced builtins — and use it as the
reserved set. Widen getRegisteredCommands parameter to ReadonlySet since
it only calls .has().
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
Same shape of bug the reviewer flagged for custom tools: forwarding
`LoadExtensionsResult` from parent to subagent reused Extension instances
whose factories closed over the parent's `ExtensionAPI` — cwd, eventBus,
and runtime all pointed at the parent. Any tool/handler/command that
referenced `api.exec()`, `api.events`, or `api.runtime` still acted on the
parent session/worktree from inside an isolated subagent.
Forward only the path list; each session rebuilds extensions through
`loadExtensions` so factories see the right `ExtensionAPI`.
- `extensibility/extensions/loader.ts`: extract `discoverExtensionPaths`
(FS scan only) from `discoverAndLoadExtensions`. The combined helper now
composes the two. New export added to the package barrel.
- `sdk.ts`:
- Add `discoverSessionExtensionPaths()` (the `disableExtensionDiscovery`-aware
path-only counterpart of `loadSessionExtensions`).
- Add `preloadedExtensionPaths?: string[]` to `CreateAgentSessionOptions`.
Three loader branches: `preloadedExtensions` (CLI same-process reuse,
still shallow-cloned), `preloadedExtensionPaths` (subagent: skip scan,
reload locally), or full discovery.
- Document `preloadedExtensions` as same-process-only; subagent
forwarding MUST use `preloadedExtensionPaths`.
- `tools/index.ts`: `ToolSession.extensionsResult` → `extensionPaths:
string[]` for the same reason.
- `task/executor.ts` and `task/index.ts`: forward `extensionPaths`. Drop
the forward for the isolated `runSubprocess` branch — worktree cwd ≠
parent cwd, so the subagent re-discovers extensions against its own
tree.
- New `test/sdk-extensions-per-session-binding.test.ts` pins the contract:
two `loadExtensions` calls on the same path with different `cwd` and
different `EventBus` instances yield distinct Extension + runtime
objects whose factories close over the per-call bindings.
- Updated `executor-pass-through` and `sdk-preloaded-extensions-isolation`
tests for the new option name and comment context.
Refs PR review on #2193
- Derived descriptors, default-model map, env keys, login list, and refresh dispatch from one ProviderDefinition per provider.
- Disabled OpenAI Codex stream obfuscation and interrupted whitespace-only tool-call argument deltas.
- Derived auth-broker callback ports and paste-code login set from the registry.
MCPTool and DeferredMCPTool now declare approval = 'write' instead of
implicitly defaulting to 'exec'. Without this, the approval system
requires user confirmation for every MCP tool call in non-yolo modes,
but the confirmation prompt never renders in the TUI while streaming,
causing the agent to hang indefinitely.
Also propagate the approval property through customToolToDefinition()
in sdk.ts, which was silently dropping it during CustomTool ->
ToolDefinition conversion.
- Added `TUI.resetDisplay()` to force an immediate full-frame replay including native scrollback.
- Moved the persistent model selector default from Ctrl+L to Alt+M, preserving existing user remaps.
- Reserved Alt+M so extensions cannot shadow the model selector shortcut.
ExtensionRunner#getShortcuts() accepted ctrl+q because #RESERVED_SHORTCUTS
predated the new default, and InputController registers extension
shortcuts before the followUp keybinding, so the editor's custom-key
map silently overwrote the extension handler. Now ctrl+q is reserved
alongside the other built-in chords and the extension authoring docs
list it as such.
Addresses code review on #1905.
- Added `selectionMarker`, `checkedIndices`, and `markableCount` options to render radio/checkbox glyphs per row.
- Moved checkbox rendering from inline label prefixes into the selector component.
- Kept trailing control rows like "Other"/"Done" on the plain cursor.
Two review fixes for the extension-flag/initial-prompt work:
1. @file ordering — `processFileArguments` runs `process.exit(1)` on a
missing/unreadable file. It had been moved after `createSession`, which
writes the terminal breadcrumb eagerly (SessionManager.create →
#newSessionSync), so `omp @missing.md "x"` left a junk session/breadcrumb
behind before exiting.
Resolve extension-registered CLI flags BEFORE creating the session: load the
session's extensions up front (new `loadSessionExtensions` helper, the single
source of createAgentSession's discovery-branch logic), build an
ExtensionFlagSink straight from the loaded extensions + runtime, re-parse
argv, then process @file args — all before any session exists. The loaded
result is handed back to createAgentSession via `preloadedExtensions` (now
checked before `disableExtensionDiscovery`, so it can't double-load) and the
same EventBus is shared, so no extra work. This keeps the P1#1 fix
(`--flag @value` is the flag's value, not a file) while failing fast with no
session side effects.
2. "Can we avoid the big list of names?" — removed the hand-maintained
`BUILTIN_FLAG_NAMES` set (and its stale "rejected at registration" doc).
`applyExtensionFlags` now always falls back to recovering a flag's value from
argv when parseArgs didn't surface it; the recovery scan mirrors parseArgs's
consumption rules (flag-looking space-form values stay their own flag) and is
a no-op for flags that were absent or already surfaced, so no list of
built-in names is needed.
Adds `ExtensionRunner.aggregateFlags` (static) so getFlags and the CLI's
pre-session sink share one implementation.
Tests: pre-session flag resolution via the exact main.ts sink pattern;
list-free recovery of an arbitrary colliding built-in (`--model`); and the
flag-looking-value rule. Verified typecheck + extension/runner/acp suites.
The previous collision guard threw in registerFlag, which broke loading the
bundled plan-mode example extension (it registers `--plan`, also a built-in)
even when `--plan` was never passed — making a documented extension unusable.
Registering a built-in-named flag is a supported pattern: `--plan` is both the
built-in plan-model selector and plan-mode's boolean mode toggle, and the value
must reach both. So instead of rejecting, preserve delivery: remove the guard,
and in applyExtensionFlags recover a colliding flag's value from argv
(resolveCollidingFlag) when parseArgs routed it to the built-in branch and it
never reached unknownFlags. Non-colliding flags are unchanged (peer-* etc.).
Verified the real bundled plan-mode.ts loads with --plan registered and
delivered; replaced the reject-test with a loads-without-throwing regression
plus colliding-flag delivery coverage.
Three issues from an adversarial review, all rooted in the startup argv parse
running before extensions load:
1. Flag-looking string values (`--name --print`): the extension-aware reparse
consumed the following token as the value, disagreeing with the startup
parse that treated `--print` as the built-in flag — so the reparse could
silently flip command shape. Extension string flags now consume a following
token only in `--flag=value` form or when it is not flag-looking; pass a
flag-looking value as `--flag=value`. Keeps both parses consistent.
2. `@file` string values (`--target @notes.md`): file args were processed from
the startup parse, which misreads the value as a file and reads it into the
prompt. processFileArguments now runs on the extension-aware parse
(initialArgs.fileArgs); pipedInput stays early for mode detection.
3. Built-in collisions: an extension flag named like a built-in (e.g. `model`)
was consumed by the built-in branch and never delivered to the runner.
registerFlag now rejects names in BUILTIN_FLAG_NAMES with a clear error
(isolated per-extension by loadExtension's try/catch).
Adds tests for all three plus the documented startup-parse misclassification.
- In `resolveApproval`, yolo mode now returns the user policy directly (`allow`/`prompt`/`deny`) and ignores tool `override` prompts.
- Updated approval-mode and approval unit tests to match the new behavior for critical bash patterns under yolo and auto-approve.
- Updated docs and settings metadata to describe yolo as user-policy-driven rather than override-driven.
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
Stale comment said 'layers user config on top' which sounded like custom
mode merged defaults with config. The actual behaviour (and what the user
asked for): in custom mode user config wins; built-in defaults only fill
gaps for tools the user hasn't configured. auto and prompt modes ignore
tools.approval entirely.
New global setting under /settings -> Interaction that controls the tool
approval flow:
auto (default) Skip every approval prompt — yolo. Matches --auto-approve.
prompt Built-in per-tool defaults only. Destructive tools (bash,
edit, write, eval, ssh) require confirmation; read-only
tools auto-allow; tools.approval.<tool> overrides ignored.
custom tools.approval.<tool> config wins. Built-in defaults only
fall back for tools the user hasn't configured. Critical
safety patterns (rm -rf /, fork bombs, curl|bash) still
prompt even when the tool is user-allowed.
The CLI --auto-approve / --yolo flag always wins regardless of the setting,
preserving the automation/CI path.
Wires through ExtensionToolWrapper.execute(): the wrapper reads
tools.approvalMode from settings, derives userPolicies only for custom mode,
and feeds the existing requiresApproval() resolver. Resolution order inside
requiresApproval already places user config above built-in defaults, so
'config wins' falls out naturally in custom mode.
Adds test/tools/approval-mode.test.ts covering all three modes, the CLI
override, the built-in fallback in custom mode, and the critical-pattern
override that fires even when bash is user-allowed.
Re-introduces the per-tool approval system from luzidd's commit 39124f3 (which
is no longer reachable from main) and improves it before re-landing.
What's restored:
- ApprovalPolicy (allow/deny/prompt) plus DEFAULT_APPROVAL_POLICIES.
- ACTION_EXCEPTIONS registry (LSP read-only, bash critical patterns).
- getApprovalPolicy() six-level resolution order.
- ExtensionToolWrapper.execute() gate before extension handlers.
- --auto-approve / --yolo CLI flag and tools.approval.<tool> user config.
- docs/approval-mode.md user guide.
What's improved over the original:
- Replaced unchecked 'as any' casts with typed unknown narrowing helpers.
- Validate userConfig values: invalid strings, numbers, etc. fall through to
the built-in default instead of being silently honoured (typo no longer
locks a tool out or grants implicit approval).
- Expanded CRITICAL_BASH_PATTERNS: chmod -R /, chown -R /, bash <(curl ...),
writes to /etc/passwd|shadow|sudoers, shutdown/reboot/halt/init 0,
kill -9 1, nc -e / nc -c reverse shells. Pattern shapes require a
command-position boundary so 'npm run reboot-tests' and 'echo "shutdown the
queue"' don't false-positive.
- Added DEBUG_READONLY_ACTIONS exception so DAP inspection actions (threads,
stack_trace, variables, scopes, read_memory, …) auto-allow while
execution-side actions (launch, attach, continue, evaluate, write_memory,
set_breakpoint, …) still prompt.
- formatApprovalPrompt: labels mcp__<server>__<tool> calls as MCP server
tools, surfaces ssh host + command, recognises the modern § hashline header
for edit, and truncates >240-char fields so a heredoc-sized body cannot
blow out the confirmation dialog.
- Test suite grown from 40 to 57 cases — new coverage for invalid user
config, the extended critical-bash patterns, benign-keyword negatives,
debug exceptions, MCP/ssh prompt formatting, and command truncation.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts -> 57 pass
- bun x biome check . -> clean
- bun run check:ts across all 9 workspaces -> clean
- Removed deprecated MCP-specific type aliases and functions from tool-discovery module, consolidating to unified generic tool discovery API.
- Migrated session and SDK code to use generic filterBySource() and collectDiscoverableTools() instead of MCP-specific variants.
- Removed deprecated interface members including hasQueuedMessages(), FocusPane, AcpBuiltinCommandRuntime, and legacy settings methods.
- Updated test suites to use renamed generic discovery methods and removed back-compat test coverage for legacy MCP shapes.
Queued extension-delivered user messages when deliverAs is set and waited for session_start extension message sends before prompting subagents.
Fixes#1343
- Relocated compaction, branch-summarization, pruning, and utils from coding-agent to packages/agent/src/compaction.
- Moved OpenAI remote compaction helpers from packages/ai to the new compaction module.
- Added handoff.ts with extractHandoffDocument, createHandoffContext, and renderHandoffPrompt helpers.
- Exposed new entries.ts with standalone SessionEntry types so coding-agent no longer owns them.
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
- Added a shared session command helper that aggregates extension, prompt, and skill slash commands.
- Updated ACP, extension UI, runtime-init, and task executor extension contexts to return session command data instead of empty arrays.
- Added GoalRuntime with wall-clock and token accounting, budget steering, and lifecycle operations (create, pause, resume, drop, complete).
- Exposed goal tool as a hidden agent tool, activated only when goal mode is enabled.
- Integrated goal continuation loop in InteractiveMode with auto-submit between turns.
- Added status line segment and theme icons for goal mode state.
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
- Deferred initialize to flush queued credential_disabled events via queueMicrotask and event splicing.
- Added tests for pre-initialize credential_disabled emissions and onError propagation of handler failures.
- Replaced auth credential disable flow with CAS checks in #tryDisableAuthCredentialIfMatches and matching SQL statement.
- Retried OAuth getApiKey after disable failures; added peer-rotation race test for fresh token and active credential retention.
- Updated CHANGELOG for deferred microtask flushing plus eval import renames and diff URL/quoted-path parsing fixes.
Adds `pi.on("credential_disabled", handler)` so extensions can react to
soft-disabled credentials (e.g. OAuth invalid_grant) without regex-matching
`agent_end` errorMessages.
`AuthStorage.onCredentialDisabled(listener)` returns an unsubscribe function;
multiple listeners fire for every event with per-listener exception isolation
and FIFO buffer-and-replay (cap 32) when none are attached. The constructor
option from #991 stays as sugar for an immediate permanent subscription.
`createAgentSession()` subscribes the per-session extension runner to
`modelRegistry.authStorage` immediately after resolution and unsubscribes on
dispose / startup failure. Events are forwarded via
`ExtensionRunner.emitCredentialDisabled(event)`, which buffers (cap 32,
drop-oldest) until `runner.initialize(...)` runs in the mode controller so
extension handlers see real UI/runtime context, not the constructor no-op
defaults.
Supersedes #997. Builds on #991.
Co-Authored-By: omp <noreply@oh-my-pi.dev>
- Added a 30s timeout for extension handlers in runner.ts so stalled callbacks now emit warnings and stop.
- Consolidated duplicated extension handler error handling by routing calls through #runHandlerWithTimeout.
- Updated ExtensionUIContext, InteractiveModeContext, and InteractiveMode to require editor factories to return CustomEditor instances.
- Removed the runtime compatibility guard and warning for non-CustomEditor implementations in setEditorComponent.
- Removed the test that verified rejection of non-CustomEditor factories in interactive-mode editor-component tests.
Some extensions/plugins still import the legacy @mariozechner/pi-* package names (pi-agent-core, pi-ai, pi-coding-agent, pi-tui) instead of @oh-my-pi/pi-*. Register a single Bun.plugin on first plugin/extension load that rewrites those specifiers to the current @oh-my-pi/pi-* equivalents, including the @mariozechner/pi-coding-agent/extensibility/{extensions,hooks} sub-exports. No filesystem mutation, no symlinks, no proxy files.
Fixes#973
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
- Documented and removed `utils/oauth` from the `ai` package entrypoint, noting it as a breaking change.
- Refactored `cli`, `auth-storage`, and `utils/oauth` to load provider modules via scoped dynamic `import()` calls.
- Removed top-level provider imports and barrel exports from `utils/oauth/index.ts`, streamlining oauth module loading.
- Consolidated OAuth symbol, type, and provider imports in coding-agent and tests to `@oh-my-pi/pi-ai/utils/oauth` modules.
- Defined `DEFAULT_LOCAL_TOKEN` locally in model-registry and removed its cross-package OAuth import usage.
- Required top-level `path` in edit requests, removed per-entry `path` fields, and updated docs/tests to match.
- Updated patch/hashline/atom streaming preview generation to return a single request-level path diff preview.
- Changed edit execution to route patch/replace/atom/hashline through single-path handlers sharing `path`.
- Added `scripts/analyze-edit-formats` Go CLI with reports to audit edit-tool usage from session JSONL logs.
- Renamed the built-in `grep` content-search tool to `search` across settings, schemas, and SDK exports.
- Switched execution wiring so `Task`, `Plan`, cursor, and shell mapping now invoke `search` instead of `grep`.
- Updated prompts, plan-mode docs, and example tool lists to replace `grep`/`ls` references with `search` guidance.
- Aligned `Grep*`/`grep` event, renderer, and hook types to `Search*`/`search` across runtime and tests.
- Documented and fixed `search` result rendering budget behavior and added internal-URL/path-list transcript notes.
- Canonicalized file and CLI defaults from `read` to `open` across tool registration and prompts.
- Added `resolveToolAlias()` and applied alias-normalized tool selection so legacy `read` maps to `open`.
- Updated runtime, UI, and export layers to treat `open` as first-class while preserving `read` compatibility.
- Renamed read prompt docs to `open.md`/`open-chunk.md` and refreshed system guidance to recommend `open`.
- Updated tool-related tests and expectations from `read` to `open` (including test fixtures and aliases).
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
- Removed `SearchDb` APIs and `searchDb` fields, dropping db-backed state from native and agent sessions.
- Replaced crate export `fff` with `fd`, moving fuzzy-find bindings into `fd.rs`.
- Removed `SearchDb`/picker fast-path logic from `glob` and `grep`, simplifying scan flow and dropping db args.
- Removed `SearchDb`/`getSearchDb` wiring from extension, tool, and task context constructors across coding-agent.
- Added over-indentation validation warnings in chunk-edit normalization for suspicious `~` body line formatting.
- Removed `bytes`, `fff-grep`, `fff-search`, and `blake3` deps, adding `grep-searcher = "0.1"`.