- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
- Updated sdk-tool-activation expectations for 386385f18b: sessions
without a granted write tool keep extension/SDK tools top-level and
allocate no xd:// state instead of auto-granting write.
- Raised the unref'd-worker parent-exit repro to a 30s timeout; the 10s
ceiling SIGTERMed the wrapper (exit 143) on shared-core CI runners.
The bridge is constructed once, at session creation, and was handed the
startup `cwd` by value. The session's own cwd moves under it — `/cd`,
resume, branch restore all call `sessionManager.moveTo` — and the two
frames that confine a path themselves (the native `delete`, and a
`read_mcp_resource` carrying `download_path`) resolve against whichever cwd
the bridge holds. So after a move the primary deleted or overwrote the
relative path in the workspace the session had left, and reported success
for the path the server actually named.
The advisor bridge already passed a live resolver; this is the same
resolver on the path that was missed. Locked by a wiring test: the seam is
the session handing its handlers to the provider, so the test captures them
there, moves the session, and asserts the frame acts on the new workspace
and leaves the old file alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 079c7ac61104d017eecbf781aa1c58eebd39b0b1)
`runs advisor tools through the approval gate` built its advisor from the
`advisor` role chain, which resolves against `modelRegistry.getAvailable()`
— the models the host holds auth for. On a developer box whose environment
carries provider keys the roster resolved and the test passed; in CI, where
the suite's isolated auth storage is empty, every advisor resolved to
`no_model` and `getAdvisorAgent()` returned undefined ("expected an advisor
agent").
The advisor now names `gpt-4o-mini` outright and runs inside the file's
`withProviderAuth` helper, so the roster resolves from the granted key
rather than from whatever the machine happens to have configured.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit ef2054da5dbe3d3d1cca4025e12e5c350f47173d)
Advisor tools are built straight from the builtin table, outside the
loop that wraps every registry tool in `ExtensionToolWrapper` - which is
where the approval mode, per-tool `tools.approval.<tool>` policies and
`autoApprove` are enforced. Both the advisor's own agent loop and its
Cursor exec bridge (`pi_write`, `pi_bash`) run those instances directly,
so an advisor granted `write` or `bash` executed them regardless of a
configured `ask` or `deny`. Verified before the fix: a raw `write`
instance created the file under `tools.approval.write: deny`; the
wrapped one refuses. `bridgeToolMap` and the grep factory only ever
wrapped the two tools they build themselves.
The grep answer also now echoes `offset_applied`. Forwarding the offset
without acknowledging it leaves the server unable to distinguish a
honored page from a client that ignored the field, so it re-paginates
from the same place. Set on all three result variants (files, count,
content); absent when the frame requested no offset.
(cherry picked from commit 13c2ed565f30ff86e31c8dc9896f940faf3aa088)
`ReadMcpResourceExecResult` has a `rejected` variant carrying `reason`,
not `error`, so the pairing text's collapsed error branch did not
typecheck against the full union. Each variant is now switched
explicitly; the refusal is unreachable today (the handler answers
content or `null`) but a collapsed default would have read `undefined`
if the client ever builds one.
`bun check` type-checks the workspace projects; the union error only
surfaced through `ci:check:full`, which is what CI runs.
Also renames a loop variable that shadowed the global `escape`, which
was failing `biome check` on the full tree.
(cherry picked from commit f5dca418d91983a14492f51f2873514781a9e066)
Every native `pi_edit` failed after a session switched onto Cursor. The
replace-mode `edit` instance the frame needs was built only for sessions
CREATED on Cursor, and the tool roster is built once, at creation - a
session that started elsewhere kept its configured-mode `edit` in the
registry, which `executeTool` resolves before its fallback, so the
frame's `old_text`/`new_text` pairs failed validation against a
`hashline` schema.
The instance is now built from the `edit` grant regardless of the
initial provider, lazily so a session that never reaches Cursor never
constructs one, and `pi_edit` asks for it through a dedicated
`getEditReplaceTool` accessor rather than relying on Cursor sessions
having deleted `edit` from the registry. A session that was never
granted `edit` is still refused.
That accessor also closes an escalation the previous wiring opened up.
The session's device resolver is handed to the bridge as `getTool` and
installed as the agent loop's `resolveFallbackTool`, which runs for ANY
call outside the advertised set - so serving `edit` from it let a
hallucinated call, or one naming a tool the session deselected after
startup, execute a replace-mode edit the model was never offered. It is
device-only again.
Regressions cover both directions at the SDK level, driving a real
unadvertised `edit` through the loop and asserting the surfaced
`Tool edit not found`: an unchanged file alone would also pass if the
fallback had resolved the tool and the edit then failed validation.
(cherry picked from commit 11a28dcf7b995a9e94913269733b3199d6f4790d)
- Replaced getTool with getExecutableTool in CursorExecBridgeOptions to prioritize mounted-device permission wrappers over canonical tools.
- Updated createAgentSession to check isAutoQaEnabled against restricted tool filtering when configuring system prompts.
- Added test coverage verifying execution overrides preserve approval gates and restricted sessions omit auto-qa guidance.
- Added a suppress load option so disabled servers still claim their
capability key: a project foo with enabled:false shadows a same-named
enabled user foo again, while scope-removed entries drop fully.
- Tool-name collisions now resolve by stable server+tool origin key
instead of manager array order, so reconnect re-appends cannot flip
the routed implementation.
- Review follow-up for PR #6787.
Moved first-wins MCP tool-name deduplication and origin-aware warnings into one shared helper used by startup extension registration, SDK custom-tool assembly, and deferred refreshes.
Added an SDK startup regression proving colliding MCP proxy tools keep the first origin instead of silently overwriting it.
Fixes#6786
- Replaced single-provider preferences with ordered priority lists for web search and image generation.
- Added a `MultiSelectSubmenu` component supporting toggle and reordering interactions in settings.
- Implemented migration logic to convert legacy single-provider preferences into ordered priority lists.
- Updated setup wizard scenes, image generation fallback logic, and search provider chains to use priority lists.
Caller-provided output schemas are free-form JSON and cannot be represented by OpenAI strict tool schemas. Keep todo strict while explicitly sending task as non-strict.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
Resolved plan-mode exit overlap with #5662 (kept restore/rollback
structure, routed pending-switch clearing through
clearPendingPlanModelSwitch) and unioned additive test blocks with
#5672/#5662.
Forwarded the session xd registry into Cursor provider tool contexts.
Routed Cursor MCP execution through the mounted registry fallback and added regression coverage for built-in devices and external MCP tools.
Fixes#5650
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Transitioned vibe tools to an ephemeral registration model where they are installed only when entering `/vibe` mode and removed upon exit.
- Added `activateVibeTools` and `deactivateVibeTools` methods to `AgentSession` to manage these transient tool registrations.
- Removed vibe tools from the default global tool registry, preventing unnecessary background exposure.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
createAgentSession() removed the hidden `resolve` tool from the registry
whenever no active tool advertised `deferrable: true`. Plan mode dispatches
its plan-approval `resolve { action: "apply", extra: { title } }` call
through a standing handler installed by InteractiveMode (no deferrable tool
involved), so read-only plan-mode toolsets (e.g. `read`, `search`, `find`,
`web_search`) silently activated plan mode without `resolve`. The agent had
no callable tool to submit the finalized plan and got stuck on the post-turn
tool-decision reminder.
Keep `resolve` registered whenever `plan.enabled` is true so the standing
handler always has a callable tool. The hidden flag still prevents `resolve`
from appearing in the active tool set until plan mode (or a deferrable tool's
preview action) opts in.
Fixes#1428
Plan-mode subagents (and any subagent with an explicit `agent.tools` array)
were given the `yield` tool in the registry but not in
`agent.state.tools`. The session prompts and idle reminders still
demanded a `yield` call to terminate, so the model would reason
"there doesn't seem to be a yield tool available" and the turn went
nowhere.
`createTools` correctly appends `yield` to the registry when
`requireYieldTool: true`, but `createAgentSession` then derived the
active tool list from `options.toolNames` directly, dropping `yield`
again. Mirror the invariant already enforced in
`parseAgentFields` (discovery/helpers.ts): when `requireYieldTool` is
set and the caller passes an explicit list, append `yield` to it
before normalization.
Fixes#1408
- Removed vim edit mode and automatically map existing vim configurations to hashline mode.
- Deleted VimTool class, VimEngine implementation, and all vim-specific editing logic (2409 lines).
- Removed vim mode from EditMode union type, edit tool strategies, and configuration schemas.
- Deleted vim parser, command handler, buffer manager, and renderer modules.
- Updated documentation and tests to remove vim mode references and add deprecation mapping.
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
- Canonicalized file and CLI defaults from `read` to `open` across tool registration and prompts.
- Added `resolveToolAlias()` and applied alias-normalized tool selection so legacy `read` maps to `open`.
- Updated runtime, UI, and export layers to treat `open` as first-class while preserving `read` compatibility.
- Renamed read prompt docs to `open.md`/`open-chunk.md` and refreshed system guidance to recommend `open`.
- Updated tool-related tests and expectations from `read` to `open` (including test fixtures and aliases).
- Removed the standalone vim tool and normalized built-in/requested tooling to edit.
- Updated session and SDK tool activation to dedupe lowercase names and track edit state via the edit key.
- Added vim-mode argument detection and delegated edit rendering/execution into Vim handlers under edit.
- Updated Vim step handling to auto-reorder numeric-positioned commands, including cc/C/S/s/i/I/A cases.
- Renamed prompt/changelog text and test expectations to reflect edit-only tool naming and usage.
- Updated `AgentSession` and `createAgentSession` to resolve active edit tool names via `resolveEditToolName`.
- Filtered inactive edit variants with `filterInactiveEditToolName` so tool listings expose the active edit mode.
- Normalized requested edit-capable tool names with `normalizeToolNamesForEditMode` during activation and startup.
- Synced edit-tool mode after model changes so the active tool swaps between `edit` and `vim` automatically.
- Enabled per-file partial updates by threading `onUpdate` through `EditTool` and `executePerFile`.
- Updated `ToolExecutionComponent` to render multi-file edit `perFileResult` entries as separate boxes with pending state.
- Added `defaultInactive` property to ToolDefinition for conditional tool registration and activation control.
- Added dynamic tool activation/deactivation API `setActiveTools()` for managing experiment tools in autoresearch mode.
- Replaced single `command-start.md` workflow with separate `command-initialize.md` and `command-resume.md` prompts for autoresearch initialization and session resumption.
- Added interactive intent dialog for autoresearch optimization goals with automatic session resumption detection based on autoresearch.md presence.
- Refactored autoresearch command handler to distinguish resume vs initialize flows and dynamically activate/deactivate experiment tools based on mode and session state.