- Assert the eval tool hides disabled backends from the model-facing wire
schema (language enum + field descriptions), summary, and description by
default (rb/jl off), and advertises them once enabled — including the
enabled-subset case.
- Update the env-flag fallback test for the new rb/jl opt-in defaults and
extend the env guard to PI_RB/PI_JL so the suite is shell-independent.
- Note the opt-in default and dynamic advertising in the changelog.
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
- Removed the `readHashLines` setting to consolidate hashline display logic.
- Simplified `resolveFileDisplayMode` to derive hashline visibility solely from the active edit mode.
- Added automatic cleanup of the `readHashLines` key from existing configuration files.
- Implemented the Devin inference provider, including OAuth flow with PKCE, Connect protocol integration, and streaming support for chat requests.
- Integrated comprehensive Protobuf-based service definitions and generated TypeScript clients for Devin's API infrastructure, including model management and workspace operations.
- Updated the AI and Catalog modules to support dynamic model discovery, provider-specific configuration, and authentication.
- Standardized tool call arguments as `Record<string, unknown>` across provider implementations to ensure type safety.
- Implemented `waitForSelector` and `waitForNavigation` methods in the browser tool API.
- Introduced per-operation fail-fast budget management with dynamic timeout clamping.
- Added validation for selector engines to reject unsupported Playwright-only platform features.
- Standardized error handling to provide descriptive, named timeouts for stalled browser operations.
- Implemented `tab.ariaSnapshot()` to capture and represent page structures as ARIA-tree YAML.
- Introduced `tab.ref()` and ref-based selector parsing to enable precise element interaction via unique ARIA identifiers.
- Integrated automated script bundling for cross-environment evaluation of ARIA snapshot logic.
- Updated browser action methods to resolve and target elements using ARIA-ref handles.
`omp --approval-mode=yolo acp` was rewritten to `launch --approval-mode=yolo
acp`, swallowing `acp` as a launch prompt so the yolo override never reached the
ACP command path (the ACP permission gate from #2097 stayed in always-ask).
`resolveCliArgv` only inspected `argv[0]`, so any leading global option flag hid
the real subcommand. It now scans past leading flags using the launch parser's
value-consumption contract (a flag's value is never mistaken for the subcommand,
e.g. `--model acp`) and hoists the recognized subcommand to the front with the
flags preserved as its own argv. Genuine launch prompts are untouched.
The flag value-consumption rule is factored into `cli/flag-tables.ts`
(`flagConsumesValue` + the shared `isUnknownLongValueCandidate`) so the resolver
and the profile bootstrap share one source of truth.
Fixed permission mode not respected in ACP mode ([#2970](https://github.com/can1357/oh-my-pi/issues/2970))
Fixes#2970
The background message reader matched every incoming message against the
pending client-request map by id before checking for a `method`. Server
request ids live in the server's own id space and routinely collide with
the client's in-flight request ids, so a server-originated
`workspace/configuration` pull whose id matched a pending request (e.g. a
basedpyright pull landing while a `documentSymbol` request with the same
id was open) was swallowed as a bogus response: the client request
resolved with `undefined` and the pull was never answered, wedging
servers that gate analysis on configuration.
Route any message carrying a `method` as a server request (or
notification) before id-matching responses, so every config pull is
answered under lazy init -- parity with the warmup/reload path, which
escaped the bug only because it issues no concurrent semantic request
while the cold-start pulls drain. lsp.lazy default is unchanged.
Fixes#3001
Stopped routing internal URL directory reads through the filesystem tree renderer so vault:// and '/data/workspaces/can1357__oh-my-pi__3116/.omp-session/2026-06-20T09-33-10-396Z_019ee460-817c-7000-8139-f8e2927809dd/local' keep their custom navigable listings.
Allow file-backed internal URL handlers to return existing directories as resources so read can list them and search/find can walk their source paths.
Fixes#3116
The ASCII table renderer in `sqlite-reader.ts` shrank columns down to
`MIN_COLUMN_WIDTH=1` to fit the 120-cell budget. With ~20+ columns
(the reporter had 33) every multi-char cell collapsed to a lone `…`
and the final per-line `truncateToWidth(..., MAX_RENDER_WIDTH)` then
chopped the right edge — so the read tool returned a table of nothing
but ellipses with the rightmost cells missing entirely.
Bump the per-column floor to 3 (so cells always show at least two real
glyphs alongside the ellipsis) and, when the column count alone forces
the floor over budget, fall back to a per-row vertical block layout —
mirroring `psql`'s expanded display mode. Each row becomes a
`column: value` group with column names padded so colons align and
the value line truncated to the same 120-cell budget.
Fixes#3107
Cleared pending-preview markers when resolve has no runnable handler so a stale gate cannot keep forcing resolve after the invoker is gone.
Added regression coverage for apply and discard draining stale pending markers.
Fixes#3061
- Replaced usage of `ReturnType<typeof setTimeout>` and `ReturnType<typeof setInterval>` with the explicit `Timer` type across the codebase.
- Updated several type definitions and function signatures to use concrete types instead of inferred return types for improved clarity and maintainability.
The PR fixed the -32601 hang for the defined server->client refresh
requests but missed two real spec methods of the identical class:
workspace/inlineValue/refresh (LSP 3.17) and workspace/foldingRange/refresh.
A server emitting either still received Method not found and could stall.
Add both to the void-ack chain and extend the regression test.
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
handleServerRequest fell through to a JSON-RPC -32601 Method not found for several defined server -> client requests (window/showMessageRequest, window/showDocument, workspace/{semanticTokens,inlayHint,codeLens,codeAction,diagnostic}/refresh). Servers that block on a real reply -- the same failure class as the client/registerCapability hang fixed in #3029 -- could stall waiting for an acknowledgement that never came.
Reply with the spec no-op result instead: null for showMessageRequest / *Refresh, { success: false } for showDocument. Headless omp cannot honour the UI surface, but it still owes a defined response.
Fixes#3044
Stopped image tool registration from resolving provider credentials during session startup. Added coverage that registration stays lazy while execution still resolves credentials.
Fixes#3036
- Accepted client/registerCapability and client/unregisterCapability server requests so Expert can continue startup before semantic requests.\n- Allowed JSON-RPC string ids on LSP request/response handling.\n- Added a regression test for dynamic registration gating hover.\n\nFixes #3029
- Switch all UI components and tests from sharp box corners (`boxSharp`) to rounded ones (`boxRound`).
- Update `Theme` to re-export sharp junction symbols (tees and cross) under `boxRound` to ensure consistent divider rendering in rounded boxes.
- Remove outdated architectural notes regarding forced tool-choice queues in documentation.
- Introduced `SoftToolRequirement` to support non-invasive tool enforcement with lifecycle management and escalation.
- Added `ToolChoiceDirective` to coordinate hard and soft tool requirements within the agent loop.
- Optimized preview workflows in `coding-agent` by replacing forced tool choices with non-forcing pending invokers.
- Enhanced `CompactionSummaryMessage` to prioritize structured rendering for tool requirement reminders.
Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.
Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init
All 522 previously-failing Windows tests now pass.
- Updated log and error truncation messages to use a consistent `[...Nch elided...]` format.
- Standardized diagnostics and prompt text output to improve consistency in reporting truncated information.
- Standardized elision markers across all tool outputs and filters to use cohesive `[...N [type] elided...]`, `[...Nln elided...]`, and `[...xB elided...]` syntax.
- Updated documentation, prompts, and test expectations to reflect the unified elision format.
- Improved transcript viewer robustness by preventing content aliasing through path-inclusive signature hashing.
- Added logic to clear stale transcript content when associated session files are deleted, accompanied by verifying test cases.
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
- Swapped legacy `zodToWireSchema` for `toolWireSchema` to normalize tool schemas.
- Updated `getSchemaPropertyKeys` in `tool-index.ts` to process schemas via `toolWireSchema`.
- Refactored tool token estimation in `context-usage.ts` to utilize the new schema helper.
- Fixed ArkType assertion checks in test helpers to correctly verify instances against `arkType.errors`.
- Aligned test suites and mock specifications with raw schema-based parameters instead of manually stringified JSON structures.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Added PDF member read syntax (`doc.pdf:<member>`) and trailing-colon listing.
- Added asset-read handling to serve extracted PDF images as inline image content.
- Added PDF image extraction caching keyed by size and mtime with marker files for retries.
- Added basename validation that rejects unknown/traversal-like PDF members and shows available names.
- Removed render_mermaid from tool discovery, task definitions, and registries.
- Removed renderMermaid setting and prompt/docs references tied to the deleted tool.
- Added maxWidth and theme color options to Mermaid ASCII resolution in markdown flow.
- Re-rendered Mermaid ASCII in both directions and clipped output to available width.
Connected createAgentSession ToolSession instances to AgentSession.getImageAttachments so inspect_image can resolve Image #N and attachment://N in real sessions. Added coverage that the real inspect_image tool session sees live user attachments.
Resolved inspect_image attachment labels and attachment URIs against the latest chat image attachments before falling back to file-path loading. Added regression coverage for Image #N labels, bracketed image markers, attachment://N URIs, missing attachment diagnostics, and cwd-independent resolution.
Fixes#2787
Re-throw ToolAbortError from soft-expired issue and PR synchronous refreshes instead of falling back to stale cached content.
Cover the abort path in github-cache tests.
Fixes#2684
Refresh soft-expired issue and PR view cache rows synchronously before returning content, while keeping PR diff rows on stale-first refresh semantics.
Add stale fallback warnings when a live refresh fails and cover the cache/protocol behavior in tests.
Fixes#2684
Restricted MSYS and WSL drive alias normalization to forward-slash roots so native Windows root-relative paths like \d\logs stay on the current drive.
Fixes#2634
Mapped MSYS and WSL drive aliases before bash cwd validation and brush filesystem resolution so cd, stat-style tests, and tool cwd handling agree on Windows.
Fixes#2634
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.