- reject explicit ssh:// port 0 before connecting (Codex P2)
- keep a path-less ssh://host:port authority port out of selector peeling
- restore ResolveContext on ProtocolHandler.complete (symmetry with resolve/write)
- clarify search/read selector-parity docs + add read-side regression
- `read ssh://host/dir` lists a remote directory one level deep; `ssh://host/` lists the remote root
- add statRemotePath + listRemoteDir; resolve reads first and classifies on error (directory -> one-level listing, dirs-first, dotfiles included)
- directory resources carry isDirectory + immutable and expose no sourcePath
- search refuses a virtual (no-sourcePath) directory resource instead of grepping the listing text
- writeRemoteFile refuses a directory destination and cleans up its temp on that path
- buildSshTarget rejects destinations beginning with "-" (SSH argument-injection / local RCE guard)
- gate ssh:// read/search/write at the exec approval tier; substring scan covers search's pre-expansion delimited paths and write's hashline-wrapped paths
- validate the entire materialized buffer as UTF-8 instead of only the first 8 KiB prefix
- write peels read selectors (raw/conflicts) so it targets the same file read does, and rejects line-range/malformed selectors instead of silently stripping them
- write to a uniquely named remote temp; document symlink-replacement on write as a v1 limit
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.
- Clarify the auto-advancement logic for the in-progress pointer in the documentation and output.
- Add an overall completion count summary to the task list view.
- Update the list format to use standard checkbox indicators and explicit tags for task statuses.
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
Accepted displayed gallery lifecycle labels as --state aliases, rejected unknown values before rendering, and updated failed fixtures to render visibly failed states.
Fixes#3473
The streaming reader's NUL check only walked completed lines collected
from streamLinesFromFile, so a binary blob whose first newline lay past
the byte budget (videos, archives, packed JSON) left collectedLines
empty and slipped through to the firstLineExceedsLimit branch — which
emitted the decoded preview as text instead of the intended refusal.
Sniff firstLinePreview alongside collectedLines so the existing refusal
fires uniformly. Also added a regression test that uses a 256 KiB blob
with no 0x0A bytes to actually exercise the firstLineExceedsLimit
path — the previous 6-byte test fit in one collected line and never
covered the bug.
Fixes#3448
Routed file-backed '/data/workspaces/can1357__oh-my-pi__3448/.omp-session/2026-06-25T07-25-13-303Z_019efdab-28d7-7000-a7a4-e20282508056/local' reads through the normal filesystem reader so binary detection, document/image handling, and streaming safeguards apply before content is materialized.
Hardened the local protocol handler to return metadata-only refusals for binary/container resources instead of decoding them with Bun.file().text().
Fixes#3448
Threaded the caller's loaded skills through internal URL resolution so skill:// handlers do not depend on process-global skill state during tool execution.
Fixes#3436
- Standardized `local://` image processing to prevent file corruption during decoding.
- Refactored local path resolution logic to enforce safety constraints and path containment.
- Implemented an image fast-path in `ReadTool` to correctly render local images before text decoding.
- Added comprehensive test coverage for image rendering, text compatibility, and path security.
- Resolved an event loop hang associated with `omp --resume` operations.
- Update the eval tool to stream stdout chunks directly into the active cell's output buffer while the process is still running.
- Prevent long-running cells from appearing empty in the UI by surfacing incremental output before the backend resolves.
- Add regression tests to ensure streamed output is captured mid-execution and reconciled with final results.
- Removed deprecated eval prelude helpers `append`, `tree`, `diff`, `sort`, `uniq`, and `counter` from all supported runtimes.
- Cleaned up runtime implementations, protocol definitions, and UI rendering logic associated with the removed helpers.
- Updated project documentation, prompts, and test suites to reflect the reduced helper API surface.
- Recorded functional changes in the package changelog.
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
- Refactored `todo` tool to accept a single operation object instead of an `ops` array.
- Implemented parameter normalization to maintain backward compatibility with legacy array-based tool calls.
- Updated tool instructions, documentation, and UI rendering components to reflect the new interface.
- Added compatibility tests to verify rendering and execution for both legacy and current operation formats.
- Elevated `eval` to an essential tool to ensure availability across all discovery modes.
- Updated system and tool prompts to mandate the use of `eval` for non-trivial shell operations like conditionals, loops, heredocs, and complex pipelines.
- Restricted `bash` usage to simple binary invocations and single-fact computation to reduce shell-escaping and execution errors.
- Clamped tool output preview height to the available viewport rows to stop redundant banner commits.
- Added `outputBlockContentWidth` helper to accurately measure visual lines for scrollback budget calculations.
- Updated `bash` and `eval-render` output wrapping to account for block padding and inner content width.
- Added regression test confirming streaming tool output maintains a stable line count without duplicating headers.
Keep the ask tool on the current question when the custom answer editor is dismissed, so Escape from Other returns to the option selector instead of aborting or recording an empty answer.
Fixes#3269
- Changed default evaluation backend configuration to only enable Python and JavaScript by default.
- Implemented dynamic tool parameter generation to hide Ruby and Julia from the model's schema when they are disabled in settings.
- Updated tool summary and field descriptions to reflect the currently enabled runtime backends.
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
- Implemented `waitForSelector` and `waitForNavigation` methods in the browser tool API.
- Introduced per-operation fail-fast budget management with dynamic timeout clamping.
- Added validation for selector engines to reject unsupported Playwright-only platform features.
- Standardized error handling to provide descriptive, named timeouts for stalled browser operations.
- Promoted `write` and `find` tools to `essential` status to ensure they are always available regardless of discovery mode.
- Updated `DEFAULT_ESSENTIAL_TOOL_NAMES` to include these tools by default.
- Updated documentation and tests to reflect the change in default essential tool availability.
Fixes#3165
- Synchronize tool arguments with the component state upon receipt of `tool_execution_start` to ensure visual consistency when final update events are missed.
- Terminate active argument reveal streams to prevent late ticks from overwriting valid, fully-materialized tool arguments with stale partial data.
- Add test coverage to verify that tool UI components render finalized arguments even in the absence of intermediate streaming updates.
- Implemented `tab.ariaSnapshot()` to capture and represent page structures as ARIA-tree YAML.
- Introduced `tab.ref()` and ref-based selector parsing to enable precise element interaction via unique ARIA identifiers.
- Integrated automated script bundling for cross-environment evaluation of ARIA snapshot logic.
- Updated browser action methods to resolve and target elements using ARIA-ref handles.
- Overhauled stealth spoofing mechanisms for WebGL, screen dimensions, Web workers, and iframe contexts using prototype-aware injection.
- Centralized function string representation patching to improve mimicry of native browser behavior across global objects.
- Patched puppeteer-core to remove detectable evaluation markers and implement lazy, pull-style execution context management.
- Enabled support for capturing LLM request JSON dumps and adjusted launcher flags to improve organic request patterns.
- Remove the static "pending" hourglass icon from edit and write tool headers to reduce visual noise.
- Update multi-file status lines to use the active spinner icon directly instead of replacing a static icon, ensuring consistent liveness cues.
Stopped routing internal URL directory reads through the filesystem tree renderer so vault:// and '/data/workspaces/can1357__oh-my-pi__3116/.omp-session/2026-06-20T09-33-10-396Z_019ee460-817c-7000-8139-f8e2927809dd/local' keep their custom navigable listings.
Allow file-backed internal URL handlers to return existing directories as resources so read can list them and search/find can walk their source paths.
Fixes#3116
The ASCII table renderer in `sqlite-reader.ts` shrank columns down to
`MIN_COLUMN_WIDTH=1` to fit the 120-cell budget. With ~20+ columns
(the reporter had 33) every multi-char cell collapsed to a lone `…`
and the final per-line `truncateToWidth(..., MAX_RENDER_WIDTH)` then
chopped the right edge — so the read tool returned a table of nothing
but ellipses with the rightmost cells missing entirely.
Bump the per-column floor to 3 (so cells always show at least two real
glyphs alongside the ellipsis) and, when the column count alone forces
the floor over budget, fall back to a per-row vertical block layout —
mirroring `psql`'s expanded display mode. Each row becomes a
`column: value` group with column names padded so colons align and
the value line truncated to the same 120-cell budget.
Fixes#3107
Cleared pending-preview markers when resolve has no runnable handler so a stale gate cannot keep forcing resolve after the invoker is gone.
Added regression coverage for apply and discard draining stale pending markers.
Fixes#3061
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Simplified system and personality prompts for improved conciseness and clarity.
- Streamlined tool instruction sets and parameter descriptions across all agent modules.
- Refactored prompt documentation in `hashline` to clarify terminology and task-specific constraints.
- Updated tool metadata in TypeScript service definitions to align with reduced documentation verbosity.
- Removed the customizable `timeout` parameter from the tool schema and implementation.
- Fixed the timeout duration to 5 seconds to simplify tool usage.
- Updated documentation and error messages to advise narrowing search patterns instead of adjusting timeouts.
Stopped image tool registration from resolving provider credentials during session startup. Added coverage that registration stays lazy while execution still resolves credentials.
Fixes#3036
- Switch all UI components and tests from sharp box corners (`boxSharp`) to rounded ones (`boxRound`).
- Update `Theme` to re-export sharp junction symbols (tees and cross) under `boxRound` to ensure consistent divider rendering in rounded boxes.
- Remove outdated architectural notes regarding forced tool-choice queues in documentation.
- Introduced `SoftToolRequirement` to support non-invasive tool enforcement with lifecycle management and escalation.
- Added `ToolChoiceDirective` to coordinate hard and soft tool requirements within the agent loop.
- Optimized preview workflows in `coding-agent` by replacing forced tool choices with non-forcing pending invokers.
- Enhanced `CompactionSummaryMessage` to prioritize structured rendering for tool requirement reminders.