- Treated ZIP-based .jar/.war/.ear/.apk as zip archives in archiveFormatFromPath and parseArchivePathCandidates so read/write member access works.
- Shared one archive-extension alternation between format detection and path splitting to stop them drifting.
- Derived the markit convertible-extension set from a single source of truth (utils/markit) matching the registered converters (pdf/docx/pptx/xlsx/epub), dropping legacy .doc/.ppt/.xls/.rtf that had no converter and only produced Unsupported format errors.
- Updated read/write tool prompts to document the zip-family extensions.
Fixes#5808
Removed process-local URL response reuse so read and URL-backed search paths fetch current content.
Added regressions for repeated reads and searches.
Fixes#5803
Tried existing unique workspace suffix matches before recovering a missing cwd path from the active approved plan. Added regression coverage for the precedence rule.
Recovered the active local plan when a model rewrites its local URL as a missing same-basename cwd-root path. Real working-tree files retain precedence.
Fixes#5704
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
- Disable context expansion in raw mode to ensure verbatim content extraction.
- Enforce strict adherence to requested line ranges when raw selectors are used.
- Remove padding in range calculations to prevent indistinguishable context lines in raw output.
The v16.3.12 explicit `selector` field (ff3b0c795c) was consumed by the
read tool but never threaded into the TUI renderers: ReadRenderArgs in
both readToolRenderer (read.ts) and ReadToolGroupComponent only derived
selectors from path-embedded `:sel` suffixes, so split-arg calls like
{ path, selector: "2-3" } rendered bare paths without line ranges or
raw modifiers.
Joined the explicit selector (trimmed, leading colons stripped, non-string
guarded) back onto the display path in renderCall, renderResult error and
success branches, and the grouped read summary, keeping hyperlinks on the
base path only.
Adopted from PR #4904 (both commits squashed), minus its unrelated
workflow-notice.md prompt churn.
Fixes#4899
Models emit optional string args as empty strings; since ff3b0c795
(#4622) read/grep rejected a present-but-empty selector as invalid
instead of behaving like an omitted one. Normalize empty and
whitespace-only selector params to undefined before validation.
Adopted from PR #4881 minus unrelated prompt churn.
Fixes#4879
Two Codex bot findings from earlier PR reviews were still open.
1. local:// URL selector shadow (read.ts): the local:// branch resolved
`local://foo:1-2` and rewrote readPath to `${localFile.path}:${sel}`,
then let splitPathAndSelPreferringLiteral run on the synthesized
string. A sibling literal `${localFile.path}:${sel}` file would win
over the intended URL selector semantics. The branch now promotes the
URL selector into the explicit-selector state and sets
readPath = localFile.path, so downstream literal-preferring routing
never re-splits the concatenation.
2. Delimited expansion before literal probe (path-utils.ts): grep called
expandDelimitedPathEntries before parsePathSpecs, and
splitDelimitedPathEntry only checked whether the peeled base of the
entry resolved. A real POSIX file whose name contained a delimiter
plus a selector-shaped tail (a;b:1-2) got split into ["a", "b:1-2"]
and never reached the literal-preferring probe. splitDelimitedPathEntry
now short-circuits on probeLiteralPathExists — "missing" is the only
outcome that lets delimiter expansion run.
Added regressions: `read local://notes.md:1-2` still slices the base file
when a sibling `notes.md:1-2` literal exists, and grep searches a real
`a;b:1-2` file without semicolon-splitting.
resolveExistingReadPath treated any stat failure other than ENOENT/ENOTDIR as
"exists" and any other resolved path was considered a hit. That silently
reinterpreted a real literal path such as test:1-2 as test plus selector 1-2
whenever the raw path was a dangling symlink, sat under an unreadable parent,
or hit a transient I/O error.
The new probeLiteralPathExists returns "exists" / "missing" / "unknown" from
an lstat probe. splitPathAndSelPreferringLiteral now falls back to the strict
selector split only on "missing"; both "exists" and "unknown" keep the raw
path, so an unreachable literal is never reinterpreted. Grep and read use
the same probe: the explicit selector branch keeps the literal path when
existence is uncertain, and only a definitive ENOENT/ENOTDIR lets structured
archive/sqlite/pdf dispatch take over.
Added regressions covering probeLiteralPathExists exists/missing/dangling-
symlink cases and splitPathAndSelPreferringLiteral over a dangling symlink.
The literal-path stat fallback made selector-shaped filenames accessible, but it did not give callers a deterministic way to read or grep a range from a literal filename such as test:1-2. Encoding that as test:1-2:1-2 remained recursively ambiguous if a longer literal file later appeared.
Read now accepts an optional selector field that is parsed independently from path. When selector is present, path is treated as the exact path first, so { path: "test:1-2", selector: "1-2" } always means lines 1-2 from the literal file test:1-2. Inline :<sel> remains supported for compatibility.
Grep now accepts an optional line-range selector field with the same literal-path behavior. Explicit selectors bypass path suffix peeling, while archive/internal/URL routing still handles non-literal structured paths.
Updated read/grep tool prompts and added deterministic regressions proving that a longer literal file like test:1-2:5-6 or test:1-2:2-2 does not change the meaning of { path: "test:1-2", selector: ... }.
The prior hunk placed the literal-preferring split after resolveArchiveReadPath,
resolveSqliteReadPath, and splitPdfImageMemberReadPath, so a real POSIX file
such as data.zip:1-2 or notes.db:1-2 still got hijacked: the archive/sqlite
resolvers matched the base extension, opened data.zip / notes.db, and errored
on the phantom :1-2 member before the literal file was ever considered.
Now the async splitter runs first. When the strict grammar would have peeled
a suffix but the literal path stats successfully, all three structured
dispatchers decline. Otherwise the ordering is unchanged, so archive/sqlite/
pdf-image reads keep working when the literal file does not exist.
Regression coverage adds `data.zip:1-2` and `notes.db:1-2` cases where the
base archive/sqlite file also exists on disk, exercising the exact ordering
bug the reviewer flagged.
splitPathAndSel unconditionally peels a trailing :<sel> chunk whenever it
matches the read-tool selector grammar (raw, conflicts, N-M, N+K, ...). On
POSIX, filenames may legitimately contain colons, so a real file named
test:1-2 or log:raw was shredded to test/log before either read.ts or
grep.parsePathSpecs stated anything and both surfaced "Path not found".
Added splitPathAndSelPreferringLiteral(rawPath, cwd) alongside the strict
splitter: it only overrides the peel when fs.stat succeeds against the raw
path. Read (execute) and grep (parsePathSpecs) call the async variant for
non-URL paths; internal-URL splitting stays unchanged. Regression covers
splitter fallbacks, read/grep behavior on literal-colon files, and that
:1-2 selectors still work when the base file is the only real match.
Fixes#4618
Validated path-like renderer inputs before calling path helpers so provider-supplied arrays or objects cannot crash TUI rendering before schema validation reports the bad tool call.
Added renderer regression coverage for read, write, and edit call/result components with array and object path arguments.
Fixes#4525
- Introduced an automated retry recovery system to track, manage, and persist recovered error states within agent sessions.
- Enabled compact transcript rendering for recovered auto-retry errors by removing heuristic commit machinery.
- Improved raw read tracking and provenance in the ReadTool to support refined file snapshot recording and hashline editing.
- Excluded recovered assistant messages from default model context and updated event controllers to handle retry recovery life cycles.
Skipped the workflow notice on raw selectors so bounded raw reads stay byte-for-byte, and taught resolveToolSearchScope to request pathOnly resolution so ast_grep/ast_edit scope-only calls no longer trip the inline-content cap on large artifacts.
Fixes#4482
shortenPath() the artifact.path leaked into the read-tool 'Artifact storage' and 'Unbounded raw read blocked' notices so $HOME never appears verbatim; details.resolvedPath keeps the absolute path for tooling.
Fixes#4482
Resolved artifact:// reads to backing files before selector handling, streamed bounded reads, and blocked unbounded raw reads for large artifacts with recovery guidance.
Fixes#4482
Codex reviewer flagged that the ranged-read fallback recommended by
the patcher's over-cap reveal path (`path:N-M`) itself applies the
512-column cap and still records the displayed line numbers via
`recordSeenLinesFromBody`. On a minified wide anchor line, `read
file:N` shows only the clipped prefix but the line lands in the
tag's seenLines — a subsequent edit anchored at N then slips past
the seen-line guard.
- `packages/coding-agent/src/edit/file-snapshot-store.ts`:
`recordSeenLinesFromBody` grows an optional `excludedLines` set;
parsed line numbers matching it are filtered before recording.
- `packages/coding-agent/src/tools/read.ts`:
`#readLocalFileMultiRange` and the single-range disk path build a
`clippedLines` set alongside `columnTruncated` for both direct-range
lines and `buildLineEntriesWithBlockContext` context lines, then
pass it into `recordSeenLinesFromBody`.
- `packages/coding-agent/src/tools/grep.ts`:
same wiring for the match line via `match.truncated`, plus a
conservative length+`...`-marker heuristic for context lines
(native `crates/pi-natives/src/grep.rs` `truncate_line` doesn't
propagate a per-line flag on `contextBefore`/`contextAfter`; a
proper native-side flag is a follow-up).
- `packages/coding-agent/test/edit/seen-line-guard.test.ts`:
new case reads a 4KB single line and asserts the clipped line
number stays out of `seenLines` and the edit against it still
rejects with the seen-line guard.
- Fixed grep and ast-grep tools rejecting fuzzy url-shaped paths like schemeless www. or collapsed-scheme spellings.
- Applied the extra-CA wrapper to the model registry default fetch to respect the NODE_EXTRA_CA_CERTS environment variable.
- Moved isReadableUrlPath utility to path-utils to share recognition logic between reading and searching pipelines.
- Implemented a fallback that checks if a directory matching a fuzzy URL path exists locally before resolving it as an external URL.
- Added explicit errors for unsupported URL schemes to replace misleading local-path errors.
Five hot-path performance fixes + two low-risk allocation reductions.
No behavior change; all derived counts/orderings are identical.
- session-manager pathTo: leaf->root walk used branch.unshift() per node
(O(n^2) over branch length); now push + single reverse. Backs
getBranch(), hit at ~17 sites per turn.
- edit/modes/patch: collapseConsecutiveSharedLines / collapseRepeatedBlocks
/ trimCommonContext built shared-line sets via
new Set(oldLines.filter(l => newLines.includes(l))) -> O(old*new) per
hunk. Precompute new Set(newLines) and use .has() -> O(old+new).
- task/executor appendRecentOutputTail: re-split + filter + slice + reverse
of the full (up to 8KB) recentOutputTail on every text_delta token. Fast
path extends the current last line in place; full recompute only when a
newline boundary or truncation changes the window. tailLastLineRepresentable
flag guards the trailing-whitespace-only-line edge case.
- task/render renderResult: header booleans (3x .some) + footer counts
(3x .filter) + request total (.reduce) re-scanned details.results ~30x/sec
via the spinner. Single pass derives aborted/failed/mergeFailed/success
counts + requestTotal; booleans derived from counts.
- task/render extractIncrementalReviewResult: re-called normalizeYieldData
internally though both callers had already normalized the same yield data.
Signature now takes pre-normalized RenderYieldItem[].
Honorable mentions (allocation reduction, no algorithmic change):
- config/model-resolver: hoist case-folded pattern out of matchModel filter
passes; build the O(n) preference context once per role in
resolveModelRoleValue and reuse across fallback patterns.
- tools/read countTextLines: count newlines directly instead of allocating
via split("\n"); hashline formatter reuses the line count instead of
recomputing.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Adjusted hashline header formatting to preserve absolute file paths instead of truncating them to basenames.
- Ensured absolute paths are passed through shortenPath to allow resolution while keeping home directory references concise.
- Prevented edit failures when reading files outside the workspace by ensuring tags remain resolvable.
- Updated `formatReadHashlineHeader` to use `path.basename` instead of the full path.
- Applied the new formatter across record and display functions in the read tool.
- Added and updated tests to verify that hashline headers now contain only the filename for nested files.
- buildSshTarget rejects destinations beginning with "-" (SSH argument-injection / local RCE guard)
- gate ssh:// read/search/write at the exec approval tier; substring scan covers search's pre-expansion delimited paths and write's hashline-wrapped paths
- validate the entire materialized buffer as UTF-8 instead of only the first 8 KiB prefix
- write peels read selectors (raw/conflicts) so it targets the same file read does, and rejects line-range/malformed selectors instead of silently stripping them
- write to a uniquely named remote temp; document symlink-replacement on write as a v1 limit
The streaming reader's NUL check only walked completed lines collected
from streamLinesFromFile, so a binary blob whose first newline lay past
the byte budget (videos, archives, packed JSON) left collectedLines
empty and slipped through to the firstLineExceedsLimit branch — which
emitted the decoded preview as text instead of the intended refusal.
Sniff firstLinePreview alongside collectedLines so the existing refusal
fires uniformly. Also added a regression test that uses a 256 KiB blob
with no 0x0A bytes to actually exercise the firstLineExceedsLimit
path — the previous 6-byte test fit in one collected line and never
covered the bug.
Fixes#3448
Routed file-backed '/data/workspaces/can1357__oh-my-pi__3448/.omp-session/2026-06-25T07-25-13-303Z_019efdab-28d7-7000-a7a4-e20282508056/local' reads through the normal filesystem reader so binary detection, document/image handling, and streaming safeguards apply before content is materialized.
Hardened the local protocol handler to return metadata-only refusals for binary/container resources instead of decoding them with Bun.file().text().
Fixes#3448
Threaded the caller's loaded skills through internal URL resolution so skill:// handlers do not depend on process-global skill state during tool execution.
Fixes#3436
- Standardized `local://` image processing to prevent file corruption during decoding.
- Refactored local path resolution logic to enforce safety constraints and path containment.
- Implemented an image fast-path in `ReadTool` to correctly render local images before text decoding.
- Added comprehensive test coverage for image rendering, text compatibility, and path security.
- Resolved an event loop hang associated with `omp --resume` operations.
- Standardized elision markers across all tool outputs and filters to use cohesive `[...N [type] elided...]`, `[...Nln elided...]`, and `[...xB elided...]` syntax.
- Updated documentation, prompts, and test expectations to reflect the unified elision format.
- Improved transcript viewer robustness by preventing content aliasing through path-inclusive signature hashing.
- Added logic to clear stale transcript content when associated session files are deleted, accompanied by verifying test cases.
- Centralized archive operations into a new `utils/zip.ts` module with unified support for ZIP, tar, and tar.gz formats.
- Optimized ZIP reading using lazy, ranged central-directory access and implemented ZIP64 support for large files.
- Hardened archive extraction with directory traversal protection and configured memory limits for loading and extraction.
- Refactored tool-specific logic to utilize the new centralized utility and deleted the redundant `archive-reader.ts`.
- Added support for extracting and caching images from PDF documents during Markit conversion.
- Implemented a rewriting step to replace PDF markdown image placeholders with clickable reference commands.
- Introduced specific sub-path routing (`file.pdf:image.png`) allowing direct reading of extracted PDF images.
- Implemented caching for extracted PDF images utilizing unique keys and a `.extracted` marker file.
- Added explicit ArkType schema descriptions across all coding agent tool definitions.
- Updated schema definitions in autoresearch and commit tools with descriptive wrappers.
- Documented tool schema enhancements in the packages/coding-agent CHANGELOG.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
- Added PDF member read syntax (`doc.pdf:<member>`) and trailing-colon listing.
- Added asset-read handling to serve extracted PDF images as inline image content.
- Added PDF image extraction caching keyed by size and mtime with marker files for retries.
- Added basename validation that rejects unknown/traversal-like PDF members and shows available names.
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
- Read tool result details now include optional per-line number arrays so displayed content can retain original line mappings.
- Read tool group and renderer now forward code start-line and line-number metadata into code-cell rendering.
- Code cell rendering now draws aligned line-number gutters and pads hidden-line hints to keep gutter alignment.
Merge displayed bridge-backed range and multi-range read lines into the existing hashline snapshot provenance so INS.POST anchors pass visible-line validation after ACP reads.
Fixes#2773