Validated path-like renderer inputs before calling path helpers so provider-supplied arrays or objects cannot crash TUI rendering before schema validation reports the bad tool call.
Added renderer regression coverage for read, write, and edit call/result components with array and object path arguments.
Fixes#4525
- Introduced an automated retry recovery system to track, manage, and persist recovered error states within agent sessions.
- Enabled compact transcript rendering for recovered auto-retry errors by removing heuristic commit machinery.
- Improved raw read tracking and provenance in the ReadTool to support refined file snapshot recording and hashline editing.
- Excluded recovered assistant messages from default model context and updated event controllers to handle retry recovery life cycles.
Skipped the workflow notice on raw selectors so bounded raw reads stay byte-for-byte, and taught resolveToolSearchScope to request pathOnly resolution so ast_grep/ast_edit scope-only calls no longer trip the inline-content cap on large artifacts.
Fixes#4482
shortenPath() the artifact.path leaked into the read-tool 'Artifact storage' and 'Unbounded raw read blocked' notices so $HOME never appears verbatim; details.resolvedPath keeps the absolute path for tooling.
Fixes#4482
Resolved artifact:// reads to backing files before selector handling, streamed bounded reads, and blocked unbounded raw reads for large artifacts with recovery guidance.
Fixes#4482
Codex reviewer flagged that the ranged-read fallback recommended by
the patcher's over-cap reveal path (`path:N-M`) itself applies the
512-column cap and still records the displayed line numbers via
`recordSeenLinesFromBody`. On a minified wide anchor line, `read
file:N` shows only the clipped prefix but the line lands in the
tag's seenLines — a subsequent edit anchored at N then slips past
the seen-line guard.
- `packages/coding-agent/src/edit/file-snapshot-store.ts`:
`recordSeenLinesFromBody` grows an optional `excludedLines` set;
parsed line numbers matching it are filtered before recording.
- `packages/coding-agent/src/tools/read.ts`:
`#readLocalFileMultiRange` and the single-range disk path build a
`clippedLines` set alongside `columnTruncated` for both direct-range
lines and `buildLineEntriesWithBlockContext` context lines, then
pass it into `recordSeenLinesFromBody`.
- `packages/coding-agent/src/tools/grep.ts`:
same wiring for the match line via `match.truncated`, plus a
conservative length+`...`-marker heuristic for context lines
(native `crates/pi-natives/src/grep.rs` `truncate_line` doesn't
propagate a per-line flag on `contextBefore`/`contextAfter`; a
proper native-side flag is a follow-up).
- `packages/coding-agent/test/edit/seen-line-guard.test.ts`:
new case reads a 4KB single line and asserts the clipped line
number stays out of `seenLines` and the edit against it still
rejects with the seen-line guard.
- Fixed grep and ast-grep tools rejecting fuzzy url-shaped paths like schemeless www. or collapsed-scheme spellings.
- Applied the extra-CA wrapper to the model registry default fetch to respect the NODE_EXTRA_CA_CERTS environment variable.
- Moved isReadableUrlPath utility to path-utils to share recognition logic between reading and searching pipelines.
- Implemented a fallback that checks if a directory matching a fuzzy URL path exists locally before resolving it as an external URL.
- Added explicit errors for unsupported URL schemes to replace misleading local-path errors.
Five hot-path performance fixes + two low-risk allocation reductions.
No behavior change; all derived counts/orderings are identical.
- session-manager pathTo: leaf->root walk used branch.unshift() per node
(O(n^2) over branch length); now push + single reverse. Backs
getBranch(), hit at ~17 sites per turn.
- edit/modes/patch: collapseConsecutiveSharedLines / collapseRepeatedBlocks
/ trimCommonContext built shared-line sets via
new Set(oldLines.filter(l => newLines.includes(l))) -> O(old*new) per
hunk. Precompute new Set(newLines) and use .has() -> O(old+new).
- task/executor appendRecentOutputTail: re-split + filter + slice + reverse
of the full (up to 8KB) recentOutputTail on every text_delta token. Fast
path extends the current last line in place; full recompute only when a
newline boundary or truncation changes the window. tailLastLineRepresentable
flag guards the trailing-whitespace-only-line edge case.
- task/render renderResult: header booleans (3x .some) + footer counts
(3x .filter) + request total (.reduce) re-scanned details.results ~30x/sec
via the spinner. Single pass derives aborted/failed/mergeFailed/success
counts + requestTotal; booleans derived from counts.
- task/render extractIncrementalReviewResult: re-called normalizeYieldData
internally though both callers had already normalized the same yield data.
Signature now takes pre-normalized RenderYieldItem[].
Honorable mentions (allocation reduction, no algorithmic change):
- config/model-resolver: hoist case-folded pattern out of matchModel filter
passes; build the O(n) preference context once per role in
resolveModelRoleValue and reuse across fallback patterns.
- tools/read countTextLines: count newlines directly instead of allocating
via split("\n"); hashline formatter reuses the line count instead of
recomputing.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Adjusted hashline header formatting to preserve absolute file paths instead of truncating them to basenames.
- Ensured absolute paths are passed through shortenPath to allow resolution while keeping home directory references concise.
- Prevented edit failures when reading files outside the workspace by ensuring tags remain resolvable.
- Updated `formatReadHashlineHeader` to use `path.basename` instead of the full path.
- Applied the new formatter across record and display functions in the read tool.
- Added and updated tests to verify that hashline headers now contain only the filename for nested files.
- buildSshTarget rejects destinations beginning with "-" (SSH argument-injection / local RCE guard)
- gate ssh:// read/search/write at the exec approval tier; substring scan covers search's pre-expansion delimited paths and write's hashline-wrapped paths
- validate the entire materialized buffer as UTF-8 instead of only the first 8 KiB prefix
- write peels read selectors (raw/conflicts) so it targets the same file read does, and rejects line-range/malformed selectors instead of silently stripping them
- write to a uniquely named remote temp; document symlink-replacement on write as a v1 limit
The streaming reader's NUL check only walked completed lines collected
from streamLinesFromFile, so a binary blob whose first newline lay past
the byte budget (videos, archives, packed JSON) left collectedLines
empty and slipped through to the firstLineExceedsLimit branch — which
emitted the decoded preview as text instead of the intended refusal.
Sniff firstLinePreview alongside collectedLines so the existing refusal
fires uniformly. Also added a regression test that uses a 256 KiB blob
with no 0x0A bytes to actually exercise the firstLineExceedsLimit
path — the previous 6-byte test fit in one collected line and never
covered the bug.
Fixes#3448
Routed file-backed '/data/workspaces/can1357__oh-my-pi__3448/.omp-session/2026-06-25T07-25-13-303Z_019efdab-28d7-7000-a7a4-e20282508056/local' reads through the normal filesystem reader so binary detection, document/image handling, and streaming safeguards apply before content is materialized.
Hardened the local protocol handler to return metadata-only refusals for binary/container resources instead of decoding them with Bun.file().text().
Fixes#3448
Threaded the caller's loaded skills through internal URL resolution so skill:// handlers do not depend on process-global skill state during tool execution.
Fixes#3436
- Standardized `local://` image processing to prevent file corruption during decoding.
- Refactored local path resolution logic to enforce safety constraints and path containment.
- Implemented an image fast-path in `ReadTool` to correctly render local images before text decoding.
- Added comprehensive test coverage for image rendering, text compatibility, and path security.
- Resolved an event loop hang associated with `omp --resume` operations.
- Standardized elision markers across all tool outputs and filters to use cohesive `[...N [type] elided...]`, `[...Nln elided...]`, and `[...xB elided...]` syntax.
- Updated documentation, prompts, and test expectations to reflect the unified elision format.
- Improved transcript viewer robustness by preventing content aliasing through path-inclusive signature hashing.
- Added logic to clear stale transcript content when associated session files are deleted, accompanied by verifying test cases.
- Centralized archive operations into a new `utils/zip.ts` module with unified support for ZIP, tar, and tar.gz formats.
- Optimized ZIP reading using lazy, ranged central-directory access and implemented ZIP64 support for large files.
- Hardened archive extraction with directory traversal protection and configured memory limits for loading and extraction.
- Refactored tool-specific logic to utilize the new centralized utility and deleted the redundant `archive-reader.ts`.
- Added support for extracting and caching images from PDF documents during Markit conversion.
- Implemented a rewriting step to replace PDF markdown image placeholders with clickable reference commands.
- Introduced specific sub-path routing (`file.pdf:image.png`) allowing direct reading of extracted PDF images.
- Implemented caching for extracted PDF images utilizing unique keys and a `.extracted` marker file.
- Added explicit ArkType schema descriptions across all coding agent tool definitions.
- Updated schema definitions in autoresearch and commit tools with descriptive wrappers.
- Documented tool schema enhancements in the packages/coding-agent CHANGELOG.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
- Added PDF member read syntax (`doc.pdf:<member>`) and trailing-colon listing.
- Added asset-read handling to serve extracted PDF images as inline image content.
- Added PDF image extraction caching keyed by size and mtime with marker files for retries.
- Added basename validation that rejects unknown/traversal-like PDF members and shows available names.
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
- Read tool result details now include optional per-line number arrays so displayed content can retain original line mappings.
- Read tool group and renderer now forward code start-line and line-number metadata into code-cell rendering.
- Code cell rendering now draws aligned line-number gutters and pads hidden-line hints to keep gutter alignment.
Merge displayed bridge-backed range and multi-range read lines into the existing hashline snapshot provenance so INS.POST anchors pass visible-line validation after ACP reads.
Fixes#2773
- Tracked seen-line provenance in snapshots and propagated it from read/search/ast-grep rows.
- Rejected hashline edits on unseen lines before patching, throwing unseen-line errors.
- Rejected single-line block anchors in strict mode and dropped them in unresolved lenient mode.
- Trimmed one-sided keeper-echo duplicates during multi-line replacements with warning output.
- Recomputed `tools.discoveryMode: "auto"` in the deferred MCP closure in `sdk.ts` once the real tool count is known: a toolset crossing the threshold now flips discovery on, registers and activates `search_tool_bm25`, and skips `activateAll` instead of force-activating every MCP tool.
- Guarded the deferred MCP task against disposed sessions: added `AgentSession.isDisposed` and `enableMCPDiscovery()`, and the late connect now calls `disconnectAll()` instead of refreshing tools onto a dead session.
- Cleared `#fastPathKey`/`#fastPathItems` in `AssistantMessageComponent.invalidate()` so theme/symbol changes rebuild reused Markdown children instead of keeping stale captured themes.
- Memoized unusable read summaries as a `false` sentinel in `read.ts` so the per-session LRU no longer retains full sources of unsummarizable files.
- Broadened `HAS_REF_DEF` in `markdown.ts` to match backslash-escaped reference labels (`[a\]b]: x`) and cleared frozen stream-lex state on blank `setText()`.
- Added regression tests: deferred auto-discovery flip and mid-connect dispose (`sdk-mcp-auto-discovery.test.ts` + `many-tools-mcp.ts` fixture), fast-path child rebuild on invalidate, and escaped-ref-def incremental-lex equivalence.
Tracks successful file reads per resolved base path (selector stripped) for the session; once a path has been whole-file-read three times, every subsequent read for that path appends a one-line nudge suggesting narrower line-range re-reads or the context echoed in edit results. Non-file sources (URLs, internal resources, directories, archives, SQLite, images) are never counted.
Updated the read tool's provider-visible path schema, prompt docs, CLI help, and internal URL docs so URL and internal URI targets are advertised consistently. Added schema coverage for the read path description.\n\nFixes #2215
tar/tgz stat-gated at 256MB, zip entries reject oversized declared sizes; raw ?q= sqlite capped at 1000 rows; giant-file reads stop scanning to EOF; multi-range reads slice one pass; malformed URL selectors error instead of dumping; archive-root selectors, member tag immutability, case-insensitive selector tokens, session-pinned artifact lookups, shared+escaped suffix globs; archive dir listings honor offsets; binary files get a NUL-sniff notice.
- Defer MCP server discovery off the first-paint critical path for UI
sessions; tools and slash commands stream in through the existing
live-refresh channel once each server connects (non-UI modes keep the
blocking path). ~290ms off first paint with MCP servers configured.
- Build the model catalog's canonical-equivalence index lazily on first
read instead of eagerly in the ModelRegistry constructor. A default
interactive launch never reads it pre-paint, moving the ~210ms build
(over ~3,200 models) off the critical path: ~244ms (~16%) off cold boot.
- Memoize per-session read summaries (tree-sitter parse) on the content
hash of the freshly-read bytes; the file is still read fresh each call
so results stay correct. Repeat same-file summary read 17ms -> 2.4ms.
- Reuse the Markdown subtree across streaming reveal ticks, memoize
grapheme counting, and stop re-highlighting finalized thinking blocks.
- Attribute the previously-unlabeled synchronous boot region in the
PI_TIMING table and add a bench:guard boot-regression target.
- Added tree-sitter `enclosing_block_boundaries` API with line range models.
- Added N-API `enclosingBlockBoundaries` bridge and exported JS declarations.
- Replaced matching-bracket context resolution with source-aware block context in read and diff flows.
- Passed source path through diff/read generators to surface native block boundary previews.
- Added matching-bracket utilities to locate partner lines for visible spans.
- Added read tool output to use displayContent text and startLine for bracket-aware previews.
- Added matching-bracket context rows to generateDiffString and generateUnifiedDiffString.
- Adjusted diff row insertion to deduplicate and keep contiguous changed groups together.
- Enabled clickable read-path output for result rows, summaries, and previews.
- Resolved read links from result paths, source metadata, internal URLs, and absolutes.
- Preserved selector suffixes while rendering line-anchor hyperlinks.
- Ignored aborted signals for plain-file and directory reads while keeping conflicts-cancel behavior.
- Added tests for non-abortable read behavior and link-label rendering regression coverage.
- Updated `executeToolCalls` to always pass the active `toolSignal` into `tool.execute` rather than bypassing it for non-abortable tools.
- Removed the `nonAbortable` option from `AgentTool` and from the read, write, and edit tools so they can no longer opt out of abort handling.
- Documented the cancellation behavior change in the affected tool docs and package changelogs.
- Split top-level and delimited read selectors into separate rows before grouping.
- Merged duplicate same-file read selectors into one summarized row with ellipsis truncation.
- Computed grouped-read status and totals from aggregated rows for accurate summaries.
- Updated changelog entries and fixtures to document/read expectations.
- Added archive format sniffing to identify ZIP, TAR, and TAR.GZ from file bytes.
- Added MIME/extension and header-based routing for notebook, sqlite, and archive payloads.
- Added archive entry rendering with slash-terminated dirs and size suffixes.
- Added tests for archive, sqlite, notebook, and fallback binary dispatch scenarios.
- Fixed edit/read/search/ast-edit/ast-grep outputs to resolve OSC8 links from session cwd.
- Fixed grouped-file output classification to honor headerBase and fileScope for parent path resolution.
- Fixed read and write renderers to use resolved source/resolved paths as hyperlink targets.