Files
oh-my-pi/packages/utils/src
Sunil Srivatsa 301908560e perf(read): materialize a local file once per read
The local text read path opened the same file for every consumer. A ranged
read of a file within the snapshot cap cost four opens and three decodes:
an 8KiB binary sniff, a streaming scan for the rendered window, a whole-file
read for bracket context, and another whole-file read to hash the snapshot.
Whole-file reads under the structural summarizer paid a fifth. Two of those
readers also ran normalizeToLF over the same bytes.

Read the bytes once at or below SNAPSHOT_MAX_BYTES and derive every view
from them: sniff the leading 8KiB of the buffer, slice the rendered window
out of it under the identical line and byte budgets, index bracket context
into its addressable lines, and hand the normalized text to the snapshot
store and the summarizer. Past the cap nothing wants the whole file, so the
streaming reader stays.

Line byte lengths are walked out of the buffer rather than measured on the
decoded strings, so reported byte counts and the truncation boundary stay
exact for content that is not valid UTF-8.

The buffered text is BOM-stripped for hashing, matching the decoder the
patcher's live read uses. A whole-file read of a BOM file previously hashed
its tag from BOM-bearing text, so the following edit only applied through
stale-hash recovery and told the model the file had changed externally when
it had not.

Also stop rejoining lines into a fresh whole-file string on the way to
tree-sitter when the caller still holds that text, and drop the unread
selectedBytesTotal accounting from the streaming reader.

Measured on 966 differential cases across CRLF, BOM, lone-CR, invalid-UTF-8,
oversized-line, empty, no-trailing-newline, multi-range and raw shapes: the
spurious recovery warning is the only behavioral difference. Raw reads, which
skip the tree-sitter parse that dominates everything else, get 30-45% faster
(2.7MB: 10.3ms -> 5.7ms); non-raw reads 1-3%.
2026-08-17 14:15:15 -07:00
..