- Aligned hashline and chunk token handling to `HASHLINE_BIGRAMS` in `computeLineHash` and parsing regexes. - Removed hashline file-level `move`/`delete` options from schemas and execution so edits are in-place updates only. - Hardened hashline prefix parsing to reject non-bigram IDs, preventing false stripping of `#` comment lines. - Updated `read()`/prompt wording so raw reads skip prefixes and examples now show anchors like `123#th`. - Updated hashline chunk tests for new `HASHLINE_BIGRAMS` IDs and replaced stale hash constants.
5.6 KiB
Reads files using syntax-aware chunks. Also inspects directories, archives, SQLite databases, images, documents (PDF/DOCX/PPTX/XLSX/RTF/EPUB/ipynb), and URLs.
The chunk-aware `read` variant returns AST-scoped chunks with current checksum IDs for structural editing, and otherwise behaves like `open` for non-code content. - You **MUST** parallelize calls when exploring related files - For URLs, `read` fetches the page and returns clean extracted text/markdown by default (reader-mode). It handles HTML pages, GitHub issues/PRs, Stack Overflow, Wikipedia, Reddit, NPM, arXiv, RSS/Atom, JSON endpoints, PDFs, etc. You **SHOULD** reach for `read` — not a browser/puppeteer tool — for fetching and inspecting web content.Parameters
path— file path or URL; may include:selectorsuffix (required)sel— optional selector for chunks, line ranges, listing, or raw modetimeout— seconds, for URLs only
Selectors
sel value |
Behavior |
|---|---|
| (omitted) | Read full file as chunks (up to {{DEFAULT_LIMIT}} lines) |
class_Foo |
Read a specific chunk |
class_Foo.fn_bar#thth~ |
Read a chunk region (body ~ / head ^) by ID |
? |
List all chunk paths with IDs |
L50 |
Read from line 50 onward (shorthand for L50 to EOF) |
L50-L120 |
Read lines 50 through 120 |
L20-L20 |
Read exactly one line |
raw |
Raw content without transformations (for URLs: untouched HTML) |
Max {{DEFAULT_MAX_LINES}} lines per call.
Chunks
Each anchor @full.chunk.path#thth (with - prefixes for nesting depth) in the output identifies a chunk. Use full.chunk.path#thth as-is to read truncated chunks.
If you need a canonical target list, run read(path="file", sel="?"). That listing shows chunk paths with IDs and is the safest structural discovery mode. Summary lines in this listing are orientation hints; follow a selector with read(path="file", sel="chunk#ID") or use raw when you need exact source.
Line numbers in the gutter are absolute file line numbers.
{{#if chunkAutoIndent}}
Chunk reads normalize leading indentation so copied content round-trips cleanly into chunk edits.
{{else}}
Chunk reads preserve literal leading tabs/spaces from the file. When editing, keep the same whitespace characters you see here.
{{/if}}
raw shows the file's literal whitespace. Structured chunk views may normalize or display indentation for edit round-tripping, so use raw when exact tabs/spaces matter, especially inside markdown fenced code blocks.
IDs change after every edit. Use the new IDs from the edit response or refresh with sel="?" before the next write/delete. insert selectors may omit IDs, but still prefer fresh paths after structural edits.
Parser boundaries vary by language: TypeScript/JavaScript decorators and JSDoc above decorated methods may appear as sibling chunk#ID entries, Python decorators are part of the function/class head, Python docstrings are body lines, and Python enum members or nested closures may remain opaque inside their parent chunk. Decorated Python ^ writes and Python ^ deletes are rejected for safety.
Markdown sections, lists, and tables are structural chunks. Recognized pipe tables expose row_N children for row-level edits; list items and table cells are not independently addressable. Fenced code blocks with a declared language are parsed again when possible, so functions inside a markdown fence can appear as addressable nested chunks.
Chunk trees: JS, TS, TSX, Python, Rust, Go. Others use blank-line fallback.
Inspection
Extracts text from PDF, Word, PowerPoint, Excel, RTF, EPUB, and Jupyter notebook files. Can inspect images.
Directories & Archives
Directories and archive roots return a list of entries. Supports .tar, .tar.gz, .tgz, .zip. Use archive.ext:path/inside/archive to read contents.
SQLite Databases
When used against a SQLite database (.sqlite, .sqlite3, .db, .db3), returns structured database content.
file.db— list tables with row countsfile.db:table— table schema + sample rowsfile.db:table:key— single row by primary keyfile.db:table?limit=50&offset=100— paginated rowsfile.db:table?where=status='active'&order=created:desc— filtered rowsfile.db?q=SELECT …— read-only SELECT query
URLs
Extracts content from web pages, GitHub issues/PRs, Stack Overflow, Wikipedia, Reddit, NPM, arXiv, RSS/Atom feeds, JSON endpoints, PDFs at URLs, and similar text-based resources. Returns clean reader-mode text/markdown — no browser required. Use sel="raw" for untouched HTML; timeout to override the default request timeout. You SHOULD prefer read over a browser/puppeteer tool for fetching URL content; only use a browser when the page requires JS execution, authentication, or interactive actions (clicks, forms, scrolling).