- Replaced the custom MuPDF-WASM PDF extraction and rendering pipeline with the new `pdfToMarkdown` native function from `@oh-my-pi/pi-natives`.
- Removed legacy MuPDF extraction modules, WASM embedding scripts, and PDF image extraction tools.
- Added OCR warnings and browser/text redirection for unsupported PDF image reads.
- Updated native package definitions, documentation, and test suites for the new PDF inspection capability.
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
Skipped binary inline-image payload bytes after the PDF ID operator so tokenizer scanning resumes at post-image content.
Advanced past stray closing delimiters while recovering malformed content streams to avoid zero-width tokenizer iterations.
Fixes#4512
The vendored markit engine kept `mupdf` external, but a single-file
`bun --compile` binary has no node_modules to resolve it from, so the
standalone binary aborted at startup with `Cannot find package 'mupdf'`
— the otherwise-lazy import is resolved eagerly at boot. Bundle mupdf and
embed its WASM blob (scripts/embed-mupdf-wasm.ts, reset after the build);
npm and source installs still load mupdf from node_modules.
Import mupdf lazily inside the PDF converter so the bundled markit chunk's
init stays synchronous: mupdf's top-level await otherwise made the chunk
init async and bun's compiled bundler failed to await it through the
barrel, exposing the converters before their module-level const tables
initialized (undefined EXTENSIONS). Also keeps the ~10MB wasm off non-PDF
document conversions.