- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
Each 30fps streaming-reveal tick re-segmented the whole revealed prefix via sliceGraphemes (per-tick slice cost grew with the prefix). BlockUnitCounter already memoized per-block grapheme counts; apply the same idea to slicing: cache the slice boundary per block index and re-segment only the per-step delta from the boundary cluster, so per-update cost grows with the delta instead of the rendered prefix.
buildDisplayMessage gains an optional sliceOf seam (default = sliceGraphemes, so existing callers are unchanged); the controller wires its BlockUnitCounter.slice through it via a #build helper that also de-duplicates six previously-repeated buildDisplayMessage call sites.
Only an exact (text, units) hit skips segmentation; the incremental guard (text === cached.text || text.startsWith(cached.text)) && units >= cached.units re-segments from the boundary cluster, so an append that extends the final cluster (e.g. a -> a combining acute, ZWJ family merge) is never stale.
Benchmark (bench/streaming-throughput.bench.ts, 30520-grapheme message, 61 ticks/episode): 23.78ms -> 2.27ms per episode (~10x). Regression tests vs a pure Intl.Segmenter reference cover fixed-text growing units, append growth, boundary-cluster extension, multi-block indices, shrink/regrow, full replace, and a 400-step seeded fuzz.
- Defer MCP server discovery off the first-paint critical path for UI
sessions; tools and slash commands stream in through the existing
live-refresh channel once each server connects (non-UI modes keep the
blocking path). ~290ms off first paint with MCP servers configured.
- Build the model catalog's canonical-equivalence index lazily on first
read instead of eagerly in the ModelRegistry constructor. A default
interactive launch never reads it pre-paint, moving the ~210ms build
(over ~3,200 models) off the critical path: ~244ms (~16%) off cold boot.
- Memoize per-session read summaries (tree-sitter parse) on the content
hash of the freshly-read bytes; the file is still read fresh each call
so results stay correct. Repeat same-file summary read 17ms -> 2.4ms.
- Reuse the Markdown subtree across streaming reveal ticks, memoize
grapheme counting, and stop re-highlighting finalized thinking blocks.
- Attribute the previously-unlabeled synchronous boot region in the
PI_TIMING table and add a bench:guard boot-regression target.
- Migrated timing measurements from `performance.now()` to `Bun.nanoseconds()` for higher precision benchmarking across all benchmark and timing-sensitive code.
- Updated elapsed time calculations to convert nanoseconds to milliseconds using division by 1e6 to maintain consistent time units.
- Increased benchmark precision from 4 to 6 decimal places for per-operation timing measurements.
- Refactored grep benchmark to use averaged timing values instead of storing individual iteration times, reducing memory overhead.
- Added native text utilities in pi-natives (visible width, truncation, slicing, segment extraction)
- Wired pi-tui to use native implementations with short-string JS fast path
- Optimized key matching via cached parsed key ids
- Refactored visual truncation to reuse Text instances
- Reduced theme color lookups by switching from Map to plain records
- Regenerated WASM artifacts
Benchmarks (after vs before, lower is better):
- truncateToWidth/ansi: 20.37ms vs 38.13ms (2000 iters)
- sliceWithWidth/ansi: 5.12ms vs 20.39ms
- extractSegments/ansi: 8.92ms vs 27.60ms
- truncateToVisualLines: 3.51ms vs 411.55ms (500 iters)
- visibleWidth/ansi: 0.24ms vs 0.31ms