- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.
- Added a per-message estimate cache in agent-core keyed by identity, with a
settle gate (assistants cache only with real usage + terminal non-error
stopReason; streaming partials bypass) and dual option-split WeakMaps for the
default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
with an exact-repeat outer-array reuse and slice-on-growth for append-only
turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
the prewalk plan-nudge scrub, via invalidateMessageCache /
registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
convert and repeat estimate are all >10x faster with noise under 20%.
Fixes#5934
TranscriptContainer gated committed-prefix compaction on version === undefined, so version-tracked AssistantMessageComponent history never compacted and every stream tick re-walked all N sealed blocks (depth-linear compose).
Compact fully-committed finalized blocks regardless of post-finalize version tracking: their rows are immutable native scrollback the terminal owns. A post-commit mutation no longer recommits on ordinary frames (no duplication) and rehydrates on the next destructive full replay (no loss). Adds the permanent bench/transcript-compose.bench.ts fixture; ratio(N5000/N500) drops 2.30 -> 0.90.
Fixes#5930
Each 30fps streaming-reveal tick re-segmented the whole revealed prefix via sliceGraphemes (per-tick slice cost grew with the prefix). BlockUnitCounter already memoized per-block grapheme counts; apply the same idea to slicing: cache the slice boundary per block index and re-segment only the per-step delta from the boundary cluster, so per-update cost grows with the delta instead of the rendered prefix.
buildDisplayMessage gains an optional sliceOf seam (default = sliceGraphemes, so existing callers are unchanged); the controller wires its BlockUnitCounter.slice through it via a #build helper that also de-duplicates six previously-repeated buildDisplayMessage call sites.
Only an exact (text, units) hit skips segmentation; the incremental guard (text === cached.text || text.startsWith(cached.text)) && units >= cached.units re-segments from the boundary cluster, so an append that extends the final cluster (e.g. a -> a combining acute, ZWJ family merge) is never stale.
Benchmark (bench/streaming-throughput.bench.ts, 30520-grapheme message, 61 ticks/episode): 23.78ms -> 2.27ms per episode (~10x). Regression tests vs a pure Intl.Segmenter reference cover fixed-text growing units, append growth, boundary-cluster extension, multi-block indices, shrink/regrow, full replace, and a 400-step seeded fuzz.
- Defer MCP server discovery off the first-paint critical path for UI
sessions; tools and slash commands stream in through the existing
live-refresh channel once each server connects (non-UI modes keep the
blocking path). ~290ms off first paint with MCP servers configured.
- Build the model catalog's canonical-equivalence index lazily on first
read instead of eagerly in the ModelRegistry constructor. A default
interactive launch never reads it pre-paint, moving the ~210ms build
(over ~3,200 models) off the critical path: ~244ms (~16%) off cold boot.
- Memoize per-session read summaries (tree-sitter parse) on the content
hash of the freshly-read bytes; the file is still read fresh each call
so results stay correct. Repeat same-file summary read 17ms -> 2.4ms.
- Reuse the Markdown subtree across streaming reveal ticks, memoize
grapheme counting, and stop re-highlighting finalized thinking blocks.
- Attribute the previously-unlabeled synchronous boot region in the
PI_TIMING table and add a bench:guard boot-regression target.
- Added a new `edit-lsp-writethrough.bench.ts` benchmark to measure writethrough latency with and without deferred diagnostics.
- Added a slow-server case to `lsp-diagnostics-freshness.test.ts` that confirmed deferred diagnostics arrive after a prompt inline return.
Before this change, every navigateTree → renderInitialMessages call path
performed two independent O(N) session-tree walks:
1. agent-session.ts:6586 buildDisplaySessionContext() [inside navigateTree]
2. ui-helpers.ts:402 sessionManager.buildSessionContext() [inside renderInitialMessages]
Changes:
- agent-session.ts: navigateTree() now calls sessionManager.buildSessionContext()
once, derives the display (deobfuscated) context from the raw result, and
returns the raw SessionContext in the result object.
- ui-helpers.ts: renderInitialMessages() accepts an optional prebuiltContext
parameter; reuses it when provided, falls back to buildSessionContext() otherwise.
- interactive-mode.ts: forwards prebuiltContext through the wrapper.
- modes/types.ts: updates InteractiveModeContext interface to match.
- selector-controller.ts: passes result.sessionContext from navigateTree into
renderInitialMessages(), closing the deduplication loop.
Bench (100-msg session, 200 iterations):
two walks [BEFORE]: 0.0702ms/op
one walk [AFTER]: 0.0298ms/op
Saved: 0.0404ms/navigation (57.5% reduction per navigate)
Tests: render-initial-messages-dedupe.test.ts asserts buildSessionContext is
called 0 times when a prebuilt context is passed, 1 time as fallback.
- Migrated timing measurements from `performance.now()` to `Bun.nanoseconds()` for higher precision benchmarking across all benchmark and timing-sensitive code.
- Updated elapsed time calculations to convert nanoseconds to milliseconds using division by 1e6 to maintain consistent time units.
- Increased benchmark precision from 4 to 6 decimal places for per-operation timing measurements.
- Refactored grep benchmark to use averaged timing values instead of storing individual iteration times, reducing memory overhead.
- Added native text utilities in pi-natives (visible width, truncation, slicing, segment extraction)
- Wired pi-tui to use native implementations with short-string JS fast path
- Optimized key matching via cached parsed key ids
- Refactored visual truncation to reuse Text instances
- Reduced theme color lookups by switching from Map to plain records
- Regenerated WASM artifacts
Benchmarks (after vs before, lower is better):
- truncateToWidth/ansi: 20.37ms vs 38.13ms (2000 iters)
- sliceWithWidth/ansi: 5.12ms vs 20.39ms
- extractSegments/ansi: 8.92ms vs 27.60ms
- truncateToVisualLines: 3.51ms vs 411.55ms (500 iters)
- visibleWidth/ansi: 0.24ms vs 0.31ms