12 Commits

Author SHA1 Message Date
can1357 12238f55ca feat: implemented native ctok tokenization engine with model scopes
- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
2026-08-19 23:27:29 +02:00
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
can1357 65534c5c07 Merge PR #5947: perf(session): memoize convertToLlm and estimateTokens over settled history (@roboomp) 2026-07-18 19:42:52 +02:00
roboomp a28eb0f470 perf(session): memoized convertToLlm and estimateTokens over settled history
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.

- Added a per-message estimate cache in agent-core keyed by identity, with a
  settle gate (assistants cache only with real usage + terminal non-error
  stopReason; streaming partials bypass) and dual option-split WeakMaps for the
  default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
  with an exact-repeat outer-array reuse and slice-on-growth for append-only
  turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
  the prewalk plan-nudge scrub, via invalidateMessageCache /
  registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
  convert and repeat estimate are all >10x faster with noise under 20%.

Fixes #5934
2026-07-18 02:04:09 +00:00
roboomp 84c6b6aa58 fix(tui): compact committed finalized transcript history
TranscriptContainer gated committed-prefix compaction on version === undefined, so version-tracked AssistantMessageComponent history never compacted and every stream tick re-walked all N sealed blocks (depth-linear compose).

Compact fully-committed finalized blocks regardless of post-finalize version tracking: their rows are immutable native scrollback the terminal owns. A post-commit mutation no longer recommits on ordinary frames (no duplication) and rehydrates on the next destructive full replay (no loss). Adds the permanent bench/transcript-compose.bench.ts fixture; ratio(N5000/N500) drops 2.30 -> 0.90.

Fixes #5930
2026-07-18 00:52:57 +00:00
can1357 4054ca896a fix(coding-agent): wire reveal bench callbacks correctly 2026-07-01 21:47:54 +02:00
oldschoola 718c7cea2f perf(coding-agent): memoize incremental grapheme slicing in streaming reveal
Each 30fps streaming-reveal tick re-segmented the whole revealed prefix via sliceGraphemes (per-tick slice cost grew with the prefix). BlockUnitCounter already memoized per-block grapheme counts; apply the same idea to slicing: cache the slice boundary per block index and re-segment only the per-step delta from the boundary cluster, so per-update cost grows with the delta instead of the rendered prefix.

buildDisplayMessage gains an optional sliceOf seam (default = sliceGraphemes, so existing callers are unchanged); the controller wires its BlockUnitCounter.slice through it via a #build helper that also de-duplicates six previously-repeated buildDisplayMessage call sites.

Only an exact (text, units) hit skips segmentation; the incremental guard (text === cached.text || text.startsWith(cached.text)) && units >= cached.units re-segments from the boundary cluster, so an append that extends the final cluster (e.g. a -> a combining acute, ZWJ family merge) is never stale.

Benchmark (bench/streaming-throughput.bench.ts, 30520-grapheme message, 61 ticks/episode): 23.78ms -> 2.27ms per episode (~10x). Regression tests vs a pure Intl.Segmenter reference cover fixed-text growing units, append growth, boundary-cluster extension, multi-block indices, shrink/regrow, full replace, and a 400-step seeded fuzz.
2026-06-29 17:47:22 -07:00
metaphorics a3e3a905ac perf(coding-agent): defer boot work and cut streaming/read CPU
- Defer MCP server discovery off the first-paint critical path for UI
  sessions; tools and slash commands stream in through the existing
  live-refresh channel once each server connects (non-UI modes keep the
  blocking path). ~290ms off first paint with MCP servers configured.
- Build the model catalog's canonical-equivalence index lazily on first
  read instead of eagerly in the ModelRegistry constructor. A default
  interactive launch never reads it pre-paint, moving the ~210ms build
  (over ~3,200 models) off the critical path: ~244ms (~16%) off cold boot.
- Memoize per-session read summaries (tree-sitter parse) on the content
  hash of the freshly-read bytes; the file is still read fresh each call
  so results stay correct. Repeat same-file summary read 17ms -> 2.4ms.
- Reuse the Markdown subtree across streaming reveal ticks, memoize
  grapheme counting, and stop re-highlighting finalized thinking blocks.
- Attribute the previously-unlabeled synchronous boot region in the
  PI_TIMING table and add a bench:guard boot-regression target.
2026-06-09 18:58:22 +09:00
can1357 4f6ea5ca24 test(coding-agent): added writethrough deferred diagnostics benchmark and test coverage
- Added a new `edit-lsp-writethrough.bench.ts` benchmark to measure writethrough latency with and without deferred diagnostics.
- Added a slow-server case to `lsp-diagnostics-freshness.test.ts` that confirmed deferred diagnostics arrive after a prompt inline return.
2026-06-08 05:28:06 +02:00
cognitive 68baba82d2 perf(coding-agent): dedupe buildSessionContext walk on session-tree navigation
Before this change, every navigateTree → renderInitialMessages call path
performed two independent O(N) session-tree walks:

  1. agent-session.ts:6586  buildDisplaySessionContext()   [inside navigateTree]
  2. ui-helpers.ts:402      sessionManager.buildSessionContext()  [inside renderInitialMessages]

Changes:
- agent-session.ts: navigateTree() now calls sessionManager.buildSessionContext()
  once, derives the display (deobfuscated) context from the raw result, and
  returns the raw SessionContext in the result object.
- ui-helpers.ts: renderInitialMessages() accepts an optional prebuiltContext
  parameter; reuses it when provided, falls back to buildSessionContext() otherwise.
- interactive-mode.ts: forwards prebuiltContext through the wrapper.
- modes/types.ts: updates InteractiveModeContext interface to match.
- selector-controller.ts: passes result.sessionContext from navigateTree into
  renderInitialMessages(), closing the deduplication loop.

Bench (100-msg session, 200 iterations):
  two walks [BEFORE]: 0.0702ms/op
  one walk  [AFTER]:  0.0298ms/op
  Saved: 0.0404ms/navigation (57.5% reduction per navigate)

Tests: render-initial-messages-dedupe.test.ts asserts buildSessionContext is
called 0 times when a prebuilt context is passed, 1 time as fallback.
2026-04-27 03:55:14 +00:00
can1357 37c74bb19d perf: migrated timing measurements to Bun.nanoseconds() for higher precision
- Migrated timing measurements from `performance.now()` to `Bun.nanoseconds()` for higher precision benchmarking across all benchmark and timing-sensitive code.
- Updated elapsed time calculations to convert nanoseconds to milliseconds using division by 1e6 to maintain consistent time units.
- Increased benchmark precision from 4 to 6 decimal places for per-operation timing measurements.
- Refactored grep benchmark to use averaged timing values instead of storing individual iteration times, reducing memory overhead.
2026-02-05 06:19:49 +01:00
can1357 8ef11355e7 perf(tui,natives): added native text utilities with JS fast path for short strings
- Added native text utilities in pi-natives (visible width, truncation, slicing, segment extraction)
- Wired pi-tui to use native implementations with short-string JS fast path
- Optimized key matching via cached parsed key ids
- Refactored visual truncation to reuse Text instances
- Reduced theme color lookups by switching from Map to plain records
- Regenerated WASM artifacts

Benchmarks (after vs before, lower is better):
- truncateToWidth/ansi: 20.37ms vs 38.13ms (2000 iters)
- sliceWithWidth/ansi: 5.12ms vs 20.39ms
- extractSegments/ansi: 8.92ms vs 27.60ms
- truncateToVisualLines: 3.51ms vs 411.55ms (500 iters)
- visibleWidth/ansi: 0.24ms vs 0.31ms
2026-01-29 10:47:55 +01:00