- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
estimateTokens now charges for serialized anthropicServerTool blocks so context maintenance sees the server-tool payload replayed on the wire; excluded from the compaction floor like other encrypted reasoning.
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.
- Added a per-message estimate cache in agent-core keyed by identity, with a
settle gate (assistants cache only with real usage + terminal non-error
stopReason; streaming partials bypass) and dual option-split WeakMaps for the
default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
with an exact-repeat outer-array reuse and slice-on-growth for append-only
turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
the prewalk plan-nudge scrub, via invalidateMessageCache /
registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
convert and repeat estimate are all >10x faster with noise under 20%.
Fixes#5934