- Added a new SQuAD-based context-compression benchmark script for evaluating recall conditions.
- Added model and shape command-line arguments along with pricing and shape configurations.
- Updated default model variants and added an unknown billing family to shape resolution.
- Updated shape resolution and model resolution tests to verify the new defaults.
- Collected distinct rendered frame widths for resume summaries.
- Covered mixed HQ/LQ archives with the existing foveation regression.
- Documented the fix in the snapcompact changelog.
Fixes#6712
- Re-compaction unfolds the prior archive's kept source verbatim, so archives written before includeThinking existed kept replaying reasoning to Claude (reasoning_extraction) even after the serializer fix; the prior text is now scrubbed when includeThinking is false, healing poisoned sessions at their next compaction.
Both compaction serializers reproduced prior assistant reasoning as text
bound for a Claude target, tripping Anthropic's reasoning_extraction
refusal and wedging Fable 5 sessions:
- context-full: serializeConversation rendered thinking verbatim inside
<thinking> tags via the anthropic dialect renderer. Now drops thinking
blocks when the summary target dialect is anthropic; other dialects
(e.g. Harmony) keep native reasoning.
- snapcompact: emitted ¶think sections baked into replayed archive
frames. Added an includeThinking serialize option (default true) and
wired the agent session to disable it for Anthropic-dialect models.
Fixes#6093
A branch whose last entry is a snapcompact CompactionEntry billed past the
compaction threshold (FRAME_TOKEN_ESTIMATE x frames) dead-ended on every
resume: prepareCompaction returns undefined (nothing after the entry to
summarize), and the #4786 elide/image rescue tiers only inspect
"message"/"custom_message" entries, so a type:"compaction" tail escaped both
and the "Compaction freed too little context" warning re-fired forever.
Add a dedicated first rescue tier that rebuilds the SAME archive locally (no
LLM, no network) by re-running snapcompact.compact() over the entry's
carried-forward source text at a maxFrames derived from the trigger
threshold's recovery band instead of the window-fit budget: planArchive
truncates the oldest chars to fit, so the rebuilt entry genuinely shrinks.
Persisting through appendCompaction lets the write-time
superseded-compaction elision drop the stale frame payload from the JSONL,
and the pass skips the misleading no-progress warning.
Fixes the loop reported in
https://github.com/can1357/oh-my-pi/issues/4786#issuecomment-5056055342
Claude-Session: https://claude.ai/code/session_014rh4JyWFkxgMhgFaEf8VBY
- An eval-worktree cherry-pick swept 16 packages/*/node_modules symlinks into the index; 'node_modules/' with a trailing slash only matches directories, so symlinked installs bypassed the ignore. Dropped the slash and removed the tracked links.
Aligns the shim's runtime safeParse/__validator with the wire/tool-call
path, so legacy draft-07 documents (tuple items) accept the same values
validateToolArguments does. Adds a regression test.
- Added `tui.scrollbackRebuild` configuration with interactive startup/controller wiring to apply `setScrollbackRebuild`.
- Exposed prewalk session state in `SegmentContext` and rendered a dedicated prewalk segment/icon in the status line.
- Added divergence-aware TUI full-paint logic that enables scrollback erase-and-replay rebuilds for non-multiplexer divergence cases.
- Updated rendering and streaming tests to verify rebuild behavior (`3J`) and eliminate stale marker expectations under drift scenarios.
- Implemented a unified benchmark normalization layer to handle metrics and artifacts across harbor, edit, and snapcompact benchmarks.
- Integrated the TypeScript edit benchmark directly into the manager, migrating logic and deprecating the standalone package.
- Updated the API and database schema to support standardized benchmark configurations, metrics, and trace-based reporting.
- Enhanced the UI to visualize comparative performance metrics, including pass rate, cost, and latency deltas for benchmark runs.
- Replaced markdown-style headings with concise `¶user:`, `¶think:`, `¶ai:`, and `¶call:` scope markers.
- Updated the serializer to merge consecutive messages or blocks of the same type under a shared prefix.
- Updated the documentation prompt to reflect the new compact formatting and scope rules.
- Added regression tests verifying scope merging and correct formatting of tool calls and intents.