- Bound PAX sparse record memory overhead by caching sparse markers and specific keys.
- Update system prompt phrasing and tests for tool inventory and date displays.
- Added Google provider thinking configuration parameters and force-reasoning-off controls.
- Implemented MCP SSE stream resumption using Last-Event-ID and `SSEResumeError`.
- Added support for TAR old-GNU sparse extension blocks, path length checks, and archive entry overrides.
- Restricted external thinking support to specific models and added semver fallback parsing.
- Added the `--external-thinking` CLI flag alongside model capability checks to gate external thinking tool availability.
- Updated Anthropic and Google transports to honor `forceReasoningOff` for native thinking-off controls.
- Renamed the `thoughts` property and parameter to `notes` across think fixtures, tools, and tests.
- Updated system prompt instructions and test suites to verify transport-specific thinking and tool activation.
- Refactored and condensed numerous system prompts, agent instructions, and tool documentation files across packages.
- Streamlined workflow rules, formatting constraints, and execution guidelines for improved clarity and brevity.
- Updated discovery rules, recommendation criteria, and syntax standards in prompt templates.
Versionless Fable/Mythos aliases (bundled claude-fable-latest) never parse a numeric version, so the parser path missed them and they fell to the 1568px default. Restored the /claude.*(fable|mythos)/i rule ahead of the generic Claude rule, mapping to the shared high-res tier.
Added regressions for anthropic/claude-fable-latest and its ~ prefixed alias.
Replaced the local Opus version regex with the shared catalog Anthropic identity and semantic-version parser. This recognizes both kind-first and version-first model IDs, provider proxy prefixes, and multi-digit minor versions while preserving the 4.7 floor.
Added regressions for anthropic--claude-4.8-opus and claude-opus-4-10.
MODEL_VARIANTS gated the 1932px frame tier on /claude-?opus-?4[.-][7-9]/i,
so claude-opus-5 fell through to the generic /claude/i entry and rendered at
the 1568px default reserved for older lines that downscale. The 1932px tier
tracks the Anthropic 4,784 visual-token cap — a family-wide billing property
opus-5 shares with opus-4-8 — so the version bound was stale rather than a
per-version eval gap.
Widened the pattern to cover Opus 5 and later; 4.0-4.6 still keep the safe
1568px default. Documented why 1932 specifically (largest square not
downscaled under the 4,784-patch cap, kept below the 2000px per-image limit).
Added regression assertions for opus-5/opus-6 (high-res) and opus-4-6
(default).
Fixes#8256
- Added a new SQuAD-based context-compression benchmark script for evaluating recall conditions.
- Added model and shape command-line arguments along with pricing and shape configurations.
- Updated default model variants and added an unknown billing family to shape resolution.
- Updated shape resolution and model resolution tests to verify the new defaults.
- Collected distinct rendered frame widths for resume summaries.
- Covered mixed HQ/LQ archives with the existing foveation regression.
- Documented the fix in the snapcompact changelog.
Fixes#6712
- Re-compaction unfolds the prior archive's kept source verbatim, so archives written before includeThinking existed kept replaying reasoning to Claude (reasoning_extraction) even after the serializer fix; the prior text is now scrubbed when includeThinking is false, healing poisoned sessions at their next compaction.
Both compaction serializers reproduced prior assistant reasoning as text
bound for a Claude target, tripping Anthropic's reasoning_extraction
refusal and wedging Fable 5 sessions:
- context-full: serializeConversation rendered thinking verbatim inside
<thinking> tags via the anthropic dialect renderer. Now drops thinking
blocks when the summary target dialect is anthropic; other dialects
(e.g. Harmony) keep native reasoning.
- snapcompact: emitted ¶think sections baked into replayed archive
frames. Added an includeThinking serialize option (default true) and
wired the agent session to disable it for Anthropic-dialect models.
Fixes#6093
A branch whose last entry is a snapcompact CompactionEntry billed past the
compaction threshold (FRAME_TOKEN_ESTIMATE x frames) dead-ended on every
resume: prepareCompaction returns undefined (nothing after the entry to
summarize), and the #4786 elide/image rescue tiers only inspect
"message"/"custom_message" entries, so a type:"compaction" tail escaped both
and the "Compaction freed too little context" warning re-fired forever.
Add a dedicated first rescue tier that rebuilds the SAME archive locally (no
LLM, no network) by re-running snapcompact.compact() over the entry's
carried-forward source text at a maxFrames derived from the trigger
threshold's recovery band instead of the window-fit budget: planArchive
truncates the oldest chars to fit, so the rebuilt entry genuinely shrinks.
Persisting through appendCompaction lets the write-time
superseded-compaction elision drop the stale frame payload from the JSONL,
and the pass skips the misleading no-progress warning.
Fixes the loop reported in
https://github.com/can1357/oh-my-pi/issues/4786#issuecomment-5056055342
Claude-Session: https://claude.ai/code/session_014rh4JyWFkxgMhgFaEf8VBY
- An eval-worktree cherry-pick swept 16 packages/*/node_modules symlinks into the index; 'node_modules/' with a trailing slash only matches directories, so symlinked installs bypassed the ignore. Dropped the slash and removed the tracked links.
Aligns the shim's runtime safeParse/__validator with the wire/tool-call
path, so legacy draft-07 documents (tuple items) accept the same values
validateToolArguments does. Adds a regression test.