Apply the snapcompact frame byte-budget cap even when the active model has no known context window, avoiding 80-frame archives on custom vision models.
Bounded persisted snapcompact image archives by base64 byte size so large sessions stop re-sending multi-megabyte frame walls on every provider request.
Auto snapcompact now falls back to context-full summaries when rendered frame payloads exceed the byte budget, and legacy oversized archives omit over-budget frames during LLM context rebuilds.
Fixes#3792
The "skips snapcompact entirely" case asserted .rejects.toThrow() on
session.compact(), but after the maxFrames<1 skip the manual /compact path
falls through to the LLM summarizer whose outcome is provider/network
dependent (resolves when a summary lands, rejects only without
credentials/network). In a sandbox it resolves and blows the 5s default
timeout. Tolerate either outcome and pin only the deterministic skip
contract (snapcompact.compact not invoked + the kept-history notice), with
a generous explicit timeout. Addresses chatgpt-codex P2 review on #3249.
chatgpt-codex third-pass review on #3249: the 4k SUMMARY_TEXT_RESERVE
in the cap math undersized the actual textHead+textTail cost a frame-
bearing archive carries (the projection separately bills
'countTokens(summary + textHead + textTail)'). At ~120k headroom on
Anthropic 11on16-bw, the cap picked maxFrames=23, but
'23 * 5024 + 2 * 13916 chars (≈7k tokens) + 2k summary template ≈ 124.5k'
still exceeded the same 120k headroom — the cap chose a value the
projection then immediately rejected, re-opening the warning loop.
#computeSnapcompactMaxFrames now resolves the live snapcompact shape
(same call the auto/manual paths pass to snapcompact.compact) and sizes
the cap reserve from 'geometry(shape).capacity':
textEdgeTokens = ceil(2 * capacity * 1.15 / 4) // 1.15 absorbs
// tokenizer drift
capReserve = textEdgeTokens + 2000 // + summary template
For the default per-provider winners that resolves to ~10k (Anthropic
Sonnet), ~14k (Opus 4.7), ~16k (Gemini 2.x), and ~10k (OpenAI) — all
larger than the prior fixed 4k. Skip decision stays separate
(baseTokens >= totalBudget), so positive sub-reserve headroom still
runs snapcompact's text-only path.
Test 1 retuned to baseline kept-recent ≈ 100k tokens with a strengthened
assertion verifying the FULL projection invariant (frames + worst-case
text edges + summary template + base ≤ budget). Confirmed test fails
against the previous 4k-reserve helper by exactly the reviewer's
predicted margin (174,271 vs 170,000 budget = 4,271 token overshoot).
chatgpt-codex second-pass review on #3249: the previous helper folded
the 4k SUMMARY_TEXT_RESERVE into both the maxFrames cap math AND the
skip decision (return 0 when frameBudget < 0). That made any residual
headroom below 4k fall negative and force the LLM-summarizer fallback,
even though a text-only snapcompact archive (the 'text.length <= 2 *
edgeCap' short-circuit in planArchive) typically costs only a few
hundred tokens of summary lead-in and would have fit cleanly.
The two reserves now serve their own jobs:
- Skip iff 'baseTokens >= totalBudget' (kept-recent + non-message
already eats the entire window − reserve envelope). No reserve
fudge here; positive residual is always worth attempting.
- Cap reserve (4k) is applied ONLY to the maxFrames calculation so
the projection still passes once frames land. When the frame budget
goes negative under that reserve but residual headroom is positive,
the helper now returns maxFrames=1 instead of 0 so snapcompact's
frame-less planArchive branch can still produce a valid archive.
Updated regression test to pin the new contract directly: kept-recent
tuned for 1500 tokens of headroom (well below the 4k cap reserve), the
old helper returned 0 and skipped to the LLM summarizer, the new helper
invokes snapcompact with maxFrames=1.
chatgpt-codex review on #3249: the helper returned 0 when frameBudget
< FRAME_TOKEN_ESTIMATE, causing the caller to skip snapcompact entirely.
But snapcompact.planArchive has a 'text.length <= 2 * edgeCap' short-
circuit that produces a valid frames:[] archive when the discarded
history is small enough — and the projection charges 0 for that. Hard
return-0 blocked that opportunity, forcing the LLM summarizer fallback
in offline/no-credential sessions where the text-only path would have
landed cleanly.
#computeSnapcompactMaxFrames now distinguishes two near-full cases:
- frameBudget < 0 → return 0 (kept-recent already exhausted budget;
no text-only summary can fit either) → caller still skips outright.
- 0 ≤ frameBudget < FRAME_TOKEN_ESTIMATE → return 1 → snapcompact runs
and picks the frame-less planArchive branch automatically for small
discarded histories; the projection guard rejects any actual
frame-bearing archive that overflows.
Added regression test pinning maxFrames=1 (not 0) in the near-full
window case.
Snapcompact's bundled MAX_FRAMES_DEFAULT (80) × FRAME_TOKEN_ESTIMATE (5024)
≈ 402k tokens worth of frames. AgentSession was calling snapcompact.compact()
with no maxFrames override, so the post-render projection inside #runAuto
Compaction / compact() always overflowed the budget on any sub-1M-token
window (Claude Sonnet 4.5's 200k = 170k usable, the 80-frame projection
alone clears that 2.4×), looping the 'snapcompact could not bring the
context under the limit — using an LLM summary instead' warning on every
threshold tick.
AgentSession.#computeSnapcompactMaxFrames now sizes the frame cap from
the resolved budget — (window − reserve − non-message − kept-recent −
summary-text reserve) / FRAME_TOKEN_ESTIMATE, clamped to MAX_FRAMES_DEFAULT
— and threads it into snapcompact.compact() in both the auto-compaction
and manual /compact paths. When the kept-recent slice already exceeds the
budget, snapcompact is skipped outright instead of running just to be
rejected: the projection guard remains as a defensive check.
Fixes#3247