diff --git a/docs/compaction.md b/docs/compaction.md index 54d589b25..767e51ce7 100644 --- a/docs/compaction.md +++ b/docs/compaction.md @@ -131,7 +131,7 @@ The automatic paths are intentionally different: `compaction.strategy: "snapcompact"` replaces the LLM summarization call with a local, deterministic archival pass (`compact` from `@oh-my-pi/snapcompact`): -- The discarded history is serialized, whitespace-collapsed, and printed onto model-aware PNG frames (frame width fixed per shape; frame height hugs the rows actually printed) using bundled public-domain pixel fonts. The shape resolves from the **model id** when the model line was measured — Claude reads X.org `6x12` glyphs with dimmed stopwords (`6x12-dim`), Gemini reads two word-wrapped columns of `8x13` glyphs with sentence-hue ink and dimmed stopwords (`doc-8on16-sent-dim`), GPT/Kimi/GLM read `8x13` glyphs on a 16px pitch (`8on16-bw`) — so a Claude routed through Vertex or OpenRouter keeps its Claude shape. Unmeasured models fall back to their wire API family (Anthropic-family/unknown → `6x12-dim`, Google → `doc-8on16-sent-dim`, OpenAI-compatible → `8on16-bw`); billing (token estimate, OpenAI's `detail: "original"` hint) always follows the API carrying the request. The `snapcompact.shape` setting (default `auto`) forces one of the research-eval variants instead: square grids (`8x8r`/`8x8u`/`6x6u`/`5x8` × sentence-hue/black ink) or the per-model eval winners (`6x12-dim`, `8x13-bw`, `8on16-bw`, and the two-column word-wrapped `doc-8on16-bw`/`-sent`/`-sent-dim`, where `dim` prints stopwords in gray). A forced variant keeps its geometry but is re-priced for the target provider's image billing. The same setting governs inline system-prompt/tool-result imaging (`snapcompact.systemPrompt`, `snapcompact.toolResults`). +- The discarded history is serialized, whitespace-collapsed, and printed onto model-aware PNG frames (frame width fixed per shape; frame height hugs the rows actually printed) using bundled public-domain pixel fonts. The shape — and frame size — resolve from the **model id** when the model line was measured: Claude reads X.org `6x12` glyphs with dimmed stopwords (`6x12-dim`; high-res lines — Opus 4.7+, Fable, Mythos — get 1932px frames under Anthropic's 4,784 visual-token cap, older lines stay at 1568px), Gemini reads two word-wrapped columns of `8x13` glyphs with sentence-hue ink and dimmed stopwords (`doc-8on16-sent-dim` at 2048px — Gemini 3.x bills a fixed 1,120-token budget per image at any pixel size), GPT/Kimi/GLM read `8x13` glyphs on a 16px pitch (`8on16-bw` at 1568px — patch billing is area-proportional, and kimi's processor downscales past 1792px). A Claude routed through Vertex or OpenRouter keeps its Claude shape. Unmeasured models fall back to their wire API family (Anthropic-family/unknown → `6x12-dim`, Google → `doc-8on16-sent-dim`, OpenAI-compatible → `8on16-bw`); billing (per-family patch/budget formulas, OpenAI's `detail: "original"` hint) always follows the API carrying the request, computed for the resolved frame size. The `snapcompact.shape` setting (default `auto`) forces one of the research-eval variants instead: square grids (`8x8r`/`8x8u`/`6x6u`/`5x8` × sentence-hue/black ink) or the per-model eval winners (`6x12-dim`, `8x13-bw`, `8on16-bw`, and the two-column word-wrapped `doc-8on16-bw`/`-sent`/`-sent-dim`, where `dim` prints stopwords in gray). A forced variant keeps its geometry but is re-priced for the target provider's image billing. The same setting governs inline system-prompt/tool-result imaging (`snapcompact.systemPrompt`, `snapcompact.toolResults`). - Serialization keeps the archive conversation-dense: tool results are truncated head+tail (default 2,000 chars at a 0.6 head ratio), tool-call argument values are capped per value (500) and per call (2,000), and tool output is printed in dim gray ink so conversation reads louder than tool noise. All budgets and the dimming are configurable via `SerializeOptions` (`toolResultMaxChars`, `toolArgMaxChars`, `toolCallMaxChars`, `truncateHeadRatio`, `dimToolResults`). - Frames persist under `CompactionEntry.preserveData.snapcompact` and are re-attached to the `compactionSummary` message as image blocks on every context rebuild; the entry's `summary` is a deterministic reading guide (grid geometry, role tags, truncation notes) plus the usual file-operation lists. - Later compactions carry earlier frames forward. The frame budget is provider-aware (`providerFrameBudget`): the per-provider image cap clamped to 8 (`MAX_FRAMES`) — OpenRouter hard-caps requests at 8 images and silently drops the excess, unknown providers get a safe floor of 5. Beyond the budget the archive fades from the middle out: the earliest frame (session head — the original request, or the filmed summary of older history) is pinned, and the oldest *unpinned* frames are evicted. Pages of the *current* compaction that no longer fit are never rendered or dropped — the newest unframed slice survives verbatim as a text tail on the summary (`Archive.textTail`, capped at two frame capacities with middle elision) and is folded back into frames by the next compaction. If the previous compaction was text-based, its summary is printed at the head of the frame archive as `[Summary of earlier history]`. diff --git a/packages/coding-agent/test/snapcompact-inline.test.ts b/packages/coding-agent/test/snapcompact-inline.test.ts index 963865ecd..201e88239 100644 --- a/packages/coding-agent/test/snapcompact-inline.test.ts +++ b/packages/coding-agent/test/snapcompact-inline.test.ts @@ -426,7 +426,7 @@ describe("estimateInlineSavings", () => { expect(estimate.visionCapable).toBe(true); expect(estimate.systemPrompt?.applied).toBe(true); expect(estimate.systemPrompt?.frames).toBe(2); - expect(estimate.systemPrompt?.imageTokens).toBe(2 * 3300); + expect(estimate.systemPrompt?.imageTokens).toBe(2 * snapcompact.SHAPES.anthropic.frameTokenEstimate); expect(estimate.systemPrompt?.savedTokens).toBe( estimate.systemPrompt!.textTokens - estimate.systemPrompt!.imageTokens, ); diff --git a/packages/snapcompact/CHANGELOG.md b/packages/snapcompact/CHANGELOG.md index da095ab62..7eae2f0ef 100644 --- a/packages/snapcompact/CHANGELOG.md +++ b/packages/snapcompact/CHANGELOG.md @@ -9,6 +9,7 @@ - Added the six research-eval winning frame variants to `SHAPE_VARIANTS`: `6x12-dim` (Claude fable), `8x13-bw` (Opus), `8on16-bw` (GPT grid runner-up), `doc-8on16-bw` (GPT), `doc-8on16-sent` (GLM), and `doc-8on16-sent-dim` (Gemini/Kimi), backed by new `Shape` fields `stretch` (disable Lanczos stretch: natural glyphs on a larger cell pitch), `columns` (two word-wrapped newspaper columns), `stopwordDim`, and the X.org `6x12`/`8x13` fonts - Added `dimStopwords()`, which prints high-frequency function words in dim ink via zero-width markers (skipping spans that are already dim), and `wrap()`, the greedy word-wrap used to typeset doc-layout pages; `geometry`/`render`/`renderMany`/`frames`/`compact` understand doc shapes (wrap once, paginate into `2 * rows`-line pages), and compaction frames persist `columns`/`stopwordDim` for mixed-shape detection - `resolveShape` now takes a `ShapeTarget` (`{ api, id }` — a pi-ai `Model` works as-is) and detects the ideal shape from the **model id**, not just the wire API: a Claude routed through Vertex or an OpenAI-compatible gateway keeps its Claude shape, with billing still priced by the API family actually carrying the request. `idealShapeVariant(modelId)` exposes the model-line table; unmeasured models fall back to the API family's winner +- `resolveShape` now also resolves an ideal **frame size** per model line, and billing estimates come from verified per-family formulas instead of flat 1568px constants: Anthropic bills 28px patches capped at 4,784 visual tokens (+5% margin), Gemini 3.x bills a fixed 1,120-token `media_resolution` budget per image at any pixel size, and OpenAI bills 32px patches × 1.2 under the 10,000-patch `detail: "original"` budget. High-res Claude lines (Opus 4.7+, Fable, Mythos — native 2576px-edge ingestion) get 1932px frames (same recall and cost, a third fewer frames); Gemini gets 2048px frames (+70% chars per frame at the same bill); GPT and Kimi stay at 1568px (area-proportional billing and a model-side 1792px processor cap, respectively). `idealShapeVariant` now returns an `IdealShape` (`{ variant, frameSize? }`) - Added per-provider image-count budgets: `PROVIDER_IMAGE_BUDGETS`, `DEFAULT_PROVIDER_IMAGE_BUDGET`, `providerImageBudget()`, and `providerFrameBudget()` (the image budget clamped to `MAX_FRAMES`). OpenRouter is capped at its measured hard limit of 8 images per request (excess images are silently dropped with no error); unknown providers get a safe floor of 5 - Added `Archive.textTail`: archive content past the frame budget is no longer dropped — `compact()` stops rendering at the budget and keeps the newest unframed slice as verbatim text on the summary (capped at two frame capacities with middle elision, counted into `truncatedChars` when elided). The tail persists in `preserveData` and is folded back into frames by the next compaction diff --git a/packages/snapcompact/research/bench_gemini.py b/packages/snapcompact/research/bench_gemini.py new file mode 100644 index 000000000..6c3f636e0 --- /dev/null +++ b/packages/snapcompact/research/bench_gemini.py @@ -0,0 +1,192 @@ +# /// script +# requires-python = ">=3.10" +# dependencies = ["pillow"] +# /// +"""Direct Gemini API bench: per-part media_resolution (Gemini 3 only knob). + +Same protocol as mono_prod.py (production frames via render_pages.ts, SQuAD +flow, seed 42), but calls generativelanguage.googleapis.com v1alpha directly +so we can set per-part `media_resolution` (e.g. MEDIA_RESOLUTION_ULTRA_HIGH = +2240 tokens/image), which OpenRouter does not forward. + + uv run --with pillow python bench_gemini.py --resolution MEDIA_RESOLUTION_ULTRA_HIGH \ + --shape-json '{...}' --name ultra-3072 --chars 400000 --questions 25 +""" + +import argparse +import base64 +import json +import subprocess +import sys +import time +import urllib.error +import urllib.request +from pathlib import Path + +HERE = Path(__file__).resolve().parent +sys.path.insert(0, str(HERE)) + +import squad # noqa: E402 +from final import cached # noqa: E402 +from providers import load_env_key # noqa: E402 +from run import CACHE, RESULTS, load_prompt, sha8 # noqa: E402 + +GEMINI_MODEL = "gemini-3.5-flash" +GEMINI_URL = f"https://generativelanguage.googleapis.com/v1alpha/models/{GEMINI_MODEL}:generateContent" +PRICE_IN, PRICE_OUT = 0.6, 4.0 # $/M, matches final.MODELS google/gemini-3.5-flash + + +def _post(body: dict, api_key: str, retries: int = 4) -> dict: + payload = json.dumps(body).encode() + req = urllib.request.Request( + GEMINI_URL, data=payload, + headers={"content-type": "application/json", "x-goog-api-key": api_key}, + ) + for attempt in range(retries + 1): + try: + with urllib.request.urlopen(req, timeout=600) as resp: + return json.loads(resp.read()) + except urllib.error.HTTPError as err: + detail = err.read().decode(errors="replace")[:500] + if err.code in (408, 429, 500, 502, 503) and attempt < retries: + wait = 2.0 * 2**attempt + print(f" HTTP {err.code}, retrying in {wait:.0f}s: {detail[:120]}") + time.sleep(wait) + continue + raise SystemExit(f"Gemini API error {err.code}: {detail}") from err + except (json.JSONDecodeError, TimeoutError, urllib.error.URLError) as err: + if attempt < retries: + wait = 2.0 * 2**attempt + print(f" bad response ({type(err).__name__}), retrying in {wait:.0f}s") + time.sleep(wait) + continue + raise + raise AssertionError("unreachable") + + +def gemini_complete(api_key: str, blocks: list[dict], resolution: str | None, max_tokens: int) -> dict: + """blocks: [{"text": str} | {"image_path": Path}]; returns {"text", "usage", "stop"}.""" + parts = [] + for b in blocks: + if "text" in b: + parts.append({"text": b["text"]}) + else: + part: dict = { + "inline_data": { + "mime_type": "image/png", + "data": base64.b64encode(Path(b["image_path"]).read_bytes()).decode(), + } + } + if resolution: + part["media_resolution"] = {"level": resolution} + parts.append(part) + body = { + "contents": [{"role": "user", "parts": parts}], + "generationConfig": {"maxOutputTokens": max_tokens}, + } + out = _post(body, api_key) + cand = (out.get("candidates") or [{}])[0] + text = "".join( + p.get("text", "") + for p in (cand.get("content") or {}).get("parts", []) + if not p.get("thought") + ) + u = out.get("usageMetadata", {}) + usage = { + "in": u.get("promptTokenCount", 0) - u.get("cachedContentTokenCount", 0), + "out": u.get("candidatesTokenCount", 0) + u.get("thoughtsTokenCount", 0), + "cache_w": 0, + "cache_r": u.get("cachedContentTokenCount", 0), + "reasoning": u.get("thoughtsTokenCount", 0), + } + stop = "max_tokens" if cand.get("finishReason") == "MAX_TOKENS" else (cand.get("finishReason") or "").lower() + return {"text": text, "usage": usage, "stop": stop} + + +def main() -> None: + ap = argparse.ArgumentParser() + ap.add_argument("--shape-json", required=True) + ap.add_argument("--name", required=True) + ap.add_argument("--resolution", default=None, + help="per-part media_resolution level, e.g. MEDIA_RESOLUTION_ULTRA_HIGH; omit for API default") + ap.add_argument("--chars", type=int, default=400_000) + ap.add_argument("--questions", type=int, default=25) + ap.add_argument("--qpb", type=int, default=5) + ap.add_argument("--seed", type=int, default=42) + ap.add_argument("--max-tokens", type=int, default=32768) + ap.add_argument("--env", default="~/.env") + ap.add_argument("--fresh", action="store_true") + args = ap.parse_args() + + api_key = load_env_key("GEMINI_API_KEY", args.env) + paras = squad.load_paragraphs(CACHE) + flow, offsets = squad.build_flow(paras, args.chars) + questions = squad.sample_chunk_questions(paras, offsets, 0, len(flow), args.questions, args.seed) + shape, label = json.loads(args.shape_json), args.name + cond = f"prod-{label}" + size = shape["frameSize"] + + frame_dir = CACHE / f"prod-frames-{label}-{sha8(flow, json.dumps(shape, sort_keys=True))}" + if not frame_dir.exists() or not any(frame_dir.iterdir()): + flow_file = CACHE / f"prod-flow-{sha8(flow)}.txt" + flow_file.write_text(flow) + subprocess.run( + ["bun", str(HERE / "render_pages.ts"), str(flow_file), json.dumps(shape), str(frame_dir)], + check=True, + ) + pngs = sorted(frame_dir.glob("page-*.png")) + cols = (size // shape["cellWidth"] - 3) // 2 if shape.get("columns") == 2 else size // shape["cellWidth"] + rows = size // shape["cellHeight"] // shape.get("lineRepeat", 1) + preamble = load_prompt("qa-image-multi.md").format(k=len(pngs), cols=cols, rows=rows) + if shape.get("columns") == 2: + preamble += ( + "\nNote: each image lays text out as two word-wrapped newspaper columns separated by a gutter; " + "read the left column top to bottom, then the right column." + ) + ctx_blocks = [{"text": preamble}, *({"image_path": str(p)} for p in pngs), {"text": "End of images."}] + + out_dir = RESULTS / f"mono-prod-gemini-direct-{label}" + out_dir.mkdir(parents=True, exist_ok=True) + answers, usages, stops = [], [], [] + for b in range(0, len(questions), args.qpb): + batch = questions[b : b + args.qpb] + q_block = "\n".join(f"{i + 1}. {q['q']}" for i, q in enumerate(batch)) + blocks = [*ctx_blocks, {"text": q_block}] + qa = cached( + f"gemini-direct-{GEMINI_MODEL}", "qa-mono-prod-direct", + {"blocks": [{k: str(v) for k, v in blk.items()} for blk in blocks], "resolution": args.resolution}, + lambda blk=blocks: gemini_complete(api_key, blk, args.resolution, args.max_tokens), + args.fresh, + ) + answers.extend(squad.parse_numbered(qa["text"], len(batch))) + usages.append(qa["usage"]) + stops.append(qa["stop"]) + rows_out = [ + { + "model": GEMINI_MODEL, "cond": cond, "pos_rel": q["pos_rel"], "q": q["q"], "answer": a, + "golds": q["golds"], "em": squad.exact_match(a, q["golds"]), "f1": squad.f1(a, q["golds"]), + "abstained": "unreadable" in a.lower(), + } + for q, a in zip(questions, answers) + ] + u = {k: sum(x[k] for x in usages) for k in ("in", "out", "cache_w", "cache_r", "reasoning")} + cost = (u["in"] + 0.1 * u["cache_r"]) / 1e6 * PRICE_IN + u["out"] / 1e6 * PRICE_OUT + summary = { + "cond": cond, "n": len(rows_out), "imgs": len(pngs), "resolution": args.resolution, + "em": sum(r["em"] for r in rows_out) / len(rows_out), + "f1": sum(r["f1"] for r in rows_out) / len(rows_out), + "abst": sum(r["abstained"] for r in rows_out), + "tok_in": u["in"], "tok_cached": u["cache_r"], "tok_out": u["out"], "reas": u["reasoning"], + "cost": cost, "stop": next((s for s in stops if s == "max_tokens"), stops[-1] if stops else ""), + } + (out_dir / "records.jsonl").write_text("\n".join(json.dumps(r) for r in rows_out)) + (out_dir / "summary.json").write_text(json.dumps([summary], indent=1)) + print( + f"{cond:<28} res={args.resolution or 'default'} imgs={summary['imgs']:>2} " + f"f1={summary['f1']:.3f} em={summary['em']:.3f} abst={summary['abst']} " + f"tok_in={summary['tok_in']} ${summary['cost']:.3f} stop={summary['stop']}" + ) + + +if __name__ == "__main__": + main() diff --git a/packages/snapcompact/research/bench_kimi.py b/packages/snapcompact/research/bench_kimi.py new file mode 100644 index 000000000..be14460ac --- /dev/null +++ b/packages/snapcompact/research/bench_kimi.py @@ -0,0 +1,197 @@ +# /// script +# requires-python = ">=3.10" +# dependencies = ["pillow"] +# /// +"""Kimi K2.6 chunked benchmark runner (KimiK26Bench scratch file). + +mono_prod protocol (same flow/questions/seed) but every request carries <=8 +frames: OpenRouter silently drops images after the first 8, and kimi itself +dilutes >8 genuine frames. Chunk plan = windows of 8 frames, overlap >=1, +evenly spread (reproduces the diag [(0,8),(6,14),(13,21)] plan for 21 frames +so the .973 anchor is a cache hit). Questions are routed to the chunk where +their answer position is most interior. + + uv run --with pillow python bench_kimi.py --shape-json '' --name