feat(coding-agent): added bench command and updated default compaction shapes

- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.
This commit is contained in:
can1357
2026-06-12 08:00:03 +02:00
parent a094b794bc
commit 300c1ada30
8 changed files with 492 additions and 11 deletions
+1 -1
View File
@@ -131,7 +131,7 @@ The automatic paths are intentionally different:
`compaction.strategy: "snapcompact"` replaces the LLM summarization call with a local, deterministic archival pass (`compact` from `@oh-my-pi/snapcompact`):
- The discarded history is serialized, whitespace-collapsed, and printed onto provider-aware square PNG frames using bundled public-domain pixel fonts. Anthropic-family and unknown APIs use repeated black `8x8` cells, Google uses repeated sentence-colored `8x8` cells, and OpenAI uses `8x13` glyphs on a 16px pitch (`8on16-bw`) with `detail: "original"`. The `snapcompact.shape` setting (default `auto`) forces one of the research-eval variants instead: square grids (`8x8r`/`8x8u`/`6x6u`/`5x8` × sentence-hue/black ink) or the per-model eval winners (`6x12-dim`, `8x13-bw`, `8on16-bw`, and the two-column word-wrapped `doc-8on16-bw`/`-sent`/`-sent-dim`, where `dim` prints stopwords in gray). A forced variant keeps its geometry but is re-priced for the target provider's image billing. The same setting governs inline system-prompt/tool-result imaging (`snapcompact.systemPrompt`, `snapcompact.toolResults`).
- The discarded history is serialized, whitespace-collapsed, and printed onto provider-aware square PNG frames using bundled public-domain pixel fonts. Anthropic-family and unknown APIs use repeated black `8x8` cells (`8x8r-bw`), Google uses two word-wrapped columns of `8x13` glyphs with sentence-hue ink and dimmed stopwords (`doc-8on16-sent-dim`), and OpenAI uses `8x13` glyphs on a 16px pitch (`8on16-bw`) with `detail: "original"`. The `snapcompact.shape` setting (default `auto`) forces one of the research-eval variants instead: square grids (`8x8r`/`8x8u`/`6x6u`/`5x8` × sentence-hue/black ink) or the per-model eval winners (`6x12-dim`, `8x13-bw`, `8on16-bw`, and the two-column word-wrapped `doc-8on16-bw`/`-sent`/`-sent-dim`, where `dim` prints stopwords in gray). A forced variant keeps its geometry but is re-priced for the target provider's image billing. The same setting governs inline system-prompt/tool-result imaging (`snapcompact.systemPrompt`, `snapcompact.toolResults`).
- Serialization keeps the archive conversation-dense: tool results are truncated head+tail (default 2,000 chars at a 0.6 head ratio), tool-call argument values are capped per value (500) and per call (2,000), and tool output is printed in dim gray ink so conversation reads louder than tool noise. All budgets and the dimming are configurable via `SerializeOptions` (`toolResultMaxChars`, `toolArgMaxChars`, `toolCallMaxChars`, `truncateHeadRatio`, `dimToolResults`).
- Frames persist under `CompactionEntry.preserveData.snapcompact` and are re-attached to the `compactionSummary` message as image blocks on every context rebuild; the entry's `summary` is a deterministic reading guide (grid geometry, role tags, truncation notes) plus the usual file-operation lists.
- Later compactions carry earlier frames forward. Beyond an 8-frame budget the archive fades from the middle out: the earliest frame (session head — the original request, or the filmed summary of older history) is pinned, and the oldest *unpinned* frames are evicted, so head and tail both survive. If the previous compaction was text-based, its summary is printed at the head of the frame archive as `[Summary of earlier history]`.
+4
View File
@@ -1,8 +1,12 @@
# Changelog
## [Unreleased]
### Added
- Added `bench` CLI command to benchmark model selectors with a shared prompt and report time-to-first-token and throughput
- Added `omp bench` options `--runs`, `--max-tokens`, `--prompt`, and `--json` to control run count, response length, prompt text, and machine-readable output
- Added benchmark failure reporting that shows per-run errors, flags failed runs in the summary, and exits with a non-zero status when any model benchmark fails
- Added `snapcompact.shape` setting to pick the frame variant snapcompact prints text with — `auto` (each provider's eval winner) or any research-eval variant: the square grids (`8x8r`/`8x8u`/`6x6u`/`5x8` × sentence-hue/black ink) and the per-model winners (`6x12-dim`, `8x13-bw`, `8on16-bw`, `doc-8on16-bw`, `doc-8on16-sent`, `doc-8on16-sent-dim`); applies to both the compaction archive and inline system-prompt/tool-result imaging
### Changed
@@ -16,6 +16,7 @@ export const commands: CommandEntry[] = [
{ name: "auth-broker", load: () => import("./commands/auth-broker").then(m => m.default) },
{ name: "auth-gateway", load: () => import("./commands/auth-gateway").then(m => m.default) },
{ name: "agents", load: () => import("./commands/agents").then(m => m.default) },
{ name: "bench", load: () => import("./commands/bench").then(m => m.default) },
{ name: "commit", load: () => import("./commands/commit").then(m => m.default) },
{ name: "completions", load: () => import("./commands/completions").then(m => m.default) },
{ name: "__complete", load: () => import("./commands/complete").then(m => m.default) },
+428
View File
@@ -0,0 +1,428 @@
import type { ResolvedThinkingLevel } from "@oh-my-pi/pi-agent-core";
import type {
Api,
ApiKeyResolver,
AssistantMessage,
AssistantMessageEvent,
AssistantMessageEventStream,
Context,
Effort,
Model,
SimpleStreamOptions,
} from "@oh-my-pi/pi-ai";
import { streamSimple } from "@oh-my-pi/pi-ai";
import type { CanonicalModelVariant } from "@oh-my-pi/pi-catalog/identity";
import { replaceTabs, truncateToWidth } from "@oh-my-pi/pi-tui";
import { formatDuration, getProjectDir } from "@oh-my-pi/pi-utils";
import chalk from "chalk";
import type { ApiKeyResolverModel } from "../config/api-key-resolver";
import { type CanonicalModelQueryOptions, ModelRegistry } from "../config/model-registry";
import { formatModelString, getModelMatchPreferences, resolveCliModel } from "../config/model-resolver";
import { Settings } from "../config/settings";
import benchPrompt from "../prompts/bench.md" with { type: "text" };
import { discoverAuthStorage } from "../sdk";
import { resolveThinkingLevelForModel, shouldDisableReasoning, toReasoningEffort } from "../thinking";
const DEFAULT_RUNS = 1;
const DEFAULT_MAX_TOKENS = 512;
const ERROR_WIDTH = 110;
const BENCH_PROMPT = benchPrompt.trim();
export interface BenchCommandArgs {
models: string[];
flags: {
runs?: number;
maxTokens?: number;
prompt?: string;
json?: boolean;
};
}
export interface BenchModelRegistry {
getAll(): Model<Api>[];
getApiKey(model: Model<Api>, sessionId?: string): Promise<string | undefined>;
resolver(model: ApiKeyResolverModel, sessionId?: string): ApiKeyResolver;
resolveCanonicalModel?(canonicalId: string, options?: CanonicalModelQueryOptions): Model<Api> | undefined;
getCanonicalVariants?(canonicalId: string, options?: CanonicalModelQueryOptions): CanonicalModelVariant[];
getCanonicalId?(model: Model<Api>): string | undefined;
}
export interface BenchRuntime {
modelRegistry: BenchModelRegistry;
settings?: Settings;
close?: () => void;
}
export interface BenchRunSuccess {
ok: true;
ttftMs: number;
durationMs: number;
outputTokens: number;
/** Generation throughput measured over the post-first-token window. */
tokensPerSecond: number;
}
export interface BenchRunFailure {
ok: false;
error: string;
}
export type BenchRunResult = BenchRunSuccess | BenchRunFailure;
export interface BenchAverages {
ttftMs: number;
durationMs: number;
outputTokens: number;
tokensPerSecond: number;
}
export interface BenchModelReport {
/** Selector as the user typed it (e.g. "opus" or "gemini-3.5:low"). */
selector: string;
/** Resolved `provider/id`. */
model: string;
/** Explicit thinking level from a `:level` selector suffix; undefined = provider default. */
thinking?: ResolvedThinkingLevel;
results: BenchRunResult[];
/** Averages over successful runs; null when every run failed. */
average: BenchAverages | null;
}
export interface BenchSummary {
runs: number;
maxTokens: number;
models: BenchModelReport[];
failures: number;
}
type BenchStreamSimple = (
model: Model<Api>,
context: Context,
options?: SimpleStreamOptions,
) => AssistantMessageEventStream;
export interface BenchDependencies {
createRuntime?: () => Promise<BenchRuntime>;
randomSessionId?: () => string;
writeStdout?: (text: string) => void;
writeStderr?: (text: string) => void;
setExitCode?: (code: number) => void;
streamSimple?: BenchStreamSimple;
now?: () => number;
stdoutIsTTY?: boolean;
}
function getErrorMessage(error: unknown): string {
if (error instanceof Error && error.message) return error.message;
return String(error);
}
function normalizePositiveInteger(name: string, value: number | undefined, fallback: number): number {
if (value === undefined) return fallback;
if (!Number.isInteger(value) || value <= 0) {
throw new Error(`Expected --${name} to be a positive integer, got ${value}`);
}
return value;
}
function isFirstTokenEvent(event: AssistantMessageEvent): boolean {
switch (event.type) {
case "text_delta":
case "thinking_delta":
case "toolcall_delta":
return event.delta.length > 0;
case "text_end":
case "thinking_end":
return event.content.length > 0;
default:
return false;
}
}
/**
* Tokens/s over the generation window (duration minus TTFT) so queue/prefill
* latency does not dilute throughput. Falls back to total duration when the
* response arrived as a single chunk (TTFT ~ duration).
*/
function computeTokensPerSecond(outputTokens: number, durationMs: number, ttftMs: number): number {
const decodeMs = durationMs - ttftMs;
const windowMs = decodeMs > 0 ? decodeMs : durationMs;
return windowMs > 0 ? (outputTokens * 1000) / windowMs : 0;
}
interface BenchRequestOptions {
apiKey: ApiKeyResolver;
sessionId: string;
prompt: string;
maxTokens: number;
/** Explicit effort from a `:level` selector suffix; absent = provider default. */
reasoning?: Effort;
/** Only set for an explicit `:off` suffix — some endpoints reject disablement. */
disableReasoning?: boolean;
}
async function runBenchRequest(
model: Model<Api>,
options: BenchRequestOptions,
streamFn: BenchStreamSimple,
now: () => number,
): Promise<BenchRunResult> {
const startedAt = now();
let firstTokenAt: number | undefined;
try {
const context: Context = {
messages: [{ role: "user", content: options.prompt, timestamp: Date.now(), attribution: "user" }],
};
const stream = streamFn(model, context, {
apiKey: options.apiKey,
sessionId: options.sessionId,
maxTokens:
Number.isFinite(model.maxTokens) && model.maxTokens > 0
? Math.min(options.maxTokens, model.maxTokens)
: options.maxTokens,
reasoning: options.reasoning,
disableReasoning: options.disableReasoning,
});
let message: AssistantMessage | undefined;
for await (const event of stream) {
if (firstTokenAt === undefined && isFirstTokenEvent(event)) {
firstTokenAt = now();
}
if (event.type === "error") {
return { ok: false, error: event.error.errorMessage ?? "request failed" };
}
if (event.type === "done") {
message = event.message;
}
}
message ??= await stream.result();
if (message.stopReason === "error" || message.errorMessage) {
return { ok: false, error: message.errorMessage ?? "request failed" };
}
const rawDuration = message.duration ?? now() - startedAt;
const durationMs = Number.isFinite(rawDuration) && rawDuration > 0 ? rawDuration : 0;
const rawTtft = message.ttft ?? (firstTokenAt === undefined ? durationMs : firstTokenAt - startedAt);
const ttftMs = Number.isFinite(rawTtft) && rawTtft > 0 ? rawTtft : 0;
const outputTokens = Number.isFinite(message.usage.output) && message.usage.output > 0 ? message.usage.output : 0;
return {
ok: true,
ttftMs,
durationMs,
outputTokens,
tokensPerSecond: computeTokensPerSecond(outputTokens, durationMs, ttftMs),
};
} catch (error) {
return { ok: false, error: getErrorMessage(error) };
}
}
function buildModelReport(
selector: string,
model: Model<Api>,
thinking: ResolvedThinkingLevel | undefined,
results: BenchRunResult[],
): BenchModelReport {
const successes = results.filter((result): result is BenchRunSuccess => result.ok);
const average =
successes.length === 0
? null
: {
ttftMs: successes.reduce((sum, r) => sum + r.ttftMs, 0) / successes.length,
durationMs: successes.reduce((sum, r) => sum + r.durationMs, 0) / successes.length,
outputTokens: successes.reduce((sum, r) => sum + r.outputTokens, 0) / successes.length,
tokensPerSecond: successes.reduce((sum, r) => sum + r.tokensPerSecond, 0) / successes.length,
};
return { selector, model: formatModelString(model), thinking, results, average };
}
function formatMs(ms: number): string {
return formatDuration(Math.max(0, Math.round(ms)));
}
function formatRunLine(result: BenchRunResult, index: number, total: number): string {
const prefix = chalk.dim(`run ${index + 1}/${total}`);
if (result.ok) {
return ` ${chalk.green("✓")} ${prefix} ${chalk.dim("TTFT")} ${formatMs(result.ttftMs)} ${chalk.dim("TPS")} ${result.tokensPerSecond.toFixed(1)}/s ${chalk.dim("tokens")} ${result.outputTokens} ${chalk.dim("total")} ${formatMs(result.durationMs)}`;
}
return ` ${chalk.red("✗")} ${prefix} ${chalk.red(truncateToWidth(replaceTabs(result.error).replace(/\r?\n/g, " "), ERROR_WIDTH))}`;
}
export function formatBenchTable(summary: BenchSummary): string {
const ranked = [...summary.models].sort((a, b) => {
if (a.average === null && b.average === null) return 0;
if (a.average === null) return 1;
if (b.average === null) return -1;
return b.average.tokensPerSecond - a.average.tokensPerSecond;
});
const rows = ranked.map(report => ({
model: report.model,
ttft: report.average ? formatMs(report.average.ttftMs) : "-",
tps: report.average ? `${report.average.tokensPerSecond.toFixed(1)}/s` : "-",
tokens: report.average ? String(Math.round(report.average.outputTokens)) : "-",
total: report.average ? formatMs(report.average.durationMs) : "-",
failed: report.results.filter(result => !result.ok).length,
}));
const headers = { model: "model", ttft: "TTFT", tps: "TPS", tokens: "tokens", total: "total" } as const;
const width = (key: keyof typeof headers): number =>
Math.max(headers[key].length, ...rows.map(row => row[key].length));
const lines = [
[
headers.model.padEnd(width("model")),
headers.ttft.padEnd(width("ttft")),
headers.tps.padEnd(width("tps")),
headers.tokens.padEnd(width("tokens")),
headers.total.padEnd(width("total")),
]
.join(" ")
.trimEnd(),
];
for (const row of rows) {
const failedSuffix = row.failed > 0 ? ` ${chalk.red(`(${row.failed} failed)`)}` : "";
lines.push(
[
row.model.padEnd(width("model")),
row.ttft.padEnd(width("ttft")),
row.tps.padEnd(width("tps")),
row.tokens.padEnd(width("tokens")),
row.total.padEnd(width("total")),
]
.join(" ")
.trimEnd() + failedSuffix,
);
}
return `${lines.map((line, index) => (index === 0 ? chalk.dim(line) : line)).join("\n")}\n`;
}
async function createDefaultRuntime(): Promise<BenchRuntime> {
const authStorage = await discoverAuthStorage();
try {
const settings = await Settings.init({ cwd: getProjectDir() });
const modelRegistry = new ModelRegistry(authStorage);
return {
modelRegistry,
settings,
close: () => authStorage.close(),
};
} catch (error) {
authStorage.close();
throw error;
}
}
interface BenchTarget {
selector: string;
model: Model<Api>;
thinking: ResolvedThinkingLevel | undefined;
}
function resolveBenchModels(
selectors: string[],
modelRegistry: BenchModelRegistry,
settings: Settings | undefined,
writeStderr: (text: string) => void,
): BenchTarget[] {
const preferences = getModelMatchPreferences(settings);
const resolved: BenchTarget[] = [];
const errors: string[] = [];
for (const selector of selectors) {
const result = resolveCliModel({ cliModel: selector, modelRegistry, preferences });
if (result.error) {
errors.push(`${selector}: ${result.error}`);
continue;
}
if (!result.model) {
errors.push(`${selector}: model not found`);
continue;
}
if (result.warning) writeStderr(`${chalk.yellow(`Warning: ${result.warning}`)}\n`);
resolved.push({
selector,
model: result.model,
thinking: resolveThinkingLevelForModel(result.model, result.thinkingLevel),
});
}
if (errors.length > 0) {
throw new Error(`Could not resolve ${errors.length === 1 ? "model" : "models"}:\n${errors.join("\n")}`);
}
return resolved;
}
export async function runBenchCommand(command: BenchCommandArgs, deps: BenchDependencies = {}): Promise<BenchSummary> {
const runs = normalizePositiveInteger("runs", command.flags.runs, DEFAULT_RUNS);
const maxTokens = normalizePositiveInteger("max-tokens", command.flags.maxTokens, DEFAULT_MAX_TOKENS);
const prompt = command.flags.prompt?.trim() || BENCH_PROMPT;
const json = command.flags.json === true;
const randomSessionId = deps.randomSessionId ?? (() => Bun.randomUUIDv7());
const writeStdout = deps.writeStdout ?? ((text: string) => process.stdout.write(text));
const writeStderr = deps.writeStderr ?? ((text: string) => process.stderr.write(text));
const setExitCode =
deps.setExitCode ??
((code: number) => {
process.exitCode = code;
});
const streamFn = deps.streamSimple ?? streamSimple;
const now = deps.now ?? (() => performance.now());
const interactive = deps.stdoutIsTTY ?? process.stdout.isTTY === true;
if (command.models.length === 0) {
throw new Error("Pass at least one model selector, e.g. `omp bench opus gpt-5.2`");
}
const runtime = await (deps.createRuntime ?? createDefaultRuntime)();
try {
const targets = resolveBenchModels(command.models, runtime.modelRegistry, runtime.settings, writeStderr);
const reports: BenchModelReport[] = [];
for (const { selector, model, thinking } of targets) {
if (!json) {
const resolvedNote = selector === formatModelString(model) ? "" : chalk.dim(` (${selector})`);
writeStdout(`${chalk.bold(formatModelString(model))}${resolvedNote}\n`);
}
const results: BenchRunResult[] = [];
for (let index = 0; index < runs; index++) {
const sessionId = randomSessionId();
const initialKey = await runtime.modelRegistry.getApiKey(model, sessionId);
if (!initialKey) {
const failure: BenchRunFailure = {
ok: false,
error: `No credentials for provider "${model.provider}". Run \`omp\` and use /login, or set the provider API key.`,
};
results.push(failure);
if (!json) writeStdout(`${formatRunLine(failure, index, runs)}\n`);
break; // remaining runs would fail identically
}
if (!json && interactive) {
writeStdout(chalk.dim(` … run ${index + 1}/${runs} streaming`));
}
const result = await runBenchRequest(
model,
{
apiKey: runtime.modelRegistry.resolver(model, sessionId),
sessionId,
prompt,
maxTokens,
reasoning: toReasoningEffort(thinking),
disableReasoning: shouldDisableReasoning(thinking) ? true : undefined,
},
streamFn,
now,
);
results.push(result);
if (!json) {
if (interactive) writeStdout("\r\x1b[2K");
writeStdout(`${formatRunLine(result, index, runs)}\n`);
}
}
reports.push(buildModelReport(selector, model, thinking, results));
}
const failures = reports.reduce((sum, report) => sum + report.results.filter(result => !result.ok).length, 0);
const summary: BenchSummary = { runs, maxTokens, models: reports, failures };
if (json) {
writeStdout(`${JSON.stringify(summary, null, 2)}\n`);
} else if (reports.length > 1 || runs > 1) {
writeStdout(`\n${formatBenchTable(summary)}`);
}
if (failures > 0) setExitCode(1);
return summary;
} finally {
runtime.close?.();
}
}
@@ -0,0 +1,42 @@
import { Args, Command, Flags } from "@oh-my-pi/pi-utils/cli";
import { runBenchCommand } from "../cli/bench-cli";
export default class Bench extends Command {
static description =
"Benchmark models with the same prompt: time-to-first-token and generation throughput (tokens/s)";
static args = {
models: Args.string({
description: "Model selectors (provider/model or fuzzy id, e.g. opus)",
required: true,
multiple: true,
}),
};
static flags = {
runs: Flags.integer({ description: "Requests per model (results are averaged)", default: 1 }),
"max-tokens": Flags.integer({ description: "Max output tokens per request", default: 512 }),
prompt: Flags.string({ description: "Custom prompt text (default: bundled bench prompt)" }),
json: Flags.boolean({ description: "Output JSON" }),
};
static examples = [
"# Compare two models\n omp bench anthropic/claude-opus-4-5 openai/gpt-5.2",
"# Fuzzy selectors work\n omp bench opus sonnet",
"# Average over 3 runs each\n omp bench opus gpt-5.2 --runs 3",
"# Machine-readable output\n omp bench opus --json",
];
async run(): Promise<void> {
const { args, flags } = await this.parse(Bench);
await runBenchCommand({
models: args.models ?? [],
flags: {
runs: flags.runs,
maxTokens: flags["max-tokens"],
prompt: flags.prompt,
json: flags.json,
},
});
}
}
@@ -1624,7 +1624,8 @@ export const SETTINGS_SCHEMA = {
{
value: "auto",
label: "Auto",
description: "Provider's eval winner: 8x8r-bw on Anthropic, 8x8r-sent on Gemini, 8on16-bw on OpenAI.",
description:
"Provider's eval winner: 8x8r-bw on Anthropic, doc-8on16-sent-dim on Gemini, 8on16-bw on OpenAI.",
},
{
value: "8x8r-bw",
@@ -1698,7 +1699,7 @@ export const SETTINGS_SCHEMA = {
value: "doc-8on16-sent-dim",
label: "Doc 8on16, sentence hues + dimmed stopwords",
description:
"Two-column doc layout, sentence-hue ink, function words dimmed gray. Gemini/Kimi eval winner.",
"Two-column doc layout, sentence-hue ink, function words dimmed gray. Gemini eval winner (mono .900 vs .853 for the repeated grid) and the Gemini auto default.",
},
],
},
+1
View File
@@ -12,6 +12,7 @@
### Changed
- **Changed the OpenAI default shape from `6x6u-sent` to `8on16-bw`.** A production-regime mono eval (gpt-5.5, the full 800k-char SQuAD flow in one request, n=50) scored the old dense default f1 .602 vs .851 for `8on16-bw` rendered by the production pipeline, at near-equal total cost (the dense cells burned the frame savings on reasoning tokens); chunked exp14 had already scored `8on16-bw` .906. `SHAPES.openaiDense` is renamed to `SHAPES.openai`
- **Changed the Google default shape from `8x8r-sent` to `doc-8on16-sent-dim`.** Production-rendered mono eval on gemini-3.5-flash (400k chars, one request, n=25): f1 .900 vs .853 for the repeated grid at lower cost, agreeing with the chunked round-2 winner. The Anthropic default stayed `8x8r-bw` — it beat the chunked research winners in the same production regime on both claude-fable (.877 vs .840 for `6x12-dim`) and opus (.833 vs .793 for `8x13-bw`)
- `normalize()` now keeps line structure: whitespace runs containing a line break collapse to `NEWLINE_GLYPH` (U+2588 FULL BLOCK, drawn by the native renderer as a pitch-black cell one character wide) instead of a plain space; leading/trailing breaks are trimmed, and the frame-reading prompt explains the marker
- `normalize()` now skips characters the fonts cannot render instead of printing `?` blanks: whole ANSI escape sequences are stripped, and bare control characters, zero-width format characters (ZWSP, BOM, directional marks), combining marks, and lone surrogates are dropped without occupying a cell; `?` remains the fallback for unsupported graphic characters only
+12 -8
View File
@@ -13,12 +13,14 @@
* printed twice with the copy on a pale highlight band. Read at F1 parity
* with raw text at ~2x lower cost; the colored variants drew refusals at
* scale, the repeated plain shape did not.
* - **Google** (`8x8r-sent`): same repeated grid with six-hue sentence
* coloring (0.90 F1 at ~2.9x lower cost on gemini-3.5-flash).
* - **OpenAI** (`6x6u-sent`): OpenAI bills a flat ~2.9k tokens per image, so
* image count is the only cost lever — unscii-8 Lanczos-stretched to 6x6
* cells packs the most readable chars per frame. Frames request
* `detail: "original"`; the default `auto` downscale destroys 6px glyphs.
* - **Google** (`doc-8on16-sent-dim`): two word-wrapped newspaper columns of
* 8x13 glyphs on a 16px pitch, sentence-hue ink, stopwords dimmed (0.90 F1
* on gemini-3.5-flash vs 0.85 for the repeated grid, at lower cost; same
* shape won the chunked round-2 evals).
* - **OpenAI** (`8on16-bw`): 8x13 glyphs on a patch-aligned 16px pitch, black
* ink (gpt-5.5 mono F1 0.851 vs 0.602 for the previous dense `6x6u-sent`).
* OpenAI bills a flat ~2.9k tokens per image; frames request
* `detail: "original"` since the default `auto` downscale destroys glyphs.
* - **Unknown providers** default to the Anthropic shape (most
* refusal-robust). Gateways that resize images (e.g. OpenRouter normalizes
* visual payloads to a fixed token budget) defeat any shape — optical
@@ -204,8 +206,10 @@ function priceShape(base: ShapeGeometry, family: BillingFamily): Shape {
export const SHAPES = {
/** `8x8r-bw`: unscii square, black ink, lines doubled on highlight bands. */
anthropic: priceShape(SHAPE_VARIANTS["8x8r-bw"], "anthropic"),
/** `8x8r-sent`: the repeated grid with sentence-hue ink. */
google: priceShape(SHAPE_VARIANTS["8x8r-sent"], "google"),
/** `doc-8on16-sent-dim`: two word-wrapped columns, sentence hues, dimmed
* stopwords. Production mono eval on gemini-3.5-flash: f1 .900 vs .853
* for the repeated grid, at lower cost; also the chunked round-2 winner. */
google: priceShape(SHAPE_VARIANTS["doc-8on16-sent-dim"], "google"),
/** `8on16-bw`: 8x13 X.org glyphs on a 16px pitch, black ink. Mono eval on
* gpt-5.5 (200k-token single request, n=50): f1 .851 vs .602 for the
* previous `6x6u-sent` default at near-equal total cost; chunked exp14