Files
oh-my-pi/packages/ai/test/error-transient-status-boundary.test.ts
T
PaleRoses 753c86ea72 fix(compaction): bound summarization input and stop retrying overflow
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.

Three independent defects:

1. `generateSummary` serialized the entire span into one prompt with no budget
   check. It now plans windows that fit the summarizer's context and folds them
   with the update prompt that iterative compaction already uses, so a stranded
   boundary is recovered instead of rejected. A provider that rejects a window
   the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
   beta-gated to 200k on OAuth credentials) halves what was actually sent and
   re-plans, because only the rejection knows the real cap.

2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
   the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
   appends to its own errors classified a deterministic 400 as a transient 503.
   Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.

3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
   identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
   an input that does not fit never fits; both layers now fail fast to the next
   candidate.

The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.

Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
2026-08-18 12:25:13 -07:00

37 lines
1.7 KiB
TypeScript

import { describe, expect, it } from "bun:test";
import * as AIError from "@oh-my-pi/pi-ai/error";
/**
* The transient classifier matches bare HTTP status codes in error text. Those
* digits must be a token of their own: omp appends its own
* `raw-http-request=<...>/<random-id>.json` pointer to provider errors, and a
* random id containing `503` used to make a hard 400 look retryable — which
* turned a deterministic oversized-prompt rejection into ten identical retries.
*/
describe("transient status classification", () => {
const overflowWithArtifactPointer =
'Summarization failed: 400 {"type":"error","error":{"type":"invalid_request_error",' +
'"message":"prompt is too long: 3030000 tokens > 1000000 maximum"}}\n' +
"raw-http-request=/home/u/.omp/logs/http-400-requests/1787022540720-3o503gxo48bvb.json";
it("does not call a 400 transient because an artifact id embeds a status code", () => {
const id = AIError.classify(new Error(overflowWithArtifactPointer), "anthropic-messages");
expect(AIError.is(id, AIError.Flag.ContextOverflow)).toBe(true);
expect(AIError.is(id, AIError.Flag.Transient)).toBe(false);
});
it("still classifies a real gateway status as transient", () => {
for (const text of ["503 Service Unavailable", "upstream returned 502", "HTTP 429 from provider"]) {
const id = AIError.classify(new Error(text), "anthropic-messages");
expect(AIError.is(id, AIError.Flag.Transient)).toBe(true);
}
});
it("does not treat status digits inside an identifier as transient", () => {
for (const text of ["model gpt-500x rejected the request", "request req500502 failed validation"]) {
const id = AIError.classify(new Error(text), "anthropic-messages");
expect(AIError.is(id, AIError.Flag.Transient)).toBe(false);
}
});
});