753c86ea72
A session that crossed a provider boundary compacted 90 times in three days without ever succeeding: every attempt asked the summarizer to read the whole re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a 1M cap), and every rejection was retried ten times. Three independent defects: 1. `generateSummary` serialized the entire span into one prompt with no budget check. It now plans windows that fit the summarizer's context and folds them with the update prompt that iterative compaction already uses, so a stranded boundary is recovered instead of rejected. A provider that rejects a window the catalog said would fit (claude-sonnet-4-5 advertises 1M but is beta-gated to 200k on OAuth credentials) halves what was actually sent and re-plans, because only the rejection knows the real cap. 2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp appends to its own errors classified a deterministic 400 as a transient 503. Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN. 3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30 identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so an input that does not fit never fits; both layers now fail fast to the next candidate. The boundary scan that decides which compaction entry a model can actually read is extracted as `findReadableCompactionIndex`, since the fold and `prepareCompaction` both need it. Verified by replaying the session that failed: 7,096 messages summarize in 3 calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15 calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
37 lines
1.7 KiB
TypeScript
37 lines
1.7 KiB
TypeScript
import { describe, expect, it } from "bun:test";
|
|
import * as AIError from "@oh-my-pi/pi-ai/error";
|
|
|
|
/**
|
|
* The transient classifier matches bare HTTP status codes in error text. Those
|
|
* digits must be a token of their own: omp appends its own
|
|
* `raw-http-request=<...>/<random-id>.json` pointer to provider errors, and a
|
|
* random id containing `503` used to make a hard 400 look retryable — which
|
|
* turned a deterministic oversized-prompt rejection into ten identical retries.
|
|
*/
|
|
describe("transient status classification", () => {
|
|
const overflowWithArtifactPointer =
|
|
'Summarization failed: 400 {"type":"error","error":{"type":"invalid_request_error",' +
|
|
'"message":"prompt is too long: 3030000 tokens > 1000000 maximum"}}\n' +
|
|
"raw-http-request=/home/u/.omp/logs/http-400-requests/1787022540720-3o503gxo48bvb.json";
|
|
|
|
it("does not call a 400 transient because an artifact id embeds a status code", () => {
|
|
const id = AIError.classify(new Error(overflowWithArtifactPointer), "anthropic-messages");
|
|
expect(AIError.is(id, AIError.Flag.ContextOverflow)).toBe(true);
|
|
expect(AIError.is(id, AIError.Flag.Transient)).toBe(false);
|
|
});
|
|
|
|
it("still classifies a real gateway status as transient", () => {
|
|
for (const text of ["503 Service Unavailable", "upstream returned 502", "HTTP 429 from provider"]) {
|
|
const id = AIError.classify(new Error(text), "anthropic-messages");
|
|
expect(AIError.is(id, AIError.Flag.Transient)).toBe(true);
|
|
}
|
|
});
|
|
|
|
it("does not treat status digits inside an identifier as transient", () => {
|
|
for (const text of ["model gpt-500x rejected the request", "request req500502 failed validation"]) {
|
|
const id = AIError.classify(new Error(text), "anthropic-messages");
|
|
expect(AIError.is(id, AIError.Flag.Transient)).toBe(false);
|
|
}
|
|
});
|
|
});
|