Mirrors the composer's raw-sequence fallback (editor.ts:1466) so
\x1b[13;2~ triggers summarize-and-switch instead of being dropped as
shift+f3, completing the parity requested in issue #8821.
resolveWorkerSpawnCmd no longer pins the host-entry worker cwd to the
install dir; the subprocess inherits the agent cwd (or the package root
under the bun-test fallback). Update the chdir rationale accordingly.
The matcher compares selector ids case-insensitively, but the lock's
bundled-catalog lookup was exact-case: Anthropic/Claude-Opus-5 exact-
matched OpenRouter's flat id while getBundledModel("Anthropic", ...)
missed, silently re-enabling the aggregator shadow the lock exists to
prevent. Scan the named provider's bundled ids case-insensitively.
- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
organizeImports sorts the type GeneratedProvider specifier ahead of the named imports, and the formatter collapses the preferred-match arrow chain to one line. Both are biome safe fixes; no behavior change. Resolves the two biome errors that were leaving the 8833 branch's check gate red.
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.
Three independent defects:
1. `generateSummary` serialized the entire span into one prompt with no budget
check. It now plans windows that fit the summarizer's context and folds them
with the update prompt that iterative compaction already uses, so a stranded
boundary is recovered instead of rejected. A provider that rejects a window
the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
beta-gated to 200k on OAuth credentials) halves what was actually sent and
re-plans, because only the rejection knows the real cap.
2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
appends to its own errors classified a deterministic 400 as a transient 503.
Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.
3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
an input that does not fit never fits; both layers now fail fast to the next
candidate.
The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.
Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
The composer accepts three encodings for Shift+Enter (kitty CSI-u, the
legacy \x1b[13;2~ form, and a bare LF from the iTerm2 mapping e.g. Claude
Code's /terminal-setup). The /tree selector only handled the kitty form and
silently routed a bare LF into the plain-Enter branch, so summarize-and-
switch never fired for those terminals.
Mirror the composer: fall through a bare LF to summarize-and-switch while
plain CR (or the decoded Enter key) still does a plain switch.
Fixes#8821