The advisor sends its whole Session update as a single ever-growing user
message. Provider prompt caches are prefix-based: a single user message whose
text keeps growing invalidates the entire message on every turn, so cache_read
stays pinned at the instructions/tools boundary (observed 14491 tokens in
production, 11066 in tests) instead of growing with the session.
Split the update into multiple user messages — one per source message —
delivered via a single Agent.prompt(AgentMessage[]) call, so the provider
caches each appended message incrementally. Verified end-to-end: cache_read
grows 0 -> 11126 -> 11457 -> 11583 across turns with the split, versus pinned
11066 on the old single-message behavior.
- delta-split.ts: pure renderAdvisorDeltaChunks using chunked
formatSessionHistoryMarkdown (shared toolResultIndex/consumedToolCallIds/
watchedRoleState) so toolCall/result pairing and role collapsing stay
byte-identical to the old single-block render (equivalence-tested).
- session-history-format.ts: add HistoryFormatOptions.watchedRoleState so
chunked renders collapse consecutive same-role messages exactly like the
single-block render.
- runtime.ts: #prepareBatch does a single dedup+render pass; #drain delivers
agent.prompt(preparedMessages) (array), falling back to the string.
- Keep field-selective fingerprint (candidate 1) + wip-marker-at-tail
(candidate 3) as complementary wins.
Tests: advisor suite 209 pass / 0 fail; type check clean; lint clean.
Affected subsets (342 tests) green; full suite hits WSL EMFILE fd limit.