b506c8ee65
The advisor sends its whole Session update as a single ever-growing user message. Provider prompt caches are prefix-based: a single user message whose text keeps growing invalidates the entire message on every turn, so cache_read stays pinned at the instructions/tools boundary (observed 14491 tokens in production, 11066 in tests) instead of growing with the session. Split the update into multiple user messages — one per source message — delivered via a single Agent.prompt(AgentMessage[]) call, so the provider caches each appended message incrementally. Verified end-to-end: cache_read grows 0 -> 11126 -> 11457 -> 11583 across turns with the split, versus pinned 11066 on the old single-message behavior. - delta-split.ts: pure renderAdvisorDeltaChunks using chunked formatSessionHistoryMarkdown (shared toolResultIndex/consumedToolCallIds/ watchedRoleState) so toolCall/result pairing and role collapsing stay byte-identical to the old single-block render (equivalence-tested). - session-history-format.ts: add HistoryFormatOptions.watchedRoleState so chunked renders collapse consecutive same-role messages exactly like the single-block render. - runtime.ts: #prepareBatch does a single dedup+render pass; #drain delivers agent.prompt(preparedMessages) (array), falling back to the string. - Keep field-selective fingerprint (candidate 1) + wip-marker-at-tail (candidate 3) as complementary wins. Tests: advisor suite 209 pass / 0 fail; type check clean; lint clean. Affected subsets (342 tests) green; full suite hits WSL EMFILE fd limit.
@oh-my-pi/pi-coding-agent
Core implementation package for the omp coding agent in the oh-my-pi monorepo.
For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:
Package-specific references:
Memory backends
The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):
off(default) — no memory subsystem runs.local— existing rollout-summarisation pipeline; writesmemory_summary.mdand consolidated artifacts under the agent dir.hindsight— talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposesretain,recall, andreflect.
Hindsight quickstart
- Run a Hindsight server (Cloud or
docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest). - Set
memory.backend = "hindsight"andhindsight.apiUrl = "http://localhost:8888"(or your Cloud URL). - Optional environment overrides (env wins over settings):
HINDSIGHT_API_URL,HINDSIGHT_API_TOKEN— connectionHINDSIGHT_BANK_ID,HINDSIGHT_DYNAMIC_BANK_ID,HINDSIGHT_AGENT_NAME— bank addressingHINDSIGHT_AUTO_RECALL,HINDSIGHT_AUTO_RETAIN,HINDSIGHT_RETAIN_MODE— lifecycleHINDSIGHT_RECALL_BUDGET,HINDSIGHT_RECALL_MAX_TOKENS— recall sizingHINDSIGHT_BANK_MISSION,HINDSIGHT_DEBUG
Switching backends mid-session immediately replaces the live backend, memory tools, listeners, and system-prompt context. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch; afterward, memory.backend is the sole runtime selector.