Files
oh-my-pi/packages/coding-agent
cuipengfei b506c8ee65 fix(advisor): split Session update into per-message user messages to keep prompt cache growing
The advisor sends its whole Session update as a single ever-growing user
message. Provider prompt caches are prefix-based: a single user message whose
text keeps growing invalidates the entire message on every turn, so cache_read
stays pinned at the instructions/tools boundary (observed 14491 tokens in
production, 11066 in tests) instead of growing with the session.

Split the update into multiple user messages — one per source message —
delivered via a single Agent.prompt(AgentMessage[]) call, so the provider
caches each appended message incrementally. Verified end-to-end: cache_read
grows 0 -> 11126 -> 11457 -> 11583 across turns with the split, versus pinned
11066 on the old single-message behavior.

- delta-split.ts: pure renderAdvisorDeltaChunks using chunked
  formatSessionHistoryMarkdown (shared toolResultIndex/consumedToolCallIds/
  watchedRoleState) so toolCall/result pairing and role collapsing stay
  byte-identical to the old single-block render (equivalence-tested).
- session-history-format.ts: add HistoryFormatOptions.watchedRoleState so
  chunked renders collapse consecutive same-role messages exactly like the
  single-block render.
- runtime.ts: #prepareBatch does a single dedup+render pass; #drain delivers
  agent.prompt(preparedMessages) (array), falling back to the string.
- Keep field-selective fingerprint (candidate 1) + wip-marker-at-tail
  (candidate 3) as complementary wins.

Tests: advisor suite 209 pass / 0 fail; type check clean; lint clean.
Affected subsets (342 tests) green; full suite hits WSL EMFILE fd limit.
2026-08-04 22:36:41 +08:00
..
2026-08-04 01:19:36 +02:00
2026-08-04 05:53:35 +02:00

@oh-my-pi/pi-coding-agent

Core implementation package for the omp coding agent in the oh-my-pi monorepo.

For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:

Package-specific references:

Memory backends

The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):

  • off (default) — no memory subsystem runs.
  • local — existing rollout-summarisation pipeline; writes memory_summary.md and consolidated artifacts under the agent dir.
  • hindsight — talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposes retain, recall, and reflect.

Hindsight quickstart

  1. Run a Hindsight server (Cloud or docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest).
  2. Set memory.backend = "hindsight" and hindsight.apiUrl = "http://localhost:8888" (or your Cloud URL).
  3. Optional environment overrides (env wins over settings):
    • HINDSIGHT_API_URL, HINDSIGHT_API_TOKEN — connection
    • HINDSIGHT_BANK_ID, HINDSIGHT_DYNAMIC_BANK_ID, HINDSIGHT_AGENT_NAME — bank addressing
    • HINDSIGHT_AUTO_RECALL, HINDSIGHT_AUTO_RETAIN, HINDSIGHT_RETAIN_MODE — lifecycle
    • HINDSIGHT_RECALL_BUDGET, HINDSIGHT_RECALL_MAX_TOKENS — recall sizing
    • HINDSIGHT_BANK_MISSION, HINDSIGHT_DEBUG

Switching backends mid-session immediately replaces the live backend, memory tools, listeners, and system-prompt context. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch; afterward, memory.backend is the sole runtime selector.