Files
oh-my-pi/packages/coding-agent
PaleRoses 753c86ea72 fix(compaction): bound summarization input and stop retrying overflow
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.

Three independent defects:

1. `generateSummary` serialized the entire span into one prompt with no budget
   check. It now plans windows that fit the summarizer's context and folds them
   with the update prompt that iterative compaction already uses, so a stranded
   boundary is recovered instead of rejected. A provider that rejects a window
   the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
   beta-gated to 200k on OAuth credentials) halves what was actually sent and
   re-plans, because only the rejection knows the real cap.

2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
   the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
   appends to its own errors classified a deterministic 400 as a transient 503.
   Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.

3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
   identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
   an input that does not fit never fits; both layers now fail fast to the next
   candidate.

The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.

Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
2026-08-18 12:25:13 -07:00
..
2026-08-17 23:55:09 +03:00
2026-08-17 23:55:09 +03:00
2026-08-17 22:29:25 +03:00

@oh-my-pi/pi-coding-agent

Core implementation package for the omp coding agent in the oh-my-pi monorepo.

For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:

Package-specific references:

Memory backends

The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):

  • off (default) — no memory subsystem runs.
  • local — existing rollout-summarisation pipeline; writes memory_summary.md and consolidated artifacts under the agent dir.
  • hindsight — talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposes retain, recall, and reflect.

Hindsight quickstart

  1. Run a Hindsight server (Cloud or docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest).
  2. Set memory.backend = "hindsight" and hindsight.apiUrl = "http://localhost:8888" (or your Cloud URL).
  3. Optional environment overrides (env wins over settings):
    • HINDSIGHT_API_URL, HINDSIGHT_API_TOKEN — connection
    • HINDSIGHT_BANK_ID, HINDSIGHT_DYNAMIC_BANK_ID, HINDSIGHT_AGENT_NAME — bank addressing
    • HINDSIGHT_AUTO_RECALL, HINDSIGHT_AUTO_RETAIN, HINDSIGHT_RETAIN_MODE — lifecycle
    • HINDSIGHT_RECALL_BUDGET, HINDSIGHT_RECALL_MAX_TOKENS — recall sizing
    • HINDSIGHT_BANK_MISSION, HINDSIGHT_DEBUG

Switching backends mid-session immediately replaces the live backend, memory tools, listeners, and system-prompt context. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch; afterward, memory.backend is the sole runtime selector.