753c86ea72
A session that crossed a provider boundary compacted 90 times in three days without ever succeeding: every attempt asked the summarizer to read the whole re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a 1M cap), and every rejection was retried ten times. Three independent defects: 1. `generateSummary` serialized the entire span into one prompt with no budget check. It now plans windows that fit the summarizer's context and folds them with the update prompt that iterative compaction already uses, so a stranded boundary is recovered instead of rejected. A provider that rejects a window the catalog said would fit (claude-sonnet-4-5 advertises 1M but is beta-gated to 200k on OAuth credentials) halves what was actually sent and re-plans, because only the rejection knows the real cap. 2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp appends to its own errors classified a deterministic 400 as a transient 503. Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN. 3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30 identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so an input that does not fit never fits; both layers now fail fast to the next candidate. The boundary scan that decides which compaction entry a model can actually read is extracted as `findReadableCompactionIndex`, since the fold and `prepareCompaction` both need it. Verified by replaying the session that failed: 7,096 messages summarize in 3 calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15 calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
@oh-my-pi/pi-coding-agent
Core implementation package for the omp coding agent in the oh-my-pi monorepo.
For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:
Package-specific references:
Memory backends
The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):
off(default) — no memory subsystem runs.local— existing rollout-summarisation pipeline; writesmemory_summary.mdand consolidated artifacts under the agent dir.hindsight— talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposesretain,recall, andreflect.
Hindsight quickstart
- Run a Hindsight server (Cloud or
docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest). - Set
memory.backend = "hindsight"andhindsight.apiUrl = "http://localhost:8888"(or your Cloud URL). - Optional environment overrides (env wins over settings):
HINDSIGHT_API_URL,HINDSIGHT_API_TOKEN— connectionHINDSIGHT_BANK_ID,HINDSIGHT_DYNAMIC_BANK_ID,HINDSIGHT_AGENT_NAME— bank addressingHINDSIGHT_AUTO_RECALL,HINDSIGHT_AUTO_RETAIN,HINDSIGHT_RETAIN_MODE— lifecycleHINDSIGHT_RECALL_BUDGET,HINDSIGHT_RECALL_MAX_TOKENS— recall sizingHINDSIGHT_BANK_MISSION,HINDSIGHT_DEBUG
Switching backends mid-session immediately replaces the live backend, memory tools, listeners, and system-prompt context. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch; afterward, memory.backend is the sole runtime selector.