The advisor system prompt told the watcher model "at most one advise per update" and "NEVER send the same advice twice", but nothing enforced either rule. Issue #3520 captured a session where the advisor emitted 309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker"> injections in the primary transcript and destabilizing the watched agent after the task was already complete. New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and: - Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every "Stop.", "*Stop*", "STOP!" variant keys to the same canonical form. - Drops a small allowlist of content-free self-talk filler (stop, done, complete, no issue continue, lgtm, nothing to add, no further input, carry on, ...) - silence is the correct expression of "no concerns". - Dedupes by exact normalized text across the session, FIFO-bounded at 4096 entries. - Rate-limits to one accepted advise per advisor model prompt cycle. The runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(), so the new batch starts with a fresh budget. Suppressed calls don't consume the budget - a noise call never displaces a real concern. Reset on advisor reset (compaction, session switch, /new) so a re-primed reviewer can re-raise old concerns against the rewritten transcript. Suppression is invisible to the advisor model: AdviseTool still returns "Recorded." for a dropped call. Surfacing "suppressed" risks the model rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to bypass the dedupe. Fixes #3520
@oh-my-pi/pi-coding-agent
Core implementation package for the omp coding agent in the oh-my-pi monorepo.
For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:
Package-specific references:
Memory backends
The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):
off(default) — no memory subsystem runs.local— existing rollout-summarisation pipeline; writesmemory_summary.mdand consolidated artifacts under the agent dir.hindsight— talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposesretain,recall, andreflect.
Hindsight quickstart
- Run a Hindsight server (Cloud or
docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest). - Set
memory.backend = "hindsight"andhindsight.apiUrl = "http://localhost:8888"(or your Cloud URL). - Optional environment overrides (env wins over settings):
HINDSIGHT_API_URL,HINDSIGHT_API_TOKEN— connectionHINDSIGHT_BANK_ID,HINDSIGHT_DYNAMIC_BANK_ID,HINDSIGHT_AGENT_NAME— bank addressingHINDSIGHT_AUTO_RECALL,HINDSIGHT_AUTO_RETAIN,HINDSIGHT_RETAIN_MODE— lifecycleHINDSIGHT_RECALL_BUDGET,HINDSIGHT_RECALL_MAX_TOKENS— recall sizingHINDSIGHT_BANK_MISSION,HINDSIGHT_DEBUG
Switching backends mid-session is honoured on the next system-prompt rebuild and the next /memory slash command. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch.