Files
oh-my-pi/packages/coding-agent
roboomp 7fa0eeb43d fix(agent): break shake auto-continue loop on token-metric divergence
The shake-strategy post-shake threshold check was reading
#estimatePendingPromptTokens([]) while #checkCompaction triggered on
calculateContextTokens(assistantMessage.usage). The local estimator
ignored block.thinkingSignature payloads (OpenAI Responses encrypted
reasoning items, Anthropic signed thinking blocks, etc.), so on a
thinking-heavy session the estimate sat ~0.9–2× below provider-reported
usage. Once the two straddled the threshold, the #2119 dead-loop guard
never fired, shake reported 'handled', and #scheduleAutoContinuePrompt
re-injected the auto-continue developer prompt every turn — 53 injections
in a real 25-minute repro session before an external timeout.

Thread the trigger's provider-anchored contextTokens through
#runAutoCompaction → #runAutoShake for the threshold and incomplete
paths, then evaluate residual pressure as triggerContextTokens −
result.tokensFreed with an 80% recovery-band hysteresis. Re-checking
against the raw threshold (even on the corrected metric) would still let
shake reclaim a trickle of the previous turn's elidable blocks and land
just under the line every turn; the band closes that oscillation.

As defense in depth, estimateTokens() now charges thinkingSignature and
redactedThinking.data alongside the visible thinking text so every
other site that uses the estimator (idle compaction, pre-prompt check,
status line) tracks provider usage on replay.

New regression test pins the contract; existing dispatch test bumped
its mocked tokensFreed so its happy-path scenario lands inside the new
recovery band.

Fixes #2275
2026-06-10 21:18:46 +00:00
..
2026-05-09 03:11:18 +02:00

@oh-my-pi/pi-coding-agent

Core implementation package for the omp coding agent in the oh-my-pi monorepo.

For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:

Package-specific references:

Memory backends

The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):

  • off (default) — no memory subsystem runs.
  • local — existing rollout-summarisation pipeline; writes memory_summary.md and consolidated artifacts under the agent dir.
  • hindsight — talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposes retain, recall, and reflect.

Hindsight quickstart

  1. Run a Hindsight server (Cloud or docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest).
  2. Set memory.backend = "hindsight" and hindsight.apiUrl = "http://localhost:8888" (or your Cloud URL).
  3. Optional environment overrides (env wins over settings):
    • HINDSIGHT_API_URL, HINDSIGHT_API_TOKEN — connection
    • HINDSIGHT_BANK_ID, HINDSIGHT_DYNAMIC_BANK_ID, HINDSIGHT_AGENT_NAME — bank addressing
    • HINDSIGHT_AUTO_RECALL, HINDSIGHT_AUTO_RETAIN, HINDSIGHT_RETAIN_MODE — lifecycle
    • HINDSIGHT_RECALL_BUDGET, HINDSIGHT_RECALL_MAX_TOKENS — recall sizing
    • HINDSIGHT_BANK_MISSION, HINDSIGHT_DEBUG

Switching backends mid-session is honoured on the next system-prompt rebuild and the next /memory slash command. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch.