7fa0eeb43d
The shake-strategy post-shake threshold check was reading #estimatePendingPromptTokens([]) while #checkCompaction triggered on calculateContextTokens(assistantMessage.usage). The local estimator ignored block.thinkingSignature payloads (OpenAI Responses encrypted reasoning items, Anthropic signed thinking blocks, etc.), so on a thinking-heavy session the estimate sat ~0.9–2× below provider-reported usage. Once the two straddled the threshold, the #2119 dead-loop guard never fired, shake reported 'handled', and #scheduleAutoContinuePrompt re-injected the auto-continue developer prompt every turn — 53 injections in a real 25-minute repro session before an external timeout. Thread the trigger's provider-anchored contextTokens through #runAutoCompaction → #runAutoShake for the threshold and incomplete paths, then evaluate residual pressure as triggerContextTokens − result.tokensFreed with an 80% recovery-band hysteresis. Re-checking against the raw threshold (even on the corrected metric) would still let shake reclaim a trickle of the previous turn's elidable blocks and land just under the line every turn; the band closes that oscillation. As defense in depth, estimateTokens() now charges thinkingSignature and redactedThinking.data alongside the visible thinking text so every other site that uses the estimator (idle compaction, pre-prompt check, status line) tracks provider usage on replay. New regression test pins the contract; existing dispatch test bumped its mocked tokensFreed so its happy-path scenario lands inside the new recovery band. Fixes #2275
@oh-my-pi/pi-coding-agent
Core implementation package for the omp coding agent in the oh-my-pi monorepo.
For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:
Package-specific references:
- CHANGELOG
- MCP configuration guide
- MCP runtime lifecycle
- MCP server/tool authoring
- DEVELOPMENT
- RenderMermaid guide
Memory backends
The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):
off(default) — no memory subsystem runs.local— existing rollout-summarisation pipeline; writesmemory_summary.mdand consolidated artifacts under the agent dir.hindsight— talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposesretain,recall, andreflect.
Hindsight quickstart
- Run a Hindsight server (Cloud or
docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest). - Set
memory.backend = "hindsight"andhindsight.apiUrl = "http://localhost:8888"(or your Cloud URL). - Optional environment overrides (env wins over settings):
HINDSIGHT_API_URL,HINDSIGHT_API_TOKEN— connectionHINDSIGHT_BANK_ID,HINDSIGHT_DYNAMIC_BANK_ID,HINDSIGHT_AGENT_NAME— bank addressingHINDSIGHT_AUTO_RECALL,HINDSIGHT_AUTO_RETAIN,HINDSIGHT_RETAIN_MODE— lifecycleHINDSIGHT_RECALL_BUDGET,HINDSIGHT_RECALL_MAX_TOKENS— recall sizingHINDSIGHT_BANK_MISSION,HINDSIGHT_DEBUG
Switching backends mid-session is honoured on the next system-prompt rebuild and the next /memory slash command. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch.