- Centralized JSON parsing and stream processing logic by moving utilities from `packages/ai` to the shared `@oh-my-pi/pi-utils` package.
- Standardized import paths for JSON parsing and streaming across the agent, ai, and coding-agent packages.
- Refactored SSE stream handling to use consolidated `parseStreamingJson` logic and introduced robust error recovery for malformed container-shaped tail events.
- Cleaned up legacy bundled registry references and updated related module exports and tests to reflect the new utility structure.
The post-maintenance headroom guard returned residualTokens < triggerContextTokens
as a secondary check after the recovery-band test. When stale/tool-output pruning
already drove the trigger (postMaintenanceContextTokens) below the band before this
pass, that strict-less comparison made a residual which merely held the line at/under
the band report a false no-progress, suppressing a valid auto-continue and emitting a
spurious warning even though the next turn could no longer re-trip threshold compaction.
The recovery band sits strictly under the compaction threshold, so reaching it already
guarantees the next turn cannot re-trip. Make the band authoritative (residual <= band)
and drop the trigger argument. Adds a regression covering the sub-band trigger case.
(cherry picked from commit 31cf2dc7a30c4134f7a351d61065eca76543f154)
Addresses PR #3412 review (roboomp blocking + codex P2 + Copilot nits):
- The overflow/incomplete retry no longer reuses the COMPACTION_RECOVERY_BAND
hysteresis. Reusing it required residual context below 0.8×threshold, which
turned a recoverable overflow (e.g. 150k on a 200k window, under the ~170k
threshold) into a manual dead-end. Add #compactionCreatedRetryFit, which
measures the rebuilt prompt against the usable fit budget
(contextWindow - effectiveReserveTokens) and is evaluated AFTER the failed
assistant is dropped, so the just-failed turn is excluded.
- The threshold auto-continue keeps the stricter recovery-band check
(#compactionCreatedHeadroom) — that path is the snapcompact thrash guard.
- Split the no-progress warning so each path warns on its own signal.
- Alias errorIsFromBeforeCompaction as assistantPredatesCompaction in the
threshold-usage path for clarity; de-duplicate the rationale comment.
Adds two regression tests (recoverable overflow retries; non-fitting overflow
pauses + warns), mutation-verified against the band-vs-fit split.
(cherry picked from commit f8cf45b2740f04d8776a84521ba563b779d65b25)
prepareCompaction keeps the most-recent turn verbatim (findCutPoint never
cuts at tool results), so when that single turn already exceeds the
compaction threshold the rewritten context stays above threshold. The
context-full / snapcompact auto-compaction success tail scheduled the
agent-authored auto-continue (and the overflow/incomplete retry)
unconditionally, so the next agent_end re-entered #checkCompaction over
the same oversized tail and re-fired forever.
This is the residual loop left after #3247 capped snapcompact's own frame
projection: once #computeSnapcompactMaxFrames drops the frame cap below
one frame, snapcompact is skipped and the context-full summarizer path
still creates no headroom, so the loop persists on that path.
- #runAutoCompaction now gates the continuation and the overflow/
incomplete retry on a post-maintenance headroom check
(#compactionCreatedHeadroom), reusing the shake recovery-band
hysteresis from #2275 (consolidated as the shared
COMPACTION_RECOVERY_BAND). When a pass frees too little it pauses
automatic maintenance and emits a single warning instead of looping.
- The post-turn threshold check ignores an assistant's stale
pre-compaction usage so the scheduled auto-continue can't re-trip on
the kept assistant's old high token count.
Adds a regression test covering the no-headroom (pause + warn, no
continuation) and headroom (auto-continue, no warn) paths; both
assertions are mutation-verified against the guard.
(cherry picked from commit 6fcbbe2b3075827ee621acaee8efa2c586eedd3e)
Stopped snapcompact preflight failures from falling through to the provider-backed LLM summarizer and covered manual plus auto compaction paths.
Fixes#3599
- Changed `inlineToolDescriptors` from a boolean to a three-way enum (`auto` | `on` | `off`) to allow per-model defaults.
- Implemented `auto` logic which defaults to inlining descriptors specifically for Gemini models.
- Added a migration to automatically map existing boolean values to `on` or `off` to maintain backward compatibility.
- Suppressed main-UI relay for sibling broadcast legs when the main agent is a direct target.
- Added `suppressRelay` option to `IrcBus.send` to allow selective disabling of relay rendering.
- Ensured broadcast fan-outs avoid rendering the same message twice in the main transcript.
- Updated `normalizeGeneratedTitle` to reconcile model-generated titles against the user's input instead of forcing title-case.
- Added logic to restore distinctive proper-noun casing (e.g., `TinyVMM`) and flatten model-generated camelCase artifacts (e.g., `dAemon`) that do not appear in the user's message.
- Ensured model-cased proper nouns that are not in the source message (e.g., `GitHub`) are preserved.
Passed the settled assistant message into session_stop emission so refusal-as-error turns can be pruned from replay context without hiding their stop details from extension hooks.
Expanded the refusal regression test to assert the session_stop payload still exposes the refusal as last_assistant_message.
Fixes#3591
Scoped provider-refusal filtering to live replay so compaction and snapcompact summaries retain the refused turn while outbound provider context still drops the refusal.
Fixes#3592
Removed the early return after refusal pruning so the agent_end tail still reaches `#emitSessionStopEvent`, restoring `session_stop` extension hooks (block/continue/telemetry) for refusal-as-error stops.
Regression test wires an extensionRunner with a session_stop handler and asserts it fires for both the refusal turn and the following clean turn.
Fixes#3591
API-level refusals now stay visible as terminal errors without being sent back as assistant dialogue on the next provider request. Added core and coding-agent conversion coverage for Anthropic refusal metadata.
Fixes#3592
Fixed a non-deterministic gc-cli archive test: archive-me and keep-recent
shared ageDays:90, so their mtimes tied within a millisecond on fast CI and
the retainNewestGlobal:1 'keep newest' pick fell back to readdir order,
archiving the wrong session. Give keep-recent ageDays:60 (still cold-eligible,
unambiguously newer).
OMP injects `INTENT_FIELD` (`i`) into every tool's wire schema. The direct
model tool-call path strips it via `extractIntent` in agent-loop, but the
eval `tool.*` bridge forwards args verbatim, so strict-schema MCP servers
(Linear, anything with `additionalProperties:false` / Zod `.strict()`)
rejected every call with `-32602 unrecognized_keys: ["i"]`. The eval
bridge surfaced the rejection as `hasError: true` instead of throwing, so
batch callers reading the value as success silently mutated nothing.
Move the strip to the MCP boundary so it owns the contract regardless of
caller: `MCPTool.execute` / `DeferredMCPTool.execute` route params through
a new `prepareOutboundArgs` that runs `stripHarnessIntent` before
`omitUnusedOptionalArgs`. `stripHarnessIntent` leaves `i` in place when
the server's own `inputSchema.properties` declares it, so a server that
legitimately uses `i` as a parameter is unaffected.
Fixes#3575