- Added a check to promote the context model before triggering mid-turn compaction.
- Prevented unnecessary compaction during long tool loops that cross threshold limits by enabling window expansion instead.
- Implemented a monitoring system to identify Gemini model runaway behavior during thinking steps using consecutive header detection.
- Introduced a configurable tool-call reminder prompt to inject corrective context when reasoning stalls.
- Added session-level logic to automatically interrupt and prune stalled assistant turns from the conversation history.
- Provided comprehensive test coverage for the detection logic and stream interruption scenarios.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Implemented file deletion (`REM`) and movement (`MV`) operations within the hashline grammar and parser.
- Added filesystem support for executing delete and move commands while maintaining snapshot history migration.
- Integrated file operation detection and parsing logic into the coding agent and patcher components.
- Added comprehensive integration tests and documentation for the new file-level operation syntax.
The post-maintenance headroom guard returned residualTokens < triggerContextTokens
as a secondary check after the recovery-band test. When stale/tool-output pruning
already drove the trigger (postMaintenanceContextTokens) below the band before this
pass, that strict-less comparison made a residual which merely held the line at/under
the band report a false no-progress, suppressing a valid auto-continue and emitting a
spurious warning even though the next turn could no longer re-trip threshold compaction.
The recovery band sits strictly under the compaction threshold, so reaching it already
guarantees the next turn cannot re-trip. Make the band authoritative (residual <= band)
and drop the trigger argument. Adds a regression covering the sub-band trigger case.
(cherry picked from commit 31cf2dc7a30c4134f7a351d61065eca76543f154)
Addresses PR #3412 review (roboomp blocking + codex P2 + Copilot nits):
- The overflow/incomplete retry no longer reuses the COMPACTION_RECOVERY_BAND
hysteresis. Reusing it required residual context below 0.8×threshold, which
turned a recoverable overflow (e.g. 150k on a 200k window, under the ~170k
threshold) into a manual dead-end. Add #compactionCreatedRetryFit, which
measures the rebuilt prompt against the usable fit budget
(contextWindow - effectiveReserveTokens) and is evaluated AFTER the failed
assistant is dropped, so the just-failed turn is excluded.
- The threshold auto-continue keeps the stricter recovery-band check
(#compactionCreatedHeadroom) — that path is the snapcompact thrash guard.
- Split the no-progress warning so each path warns on its own signal.
- Alias errorIsFromBeforeCompaction as assistantPredatesCompaction in the
threshold-usage path for clarity; de-duplicate the rationale comment.
Adds two regression tests (recoverable overflow retries; non-fitting overflow
pauses + warns), mutation-verified against the band-vs-fit split.
(cherry picked from commit f8cf45b2740f04d8776a84521ba563b779d65b25)
prepareCompaction keeps the most-recent turn verbatim (findCutPoint never
cuts at tool results), so when that single turn already exceeds the
compaction threshold the rewritten context stays above threshold. The
context-full / snapcompact auto-compaction success tail scheduled the
agent-authored auto-continue (and the overflow/incomplete retry)
unconditionally, so the next agent_end re-entered #checkCompaction over
the same oversized tail and re-fired forever.
This is the residual loop left after #3247 capped snapcompact's own frame
projection: once #computeSnapcompactMaxFrames drops the frame cap below
one frame, snapcompact is skipped and the context-full summarizer path
still creates no headroom, so the loop persists on that path.
- #runAutoCompaction now gates the continuation and the overflow/
incomplete retry on a post-maintenance headroom check
(#compactionCreatedHeadroom), reusing the shake recovery-band
hysteresis from #2275 (consolidated as the shared
COMPACTION_RECOVERY_BAND). When a pass frees too little it pauses
automatic maintenance and emits a single warning instead of looping.
- The post-turn threshold check ignores an assistant's stale
pre-compaction usage so the scheduled auto-continue can't re-trip on
the kept assistant's old high token count.
Adds a regression test covering the no-headroom (pause + warn, no
continuation) and headroom (auto-continue, no warn) paths; both
assertions are mutation-verified against the guard.
(cherry picked from commit 6fcbbe2b3075827ee621acaee8efa2c586eedd3e)
Stopped snapcompact preflight failures from falling through to the provider-backed LLM summarizer and covered manual plus auto compaction paths.
Fixes#3599
Passed the settled assistant message into session_stop emission so refusal-as-error turns can be pruned from replay context without hiding their stop details from extension hooks.
Expanded the refusal regression test to assert the session_stop payload still exposes the refusal as last_assistant_message.
Fixes#3591
Removed the early return after refusal pruning so the agent_end tail still reaches `#emitSessionStopEvent`, restoring `session_stop` extension hooks (block/continue/telemetry) for refusal-as-error stops.
Regression test wires an extensionRunner with a session_stop handler and asserts it fires for both the refusal turn and the following clean turn.
Fixes#3591
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.
- Fixed stale `preserveData.snapcompact` frames leaking into context-full compaction after switching from `snapcompact` to `context-full` strategy, which inflated context usage and made sessions appear to compact prematurely.
- Added secret redaction for migrated snapcompact archive plaintext (`text`/`textHead`/`textTail`) during the snapcompact->context-full transition, while preserving opaque provider-replay state byte-identical.
- Added `archiveSourceText()` and `stripPreservedArchive()` utilities to snapcompact module for archive extraction and cleanup.
- Consolidated duplicate `stripSnapcompactPreserveData` functions into `snapcompact.stripPreservedArchive`.
- Added unit tests to verify archive removal and empty state collapse behavior.
The advisor system prompt told the watcher model "at most one advise per
update" and "NEVER send the same advice twice", but nothing enforced
either rule. Issue #3520 captured a session where the advisor emitted
309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No
issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker">
injections in the primary transcript and destabilizing the watched agent
after the task was already complete.
New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and:
- Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every
"Stop.", "*Stop*", "STOP!" variant keys to the same canonical form.
- Drops a small allowlist of content-free self-talk filler (stop, done,
complete, no issue continue, lgtm, nothing to add, no further input,
carry on, ...) - silence is the correct expression of "no concerns".
- Dedupes by exact normalized text across the session, FIFO-bounded at
4096 entries.
- Rate-limits to one accepted advise per advisor model prompt cycle. The
runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(),
so the new batch starts with a fresh budget. Suppressed calls don't
consume the budget - a noise call never displaces a real concern.
Reset on advisor reset (compaction, session switch, /new) so a re-primed
reviewer can re-raise old concerns against the rewritten transcript.
Suppression is invisible to the advisor model: AdviseTool still returns
"Recorded." for a dropped call. Surfacing "suppressed" risks the model
rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to
bypass the dedupe.
Fixes#3520
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.
Added coverage that repeated advice is allowed again after the dedupe state resets.
Fixes#3511
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
The preserveCompaction abort path skipped abortCompaction() entirely,
so a manual /compact starting while auto-compaction was in flight no
longer cancelled it. Both passes could then appendCompaction/
replaceMessages, double-rewriting history (reachable via the RPC/
extension compact paths, whose only guard checks #compactionAbortController).
Preserve the just-installed manual controller but still abort the
auto-compaction controller. Adds regression coverage.
Adding `paths` to the global PRIMARY_ARG_KEYS hid the pattern for the
structural tools: ast_grep ({pat,paths}) and ast_edit ({ops,paths})
rendered scope-only (e.g. `ast_grep(src/**/*.ts)`), dropping `pat`/`ops`
— the most decision-relevant argument. Drop the global `paths` key and
special-case `find` (mirroring `search`) so find/search still surface
scope while ast_grep/ast_edit keep showing their pattern via the
existing fallback. Adds regression tests for both structural tools.
Install the manual compaction abort controller before abort teardown so input routing observes session.isCompacting during the starting window.
Fixes#3485
Added scoped path summaries for find/search tool calls in concise session history rendering, with regression coverage for JSON fallback and hidden search scope.
Fixes#3482
Reviewer caught that demoting the whole mixed payload to `user` (`@notes.md
@screenshot.png`) regressed the developer-priority treatment text-only
mentions still get for image-free turns. `generateFileMentionMessages` packs
every `@…` into one `fileMention`, so the previous `hasImage` toggle
collapsed the source-file context into the user slot whenever an image was
attached.
`convertToLlm` now returns up to two messages per `fileMention` via
`flatMap`: text-only files keep their existing `developer` envelope, and
image-bearing files emit a separate `user` envelope that carries their
`<file>` wrappers plus the `input_image` block. Pure-text and pure-image
turns still collapse to a single message.
Tests cover the mixed case (split into developer + user), the image-only case
(single user message), and the existing text-only case (single developer
message).
Fixes#3443
Codex GPT models on chatgpt.com /codex/responses rejected `@image` turns with
`Codex error event: [OneOfParam] [input[N].content[M]] [invalid_enum_value]
Invalid value: 'input_image'. Supported values are: 'input_text'.` —
`convertToLlm`'s `fileMention` arm always emitted a `developer`-role
Responses message, but a developer-role content slot only accepts
`input_text`. #3421's prior fix only suppressed the Codex Responses Lite
header on image-bearing turns; the full transport kept rejecting the same body.
`fileMention` now uses `user` role when any attached file carries an image;
text-only mentions keep `developer` so the auto-read context still rides at
instruction priority for the agent.
Fixes#3443