- Added a check to promote the context model before triggering mid-turn compaction.
- Prevented unnecessary compaction during long tool loops that cross threshold limits by enabling window expansion instead.
- Implemented a monitoring system to identify Gemini model runaway behavior during thinking steps using consecutive header detection.
- Introduced a configurable tool-call reminder prompt to inject corrective context when reasoning stalls.
- Added session-level logic to automatically interrupt and prune stalled assistant turns from the conversation history.
- Provided comprehensive test coverage for the detection logic and stream interruption scenarios.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Implemented file deletion (`REM`) and movement (`MV`) operations within the hashline grammar and parser.
- Added filesystem support for executing delete and move commands while maintaining snapshot history migration.
- Integrated file operation detection and parsing logic into the coding agent and patcher components.
- Added comprehensive integration tests and documentation for the new file-level operation syntax.
The post-maintenance headroom guard returned residualTokens < triggerContextTokens
as a secondary check after the recovery-band test. When stale/tool-output pruning
already drove the trigger (postMaintenanceContextTokens) below the band before this
pass, that strict-less comparison made a residual which merely held the line at/under
the band report a false no-progress, suppressing a valid auto-continue and emitting a
spurious warning even though the next turn could no longer re-trip threshold compaction.
The recovery band sits strictly under the compaction threshold, so reaching it already
guarantees the next turn cannot re-trip. Make the band authoritative (residual <= band)
and drop the trigger argument. Adds a regression covering the sub-band trigger case.
(cherry picked from commit 31cf2dc7a30c4134f7a351d61065eca76543f154)
Addresses PR #3412 review (roboomp blocking + codex P2 + Copilot nits):
- The overflow/incomplete retry no longer reuses the COMPACTION_RECOVERY_BAND
hysteresis. Reusing it required residual context below 0.8×threshold, which
turned a recoverable overflow (e.g. 150k on a 200k window, under the ~170k
threshold) into a manual dead-end. Add #compactionCreatedRetryFit, which
measures the rebuilt prompt against the usable fit budget
(contextWindow - effectiveReserveTokens) and is evaluated AFTER the failed
assistant is dropped, so the just-failed turn is excluded.
- The threshold auto-continue keeps the stricter recovery-band check
(#compactionCreatedHeadroom) — that path is the snapcompact thrash guard.
- Split the no-progress warning so each path warns on its own signal.
- Alias errorIsFromBeforeCompaction as assistantPredatesCompaction in the
threshold-usage path for clarity; de-duplicate the rationale comment.
Adds two regression tests (recoverable overflow retries; non-fitting overflow
pauses + warns), mutation-verified against the band-vs-fit split.
(cherry picked from commit f8cf45b2740f04d8776a84521ba563b779d65b25)
prepareCompaction keeps the most-recent turn verbatim (findCutPoint never
cuts at tool results), so when that single turn already exceeds the
compaction threshold the rewritten context stays above threshold. The
context-full / snapcompact auto-compaction success tail scheduled the
agent-authored auto-continue (and the overflow/incomplete retry)
unconditionally, so the next agent_end re-entered #checkCompaction over
the same oversized tail and re-fired forever.
This is the residual loop left after #3247 capped snapcompact's own frame
projection: once #computeSnapcompactMaxFrames drops the frame cap below
one frame, snapcompact is skipped and the context-full summarizer path
still creates no headroom, so the loop persists on that path.
- #runAutoCompaction now gates the continuation and the overflow/
incomplete retry on a post-maintenance headroom check
(#compactionCreatedHeadroom), reusing the shake recovery-band
hysteresis from #2275 (consolidated as the shared
COMPACTION_RECOVERY_BAND). When a pass frees too little it pauses
automatic maintenance and emits a single warning instead of looping.
- The post-turn threshold check ignores an assistant's stale
pre-compaction usage so the scheduled auto-continue can't re-trip on
the kept assistant's old high token count.
Adds a regression test covering the no-headroom (pause + warn, no
continuation) and headroom (auto-continue, no warn) paths; both
assertions are mutation-verified against the guard.
(cherry picked from commit 6fcbbe2b3075827ee621acaee8efa2c586eedd3e)
Stopped snapcompact preflight failures from falling through to the provider-backed LLM summarizer and covered manual plus auto compaction paths.
Fixes#3599
Passed the settled assistant message into session_stop emission so refusal-as-error turns can be pruned from replay context without hiding their stop details from extension hooks.
Expanded the refusal regression test to assert the session_stop payload still exposes the refusal as last_assistant_message.
Fixes#3591
Removed the early return after refusal pruning so the agent_end tail still reaches `#emitSessionStopEvent`, restoring `session_stop` extension hooks (block/continue/telemetry) for refusal-as-error stops.
Regression test wires an extensionRunner with a session_stop handler and asserts it fires for both the refusal turn and the following clean turn.
Fixes#3591
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.
- Fixed stale `preserveData.snapcompact` frames leaking into context-full compaction after switching from `snapcompact` to `context-full` strategy, which inflated context usage and made sessions appear to compact prematurely.
- Added secret redaction for migrated snapcompact archive plaintext (`text`/`textHead`/`textTail`) during the snapcompact->context-full transition, while preserving opaque provider-replay state byte-identical.
- Added `archiveSourceText()` and `stripPreservedArchive()` utilities to snapcompact module for archive extraction and cleanup.
- Consolidated duplicate `stripSnapcompactPreserveData` functions into `snapcompact.stripPreservedArchive`.
- Added unit tests to verify archive removal and empty state collapse behavior.
The advisor system prompt told the watcher model "at most one advise per
update" and "NEVER send the same advice twice", but nothing enforced
either rule. Issue #3520 captured a session where the advisor emitted
309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No
issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker">
injections in the primary transcript and destabilizing the watched agent
after the task was already complete.
New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and:
- Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every
"Stop.", "*Stop*", "STOP!" variant keys to the same canonical form.
- Drops a small allowlist of content-free self-talk filler (stop, done,
complete, no issue continue, lgtm, nothing to add, no further input,
carry on, ...) - silence is the correct expression of "no concerns".
- Dedupes by exact normalized text across the session, FIFO-bounded at
4096 entries.
- Rate-limits to one accepted advise per advisor model prompt cycle. The
runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(),
so the new batch starts with a fresh budget. Suppressed calls don't
consume the budget - a noise call never displaces a real concern.
Reset on advisor reset (compaction, session switch, /new) so a re-primed
reviewer can re-raise old concerns against the rewritten transcript.
Suppression is invisible to the advisor model: AdviseTool still returns
"Recorded." for a dropped call. Surfacing "suppressed" risks the model
rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to
bypass the dedupe.
Fixes#3520
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.
Added coverage that repeated advice is allowed again after the dedupe state resets.
Fixes#3511
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
The preserveCompaction abort path skipped abortCompaction() entirely,
so a manual /compact starting while auto-compaction was in flight no
longer cancelled it. Both passes could then appendCompaction/
replaceMessages, double-rewriting history (reachable via the RPC/
extension compact paths, whose only guard checks #compactionAbortController).
Preserve the just-installed manual controller but still abort the
auto-compaction controller. Adds regression coverage.
Install the manual compaction abort controller before abort teardown so input routing observes session.isCompacting during the starting window.
Fixes#3485
Replaces the old /move (which relocated the current session file) with a
new flow that starts a fresh empty session in the target directory, leaving
the previous session resumable via /resume. With no argument, /move opens
a path autocomplete overlay (type to filter, Tab to accept, Enter to
confirm). If the target directory does not exist, a confirmation prompt
offers to create it. Empty move sessions are cleaned up on shutdown.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Trailing empty assistant 'stop' arriving after a successful 'yield'
revived the already-yielded subagent. AgentSession.agent_end maintenance
compared #assistantEndedWithSuccessfulYield(msg) against the trailing
empty-stop message — not the yield-bearing one — so the empty-stop
recovery path appended a retry reminder and scheduled agent.continue().
Track a sticky #yieldTerminationPending flag set when the yield tool
finishes without error and cleared on the next #promptWithMessage. The
agent_end routing extends the existing successful-yield branch: when the
flag is set, or the current message ended with yield, short-circuit
empty-stop / unexpected-stop / compaction continuations for the rest of
the run, so a successful yield is terminal regardless of trailing stops.
Fixes#3389
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.
Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.
Fixes#3356
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
Active goal loops can stay inside one agent run while the model keeps
emitting tool calls, so the normal agent_end threshold maintenance never
runs. That lets context grow past the soft threshold until provider
overflow or user abort.
Run threshold maintenance from the per-turn onTurnEnd hook for active
goals, splice the compacted agent state back into the live loop message
array, and suppress queued continuations because the current run is
already continuing. Cover the mid-run tool-call path and the non-goal
control case.
Refs #3174
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.