Commit Graph
918 Commits
Author SHA1 Message Date
Wolfgang Schoenbergerandcan1357 699e07fea7 fix(coding-agent): separate overflow-retry fit check from auto-continue band
Addresses PR #3412 review (roboomp blocking + codex P2 + Copilot nits):

- The overflow/incomplete retry no longer reuses the COMPACTION_RECOVERY_BAND
  hysteresis. Reusing it required residual context below 0.8×threshold, which
  turned a recoverable overflow (e.g. 150k on a 200k window, under the ~170k
  threshold) into a manual dead-end. Add #compactionCreatedRetryFit, which
  measures the rebuilt prompt against the usable fit budget
  (contextWindow - effectiveReserveTokens) and is evaluated AFTER the failed
  assistant is dropped, so the just-failed turn is excluded.
- The threshold auto-continue keeps the stricter recovery-band check
  (#compactionCreatedHeadroom) — that path is the snapcompact thrash guard.
- Split the no-progress warning so each path warns on its own signal.
- Alias errorIsFromBeforeCompaction as assistantPredatesCompaction in the
  threshold-usage path for clarity; de-duplicate the rationale comment.

Adds two regression tests (recoverable overflow retries; non-fitting overflow
pauses + warns), mutation-verified against the band-vs-fit split.

(cherry picked from commit f8cf45b2740f04d8776a84521ba563b779d65b25)
2026-06-26 23:42:50 +02:00
Wolfgang Schoenbergerandcan1357 6a55205b68 fix(coding-agent): stop auto-compaction thrash when a single turn exceeds the threshold
prepareCompaction keeps the most-recent turn verbatim (findCutPoint never
cuts at tool results), so when that single turn already exceeds the
compaction threshold the rewritten context stays above threshold. The
context-full / snapcompact auto-compaction success tail scheduled the
agent-authored auto-continue (and the overflow/incomplete retry)
unconditionally, so the next agent_end re-entered #checkCompaction over
the same oversized tail and re-fired forever.

This is the residual loop left after #3247 capped snapcompact's own frame
projection: once #computeSnapcompactMaxFrames drops the frame cap below
one frame, snapcompact is skipped and the context-full summarizer path
still creates no headroom, so the loop persists on that path.

- #runAutoCompaction now gates the continuation and the overflow/
  incomplete retry on a post-maintenance headroom check
  (#compactionCreatedHeadroom), reusing the shake recovery-band
  hysteresis from #2275 (consolidated as the shared
  COMPACTION_RECOVERY_BAND). When a pass frees too little it pauses
  automatic maintenance and emits a single warning instead of looping.
- The post-turn threshold check ignores an assistant's stale
  pre-compaction usage so the scheduled auto-continue can't re-trip on
  the kept assistant's old high token count.

Adds a regression test covering the no-headroom (pause + warn, no
continuation) and headroom (auto-continue, no warn) paths; both
assertions are mutation-verified against the guard.

(cherry picked from commit 6fcbbe2b3075827ee621acaee8efa2c586eedd3e)
2026-06-26 23:42:50 +02:00
can1357 4e3a0a0f46 Merge PR #3523: fix(advisor): filter content-free advisor notes at enqueue boundary (@roboomp)
# Conflicts:
#	packages/coding-agent/src/session/agent-session.ts
2026-06-26 23:42:40 +02:00
can1357 c50af68125 Merge PR #3594: fix(session): prune classifier refusals from context (@roboomp) 2026-06-26 23:42:15 +02:00
roboomp 061c0ffb43 fix(compaction): kept snapcompact local on preflight failure
Stopped snapcompact preflight failures from falling through to the provider-backed LLM summarizer and covered manual plus auto compaction paths.

Fixes #3599
2026-06-26 20:37:50 +00:00
roboomp 11c96c9ae3 fix(session): preserved refusal stop payloads
Passed the settled assistant message into session_stop emission so refusal-as-error turns can be pruned from replay context without hiding their stop details from extension hooks.

Expanded the refusal regression test to assert the session_stop payload still exposes the refusal as last_assistant_message.

Fixes #3591
2026-06-26 18:06:09 +00:00
roboomp db61ca0274 fix(session): kept session_stop hooks on classifier refusals
Removed the early return after refusal pruning so the agent_end tail still reaches `#emitSessionStopEvent`, restoring `session_stop` extension hooks (block/continue/telemetry) for refusal-as-error stops.

Regression test wires an extensionRunner with a session_stop handler and asserts it fires for both the refusal turn and the following clean turn.

Fixes #3591
2026-06-26 17:59:26 +00:00
roboomp fbcc157b75 fix(session): pruned classifier refusals from context
Stopped API-level classifier refusal messages from being persisted or replayed as assistant dialogue after the failed turn settles.

Fixes #3591
2026-06-26 17:53:48 +00:00
can1357 5d4203d52b fix(coding-agent): preserved resume scoping after PR merges 2026-06-26 17:17:43 +02:00
can1357 d07d03d2af Merge PR #3565: compact long tool loops mid-turn (@riverpilot) 2026-06-26 17:12:09 +02:00
can1357 b6b0379a3b Merge PR #3540: manual coding-agent GC command (@Kenmege)
# Conflicts:
#	packages/coding-agent/test/issue-3461-repro.test.ts
2026-06-26 17:09:24 +02:00
can1357 80c6528c93 fix(coding-agent): resolved PR 3562 merge fallout 2026-06-26 17:08:53 +02:00
Alexander Kirilin e9530895a1 Merge remote-tracking branch 'origin/main' into fix/mid-turn-auto-compaction-3525 2026-06-26 11:02:40 -04:00
Alexander Kirilin d2b6bde571 fix(agent): preserve mid-turn tool persistence order 2026-06-26 10:57:36 -04:00
can1357 02f870fb20 feat: supported markdown section operations for block edits
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.
2026-06-26 16:14:23 +02:00
Alexander Kirilin ab513757f8 fix(agent): preserve tool turn before mid-run compaction 2026-06-26 09:35:24 -04:00
Alexander Kirilin e99b8c4128 docs(agent): clarify mid-turn continuation hook 2026-06-26 09:17:07 -04:00
Alexander Kirilinandcoderred bbe3d98375 fix(agent): compact long tool loops mid-turn
Co-authored-by: coderred <coderredlab@gmail.com>
2026-06-26 09:05:11 -04:00
can1357 415099c420 fix: stripped stale snapcompact archive state during compaction strategy migration
- Fixed stale `preserveData.snapcompact` frames leaking into context-full compaction after switching from `snapcompact` to `context-full` strategy, which inflated context usage and made sessions appear to compact prematurely.
- Added secret redaction for migrated snapcompact archive plaintext (`text`/`textHead`/`textTail`) during the snapcompact->context-full transition, while preserving opaque provider-replay state byte-identical.
- Added `archiveSourceText()` and `stripPreservedArchive()` utilities to snapcompact module for archive extraction and cleanup.
2026-06-26 14:31:50 +02:00
can1357 93b80e5b1b refactor: consolidated duplicate
- Consolidated duplicate `stripSnapcompactPreserveData` functions into `snapcompact.stripPreservedArchive`.
- Added unit tests to verify archive removal and empty state collapse behavior.
2026-06-26 14:28:49 +02:00
OpenAI GPT-5.5 531faf80a5 fix(agent): strip snapcompact state on session compaction
Co-Authored-By: OpenAI GPT-5.5 <noreply@openai.com>
2026-06-26 09:43:02 +00:00
Dr Kennedy Umege 4ead62cace feat: add manual coding-agent gc command 2026-06-26 08:39:50 +01:00
roboomp ce4f5b8615 fix(advisor): enforced one-advise-per-update + dedupe + noise filter at the enqueueAdvice boundary
The advisor system prompt told the watcher model "at most one advise per
update" and "NEVER send the same advice twice", but nothing enforced
either rule. Issue #3520 captured a session where the advisor emitted
309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No
issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker">
injections in the primary transcript and destabilizing the watched agent
after the task was already complete.

New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and:

- Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every
  "Stop.", "*Stop*", "STOP!" variant keys to the same canonical form.
- Drops a small allowlist of content-free self-talk filler (stop, done,
  complete, no issue continue, lgtm, nothing to add, no further input,
  carry on, ...) - silence is the correct expression of "no concerns".
- Dedupes by exact normalized text across the session, FIFO-bounded at
  4096 entries.
- Rate-limits to one accepted advise per advisor model prompt cycle. The
  runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(),
  so the new batch starts with a fresh budget. Suppressed calls don't
  consume the budget - a noise call never displaces a real concern.

Reset on advisor reset (compaction, session switch, /new) so a re-primed
reviewer can re-raise old concerns against the rewritten transcript.

Suppression is invisible to the advisor model: AdviseTool still returns
"Recorded." for a dropped call. Surfacing "suppressed" risks the model
rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to
bypass the dedupe.

Fixes #3520
2026-06-26 02:48:27 +00:00
roboomp cce5133bfa fix(advisor): reset advisory dedupe state
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.

Added coverage that repeated advice is allowed again after the dedupe state resets.

Fixes #3511
2026-06-26 00:31:51 +00:00
can1357 5a50b047d5 Merge PR #3428: Fix active goal compaction after yield stop (@cexll) 2026-06-25 22:39:28 +02:00
can1357 938489f3fd feat(coding-agent): added configurable service tier settings for subagents and advisor
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
2026-06-25 22:35:15 +02:00
can1357 556ced266c Merge PR #3466 into sweep 2026-06-25 20:42:31 +02:00
can1357 88783b6e5c fix(agent): cancel in-flight auto-compaction on manual compact startup
The preserveCompaction abort path skipped abortCompaction() entirely,
so a manual /compact starting while auto-compaction was in flight no
longer cancelled it. Both passes could then appendCompaction/
replaceMessages, double-rewriting history (reachable via the RPC/
extension compact paths, whose only guard checks #compactionAbortController).
Preserve the just-installed manual controller but still abort the
auto-compaction controller. Adds regression coverage.
2026-06-25 20:42:30 +02:00
can1357 5cfea2884d Merge PR #3489 into sweep 2026-06-25 20:42:30 +02:00
pr-evalandcan1357 0fb6af9354 fix(coding-agent): keep ast_grep/ast_edit patterns in transcript summaries
Adding `paths` to the global PRIMARY_ARG_KEYS hid the pattern for the
structural tools: ast_grep ({pat,paths}) and ast_edit ({ops,paths})
rendered scope-only (e.g. `ast_grep(src/**/*.ts)`), dropping `pat`/`ops`
— the most decision-relevant argument. Drop the global `paths` key and
special-case `find` (mirroring `search`) so find/search still surface
scope while ast_grep/ast_edit keep showing their pattern via the
existing fallback. Adds regression tests for both structural tools.
2026-06-25 20:42:30 +02:00
roboomp d380a9e723 fix(agent): queued manual compact startup steers
Install the manual compaction abort controller before abort teardown so input routing observes session.isCompacting during the starting window.

Fixes #3485
2026-06-25 15:49:23 +00:00
roboomp 8f63bd60be fix(coding-agent): surfaced transcript search paths
Added scoped path summaries for find/search tool calls in concise session history rendering, with regression coverage for JSON fallback and hidden search scope.

Fixes #3482
2026-06-25 15:38:59 +00:00
roboomp 80862b79da fix(agent): handled ollama-cloud task backoff
Added ollama-cloud subagent concurrency limiting, role fallback-chain inheritance, and visible empty length errors for native Ollama responses.

Fixes #3464
2026-06-25 11:57:51 +00:00
roboomp 897cce792f fix(coding-agent): split mixed file mentions so text stays on developer
Reviewer caught that demoting the whole mixed payload to `user` (`@notes.md
@screenshot.png`) regressed the developer-priority treatment text-only
mentions still get for image-free turns. `generateFileMentionMessages` packs
every `@…` into one `fileMention`, so the previous `hasImage` toggle
collapsed the source-file context into the user slot whenever an image was
attached.

`convertToLlm` now returns up to two messages per `fileMention` via
`flatMap`: text-only files keep their existing `developer` envelope, and
image-bearing files emit a separate `user` envelope that carries their
`<file>` wrappers plus the `input_image` block. Pure-text and pure-image
turns still collapse to a single message.

Tests cover the mixed case (split into developer + user), the image-only case
(single user message), and the existing text-only case (single developer
message).

Fixes #3443
2026-06-25 05:54:15 +00:00
roboomp c2174a87b2 fix(coding-agent): routed image-bearing @ mentions as user-role messages
Codex GPT models on chatgpt.com /codex/responses rejected `@image` turns with
`Codex error event: [OneOfParam] [input[N].content[M]] [invalid_enum_value]
Invalid value: 'input_image'. Supported values are: 'input_text'.` —
`convertToLlm`'s `fileMention` arm always emitted a `developer`-role
Responses message, but a developer-role content slot only accepts
`input_text`. #3421's prior fix only suppressed the Codex Responses Lite
header on image-bearing turns; the full transport kept rejecting the same body.

`fileMention` now uses `user` role when any attached file carries an image;
text-only mentions keep `developer` so the auto-read context still rides at
instruction priority for the agent.

Fixes #3443
2026-06-25 05:47:07 +00:00
ben cad49e06a0 fix yield compaction diagnostics 2026-06-25 11:27:10 +08:00
ben 872e60f056 fix active goal compaction after yield stop 2026-06-25 10:33:30 +08:00
can1357 0eb21efa1a Merge remote-tracking branch 'origin/farm/7ef98714/snapcompact-copilot-vision-gate' 2026-06-24 21:00:37 +02:00
can1357 0f2737f16e Merge remote-tracking branch 'origin/farm/e322f828/eval-agent-yield-terminal' 2026-06-24 20:59:54 +02:00
roboomp 997b2b24ff fix(session): suppressed empty-stop retry after successful yield
Trailing empty assistant 'stop' arriving after a successful 'yield'
revived the already-yielded subagent. AgentSession.agent_end maintenance
compared #assistantEndedWithSuccessfulYield(msg) against the trailing
empty-stop message — not the yield-bearing one — so the empty-stop
recovery path appended a retry reminder and scheduled agent.continue().

Track a sticky #yieldTerminationPending flag set when the yield tool
finishes without error and cleared on the next #promptWithMessage. The
agent_end routing extends the existing successful-yield branch: when the
flag is set, or the current message ended with yield, short-circuit
empty-stop / unexpected-stop / compaction continuations for the rest of
the run, so a successful yield is terminal regardless of trailing stops.

Fixes #3389
2026-06-24 17:08:49 +00:00
roboomp 714051d795 fix(catalog,coding-agent): disable vision on non-personal copilot endpoints
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.

- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
  forces input=['text'] whenever the resolved baseUrl is not the
  canonical personal-Copilot host, so the upstream's vision flag is
  honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
  OR-upgrading with the bundled reference) when the merged baseUrl
  differs from the bundled one, so a bundled spec pinned to the
  personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
  rasterizer for any github-copilot model whose baseUrl is non-personal,
  catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
  pi-catalog/wire/github-copilot so catalog and coding-agent share one
  canonical check.

Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).

Fixes #3387
2026-06-24 16:27:14 +00:00
can1357 9c56956631 fix(tui): kept recent turns visible when collapsing remote compactions
With collapseCompactedHistory the live display fell into the LLM compaction
branch, which skips the firstKeptEntryId..compaction turns whenever an OpenAI
remote-compaction replacementHistory payload is present. That payload feeds the
provider only and is not rendered, so a remotely-compacted session showed just
the summary plus post-compaction rows, hiding recent turns that were visible
before. Emit the kept SessionEntry rows in transcript mode regardless. Adds a
regression.
2026-06-24 18:26:20 +02:00
can1357 c04747c5ac Merge PR #3259: fix(tui): reduce large transcript stalls (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 3117a20ab9 Merge PR #3249: fix(agent): size snapcompact maxFrames by the live model window (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 ef67f685ac Merge PR #3380: fix(coding-agent): clamp auto thinking to undefined for models without controllable effort (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 2b42f19456 Merge PR #3232: fix(agent): clamp provider context images (@roboomp) 2026-06-24 18:23:48 +02:00
roboomp b0e07f52d2 fix(coding-agent): clamped auto thinking to undefined for models without controllable effort
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.

Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.

Fixes #3356
2026-06-24 14:12:18 +00:00
roboomp 24adad2890 fix(providers): honored llama cpp model context
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
2026-06-23 12:23:59 +00:00
can1357 44b62419c5 Merge remote-tracking branch 'origin/farm/fc89a4e3/goal-auto-compaction-still-not-triggering' 2026-06-23 00:01:47 +02:00
roboomp 29e69f5364 style: bun run fix 2026-06-22 20:02:46 +00:00