Commit Graph
909 Commits
Author SHA1 Message Date
can1357 5d4203d52b fix(coding-agent): preserved resume scoping after PR merges 2026-06-26 17:17:43 +02:00
can1357 d07d03d2af Merge PR #3565: compact long tool loops mid-turn (@riverpilot) 2026-06-26 17:12:09 +02:00
can1357 b6b0379a3b Merge PR #3540: manual coding-agent GC command (@Kenmege)
# Conflicts:
#	packages/coding-agent/test/issue-3461-repro.test.ts
2026-06-26 17:09:24 +02:00
can1357 80c6528c93 fix(coding-agent): resolved PR 3562 merge fallout 2026-06-26 17:08:53 +02:00
Alexander Kirilin e9530895a1 Merge remote-tracking branch 'origin/main' into fix/mid-turn-auto-compaction-3525 2026-06-26 11:02:40 -04:00
Alexander Kirilin d2b6bde571 fix(agent): preserve mid-turn tool persistence order 2026-06-26 10:57:36 -04:00
can1357 02f870fb20 feat: supported markdown section operations for block edits
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.
2026-06-26 16:14:23 +02:00
Alexander Kirilin ab513757f8 fix(agent): preserve tool turn before mid-run compaction 2026-06-26 09:35:24 -04:00
Alexander Kirilin e99b8c4128 docs(agent): clarify mid-turn continuation hook 2026-06-26 09:17:07 -04:00
Alexander Kirilinandcoderred bbe3d98375 fix(agent): compact long tool loops mid-turn
Co-authored-by: coderred <coderredlab@gmail.com>
2026-06-26 09:05:11 -04:00
can1357 415099c420 fix: stripped stale snapcompact archive state during compaction strategy migration
- Fixed stale `preserveData.snapcompact` frames leaking into context-full compaction after switching from `snapcompact` to `context-full` strategy, which inflated context usage and made sessions appear to compact prematurely.
- Added secret redaction for migrated snapcompact archive plaintext (`text`/`textHead`/`textTail`) during the snapcompact->context-full transition, while preserving opaque provider-replay state byte-identical.
- Added `archiveSourceText()` and `stripPreservedArchive()` utilities to snapcompact module for archive extraction and cleanup.
2026-06-26 14:31:50 +02:00
can1357 93b80e5b1b refactor: consolidated duplicate
- Consolidated duplicate `stripSnapcompactPreserveData` functions into `snapcompact.stripPreservedArchive`.
- Added unit tests to verify archive removal and empty state collapse behavior.
2026-06-26 14:28:49 +02:00
OpenAI GPT-5.5 531faf80a5 fix(agent): strip snapcompact state on session compaction
Co-Authored-By: OpenAI GPT-5.5 <noreply@openai.com>
2026-06-26 09:43:02 +00:00
Dr Kennedy Umege 4ead62cace feat: add manual coding-agent gc command 2026-06-26 08:39:50 +01:00
roboomp cce5133bfa fix(advisor): reset advisory dedupe state
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.

Added coverage that repeated advice is allowed again after the dedupe state resets.

Fixes #3511
2026-06-26 00:31:51 +00:00
can1357 5a50b047d5 Merge PR #3428: Fix active goal compaction after yield stop (@cexll) 2026-06-25 22:39:28 +02:00
can1357 938489f3fd feat(coding-agent): added configurable service tier settings for subagents and advisor
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
2026-06-25 22:35:15 +02:00
can1357 556ced266c Merge PR #3466 into sweep 2026-06-25 20:42:31 +02:00
can1357 88783b6e5c fix(agent): cancel in-flight auto-compaction on manual compact startup
The preserveCompaction abort path skipped abortCompaction() entirely,
so a manual /compact starting while auto-compaction was in flight no
longer cancelled it. Both passes could then appendCompaction/
replaceMessages, double-rewriting history (reachable via the RPC/
extension compact paths, whose only guard checks #compactionAbortController).
Preserve the just-installed manual controller but still abort the
auto-compaction controller. Adds regression coverage.
2026-06-25 20:42:30 +02:00
can1357 5cfea2884d Merge PR #3489 into sweep 2026-06-25 20:42:30 +02:00
pr-evalandcan1357 0fb6af9354 fix(coding-agent): keep ast_grep/ast_edit patterns in transcript summaries
Adding `paths` to the global PRIMARY_ARG_KEYS hid the pattern for the
structural tools: ast_grep ({pat,paths}) and ast_edit ({ops,paths})
rendered scope-only (e.g. `ast_grep(src/**/*.ts)`), dropping `pat`/`ops`
— the most decision-relevant argument. Drop the global `paths` key and
special-case `find` (mirroring `search`) so find/search still surface
scope while ast_grep/ast_edit keep showing their pattern via the
existing fallback. Adds regression tests for both structural tools.
2026-06-25 20:42:30 +02:00
roboomp d380a9e723 fix(agent): queued manual compact startup steers
Install the manual compaction abort controller before abort teardown so input routing observes session.isCompacting during the starting window.

Fixes #3485
2026-06-25 15:49:23 +00:00
roboomp 8f63bd60be fix(coding-agent): surfaced transcript search paths
Added scoped path summaries for find/search tool calls in concise session history rendering, with regression coverage for JSON fallback and hidden search scope.

Fixes #3482
2026-06-25 15:38:59 +00:00
roboomp 80862b79da fix(agent): handled ollama-cloud task backoff
Added ollama-cloud subagent concurrency limiting, role fallback-chain inheritance, and visible empty length errors for native Ollama responses.

Fixes #3464
2026-06-25 11:57:51 +00:00
roboomp 897cce792f fix(coding-agent): split mixed file mentions so text stays on developer
Reviewer caught that demoting the whole mixed payload to `user` (`@notes.md
@screenshot.png`) regressed the developer-priority treatment text-only
mentions still get for image-free turns. `generateFileMentionMessages` packs
every `@…` into one `fileMention`, so the previous `hasImage` toggle
collapsed the source-file context into the user slot whenever an image was
attached.

`convertToLlm` now returns up to two messages per `fileMention` via
`flatMap`: text-only files keep their existing `developer` envelope, and
image-bearing files emit a separate `user` envelope that carries their
`<file>` wrappers plus the `input_image` block. Pure-text and pure-image
turns still collapse to a single message.

Tests cover the mixed case (split into developer + user), the image-only case
(single user message), and the existing text-only case (single developer
message).

Fixes #3443
2026-06-25 05:54:15 +00:00
roboomp c2174a87b2 fix(coding-agent): routed image-bearing @ mentions as user-role messages
Codex GPT models on chatgpt.com /codex/responses rejected `@image` turns with
`Codex error event: [OneOfParam] [input[N].content[M]] [invalid_enum_value]
Invalid value: 'input_image'. Supported values are: 'input_text'.` —
`convertToLlm`'s `fileMention` arm always emitted a `developer`-role
Responses message, but a developer-role content slot only accepts
`input_text`. #3421's prior fix only suppressed the Codex Responses Lite
header on image-bearing turns; the full transport kept rejecting the same body.

`fileMention` now uses `user` role when any attached file carries an image;
text-only mentions keep `developer` so the auto-read context still rides at
instruction priority for the agent.

Fixes #3443
2026-06-25 05:47:07 +00:00
ben cad49e06a0 fix yield compaction diagnostics 2026-06-25 11:27:10 +08:00
ben 872e60f056 fix active goal compaction after yield stop 2026-06-25 10:33:30 +08:00
can1357 0eb21efa1a Merge remote-tracking branch 'origin/farm/7ef98714/snapcompact-copilot-vision-gate' 2026-06-24 21:00:37 +02:00
can1357 0f2737f16e Merge remote-tracking branch 'origin/farm/e322f828/eval-agent-yield-terminal' 2026-06-24 20:59:54 +02:00
roboomp 997b2b24ff fix(session): suppressed empty-stop retry after successful yield
Trailing empty assistant 'stop' arriving after a successful 'yield'
revived the already-yielded subagent. AgentSession.agent_end maintenance
compared #assistantEndedWithSuccessfulYield(msg) against the trailing
empty-stop message — not the yield-bearing one — so the empty-stop
recovery path appended a retry reminder and scheduled agent.continue().

Track a sticky #yieldTerminationPending flag set when the yield tool
finishes without error and cleared on the next #promptWithMessage. The
agent_end routing extends the existing successful-yield branch: when the
flag is set, or the current message ended with yield, short-circuit
empty-stop / unexpected-stop / compaction continuations for the rest of
the run, so a successful yield is terminal regardless of trailing stops.

Fixes #3389
2026-06-24 17:08:49 +00:00
roboomp 714051d795 fix(catalog,coding-agent): disable vision on non-personal copilot endpoints
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.

- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
  forces input=['text'] whenever the resolved baseUrl is not the
  canonical personal-Copilot host, so the upstream's vision flag is
  honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
  OR-upgrading with the bundled reference) when the merged baseUrl
  differs from the bundled one, so a bundled spec pinned to the
  personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
  rasterizer for any github-copilot model whose baseUrl is non-personal,
  catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
  pi-catalog/wire/github-copilot so catalog and coding-agent share one
  canonical check.

Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).

Fixes #3387
2026-06-24 16:27:14 +00:00
can1357 9c56956631 fix(tui): kept recent turns visible when collapsing remote compactions
With collapseCompactedHistory the live display fell into the LLM compaction
branch, which skips the firstKeptEntryId..compaction turns whenever an OpenAI
remote-compaction replacementHistory payload is present. That payload feeds the
provider only and is not rendered, so a remotely-compacted session showed just
the summary plus post-compaction rows, hiding recent turns that were visible
before. Emit the kept SessionEntry rows in transcript mode regardless. Adds a
regression.
2026-06-24 18:26:20 +02:00
can1357 c04747c5ac Merge PR #3259: fix(tui): reduce large transcript stalls (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 3117a20ab9 Merge PR #3249: fix(agent): size snapcompact maxFrames by the live model window (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 ef67f685ac Merge PR #3380: fix(coding-agent): clamp auto thinking to undefined for models without controllable effort (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 2b42f19456 Merge PR #3232: fix(agent): clamp provider context images (@roboomp) 2026-06-24 18:23:48 +02:00
roboomp b0e07f52d2 fix(coding-agent): clamped auto thinking to undefined for models without controllable effort
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.

Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.

Fixes #3356
2026-06-24 14:12:18 +00:00
roboomp 24adad2890 fix(providers): honored llama cpp model context
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
2026-06-23 12:23:59 +00:00
can1357 44b62419c5 Merge remote-tracking branch 'origin/farm/fc89a4e3/goal-auto-compaction-still-not-triggering' 2026-06-23 00:01:47 +02:00
roboomp 29e69f5364 style: bun run fix 2026-06-22 20:02:46 +00:00
roboomp a305e68a53 fix(compaction): compact goal runs between tool turns
Active goal loops can stay inside one agent run while the model keeps
emitting tool calls, so the normal agent_end threshold maintenance never
runs. That lets context grow past the soft threshold until provider
overflow or user abort.

Run threshold maintenance from the per-turn onTurnEnd hook for active
goals, splice the compacted agent state back into the live loop message
array, and suppress queued continuations because the current run is
already continuing. Cover the mid-run tool-call path and the non-goal
control case.

Refs #3174
2026-06-22 20:02:33 +00:00
can1357 5c21b28786 feat: optimized handoff generation and harden request safety
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
2026-06-22 20:05:40 +02:00
can1357 1082b6316a Merge remote-tracking branch 'origin/farm/5ce2a200/hide-secrets-blocks-advisor' 2026-06-22 17:44:48 +02:00
can1357 beb61b1831 Merge remote-tracking branch 'origin/farm/f3af36ee/clear-openai-completions-on-switch' 2026-06-22 17:43:34 +02:00
roboomp 45b200cd09 fix(coding-agent): evicted resolved completions session URLs
The completions provider stores session state under the request-time resolved base URL, which can differ from the catalog baseUrl for Moonshot, Alibaba Coding Plan, Azure deployments, and similar provider overrides. The model-switch cleanup now evicts the previous provider prefix whenever the switch leaves that completions backend, so those resolved-url keys cannot survive the switch.
2026-06-22 13:39:48 +00:00
roboomp 596a2d1317 style: bun run fix 2026-06-22 13:30:10 +00:00
roboomp 8cebe98d2f fix(coding-agent): evicted openai-completions provider session state on backend switch
`AgentSession.#closeProviderSessionsForModelSwitch` only handled
`openai-codex-responses` and `openai-responses:<provider>` keys. The
`openai-completions:<provider>:<baseUrl>:<modelId>` entries — which cache
strict-tools disable scopes and reasoning-effort fallbacks tied to the
upstream backend — survived /model switches between different providers or
base URLs, so the next request to that backend (e.g. on /model toggle
back) replayed stale decisions made against an entirely different
transport.

Switching to a model whose `(provider, baseUrl)` differs from the current
openai-completions model now evicts every cached entry sharing the old
prefix. Same-backend model toggles keep their cached state, matching the
existing codex/responses semantics.

Fixes #3260
2026-06-22 13:30:01 +00:00
roboomp 3bcbf1515d fix(tui): reduced large transcript stalls
Tail appended transcript JSONL instead of rebuilding rendered history on every poll, collapse compacted history for live chat rendering, and replace synchronous session rewrites so tailers detect historical changes.

Fixes #3258
2026-06-22 12:01:34 +00:00
roboomp 232994496d fix(agent): size snapcompact cap reserve from live shape's text-edge cost
chatgpt-codex third-pass review on #3249: the 4k SUMMARY_TEXT_RESERVE
in the cap math undersized the actual textHead+textTail cost a frame-
bearing archive carries (the projection separately bills
'countTokens(summary + textHead + textTail)'). At ~120k headroom on
Anthropic 11on16-bw, the cap picked maxFrames=23, but
'23 * 5024 + 2 * 13916 chars (≈7k tokens) + 2k summary template ≈ 124.5k'
still exceeded the same 120k headroom — the cap chose a value the
projection then immediately rejected, re-opening the warning loop.

#computeSnapcompactMaxFrames now resolves the live snapcompact shape
(same call the auto/manual paths pass to snapcompact.compact) and sizes
the cap reserve from 'geometry(shape).capacity':

  textEdgeTokens = ceil(2 * capacity * 1.15 / 4)   // 1.15 absorbs
                                                   // tokenizer drift
  capReserve     = textEdgeTokens + 2000           // + summary template

For the default per-provider winners that resolves to ~10k (Anthropic
Sonnet), ~14k (Opus 4.7), ~16k (Gemini 2.x), and ~10k (OpenAI) — all
larger than the prior fixed 4k. Skip decision stays separate
(baseTokens >= totalBudget), so positive sub-reserve headroom still
runs snapcompact's text-only path.

Test 1 retuned to baseline kept-recent ≈ 100k tokens with a strengthened
assertion verifying the FULL projection invariant (frames + worst-case
text edges + summary template + base ≤ budget). Confirmed test fails
against the previous 4k-reserve helper by exactly the reviewer's
predicted margin (174,271 vs 170,000 budget = 4,271 token overshoot).
2026-06-22 10:38:24 +00:00