Commit Graph
1248 Commits
Author SHA1 Message Date
can1357 82327af420 fix(coding-agent): require terminal user prompt for todo suppression 2026-07-14 18:45:03 +02:00
can1357 f1ade56f65 Merge PR #5097: fix(coding-agent): suppress todo reminders for interactive questions (@roboomp) 2026-07-14 18:45:02 +02:00
can1357 1dfec0d161 Merge PR #5090: fix(advisor): reduce stale advisories via delta coalescing, WIP markers, and delivery annotation (@apoc)
# Conflicts:
#	packages/coding-agent/src/advisor/__tests__/advisor.test.ts
#	packages/coding-agent/src/advisor/runtime.ts
#	packages/coding-agent/src/session/agent-session.ts
2026-07-14 18:44:57 +02:00
can1357 280d4b039b fix(coding-agent): persist pruned autolearn captures 2026-07-14 18:41:16 +02:00
can1357 af9e9a898b Merge PR #5217: fix(coding-agent): accept Auto-Learn capture empty stops (@roboomp) 2026-07-14 18:41:15 +02:00
can1357 645cf45ed6 fix(coding-agent): preserve all terminal advisor notes 2026-07-14 18:41:15 +02:00
can1357 b52a4cdede Merge PR #5186: fix(coding-agent): preserve late advisor notes after terminal answers (@roboomp)
# Conflicts:
#	packages/coding-agent/test/agent-session-advisor-suppression.test.ts
2026-07-14 18:40:35 +02:00
can1357 fcad4f22b9 Merge PR #5183: fix(advisor): quarantine unknown tool responses (@roboomp) 2026-07-14 18:39:54 +02:00
can1357 72b1ddf04c fix(session): validate pending exit diagnostics 2026-07-14 18:39:53 +02:00
can1357 7773f48ebc Merge PR #5168: fix(session): recover interrupted session turns (@paralin) 2026-07-14 18:39:53 +02:00
can1357 58b80dac42 Merge PR #5139: fix(coding-agent): flush throttled output tails (@wolfiesch) 2026-07-14 18:39:26 +02:00
can1357 da24614d5a feat: added scrollback rebuild controls and prewalk status-line visibility
- Added `tui.scrollbackRebuild` configuration with interactive startup/controller wiring to apply `setScrollbackRebuild`.
- Exposed prewalk session state in `SegmentContext` and rendered a dedicated prewalk segment/icon in the status line.
- Added divergence-aware TUI full-paint logic that enables scrollback erase-and-replay rebuilds for non-multiplexer divergence cases.
- Updated rendering and streaming tests to verify rebuild behavior (`3J`) and eliminate stale marker expectations under drift scenarios.
2026-07-14 00:39:34 +02:00
can1357 f9f6ed9e8d feat(coding-agent): replaced legacy pi/ role alias prefix with
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
2026-07-13 23:26:33 +02:00
can1357 a5673c90f8 feat(ai): removed legacy Google interactions routing from AI providers
- Removed the Google Interactions transport and deleted interaction-specific request options from the shared AI stream typing/API surface.
- Simplified Google provider routing to eliminate interactions auto-selection logic and keep `streamGoogle` on the `:streamGenerateContent` path.
- Updated Vertex request handling to use resolved stream hosts without `/interactions`/`Api-Revision` and removed related interaction constants.
- Deleted obsolete Interactions tests and updated remaining Google stream tests to no longer reference `useInteractionsApi`/`storeInteraction`/`previousInteractionId`.
2026-07-13 18:43:52 +02:00
can1357 4df6f6683d chore: finalizing the new /prewalk 2026-07-13 15:18:31 +02:00
can1357 f405525bf4 feat: removed boomerang feature and associated validation workflows
- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
2026-07-13 14:15:54 +02:00
can1357 590270ca26 feat(coding-agent): gated downshift trigger on todo initialization
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
2026-07-13 10:37:20 +02:00
can1357 95ecc61bc9 feat(coding-agent): implemented downshift boomerang flow for context handoff
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
2026-07-13 10:37:20 +02:00
can1357 4d019e5617 feat(coding-agent): trigger downshift on post-plan todo initialization
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.
2026-07-13 10:36:39 +02:00
can1357 9f1ff90a39 feat(coding-agent): implemented downshift and plan-yolo agent workflows
- Replaced legacy reasoning-slide functionality with new downshift and plan-yolo capabilities.
- Updated CLI arguments, slash commands, and configuration schemas to support the new model-switching and execution behaviors.
- Refactored agent session logic to handle downshift arming, plan-yolo headless execution, and context scrubbing.
- Renamed and added system prompts to align with the updated downshift and plan-yolo workflows.
2026-07-13 06:03:46 +02:00
can1357 0a98aa252b feat(coding-agent/modes): hardened tan fork isolation and session sync
- Clear inherited todo list state and persist empty edit at fork creation to prevent parent task reminders from affecting the tangential session.
- Re-inject the fork notice after each auto-compaction event to ensure the boundary between the parent and child session survives history summarization.
- Align the provider cache key with the parent's actual pinned key to correctly mirror cached session context.
- Update `AgentSession` to perform a full entry rewrite during tool result pruning to ensure session files match pruned state for reliable resuming and branching.
2026-07-13 06:00:30 +02:00
can1357 ac16253613 wip: rslide (experimental) 2026-07-13 04:21:54 +02:00
can1357 22f2c1947b test(session): validated auto-compaction guard and recovery behavior
- Added test coverage for auto-compaction entry warning attribution during dead-end scenarios.
- Verified automatic continuation and warning suppression when image-drop recovery successfully clears context pressure.
- Mocked session context usage and image-drop functionality to simulate high-usage state recovery.
2026-07-13 01:41:09 +02:00
can1357 aa52fa4231 feat(session): enabled visibility for elided tool calls in transcripts
- Added `StrippedToolCallsMarker` to identify assistant messages with dangling tool calls that were removed during context building.
- Updated `buildSessionContext` to preserve turn entries in transcripts even when all content is stripped, marking them with the count of elided calls.
- Updated `UiHelpers` to render a dim, italicized placeholder in the TUI when an assistant message has elided tool calls, preventing silent activity gaps.
2026-07-13 01:40:15 +02:00
can1357 4903a13511 feat(session): implemented automated recovery and status reporting
- Added `warning` field to `CompactionEntry` and `CompactionSummaryMessage` to persist dead-end status in session history.
- Introduced a multi-tier rescue mechanism in `AgentSession` that automatically performs `elide` and `dropImages` passes when maintenance fails to recover sufficient headroom.
- Integrated visual indicators for dead-ends into the `CompactionSummaryMessageComponent`, surfacing warnings directly on the compaction divider and detail block.
- Updated `AgentSession` logic to re-evaluate progress after each rescue tier and emit recovery notices, ensuring transparent reporting of automated history rewrites.
2026-07-13 01:40:14 +02:00
can1357 bf5eb3769f feat(coding-agent): rebased pending context snapshot after compaction
- Update in-flight context snapshots after historical messages are modified by compaction or elision.
- Prevent stale run-start token counts from triggering false-positive dead-end "no progress" warnings during auto-continuation.
- Add regression test to verify that prompt headroom measurements correctly account for post-compaction context sizes.
2026-07-13 01:28:18 +02:00
can1357 58d6130b50 feat(coding-agent): enabled model fallback for hard errors
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
2026-07-13 00:59:43 +02:00
can1357 20c0a2e410 test(coding-agent): fixed storage tests for schema v6 and slow CI disks
- Compat tests assert the exported SCHEMA_VERSION instead of a stale
  hardcoded 5; the perf-sample change bumped agent.db to v6.
- Backfill cap fixture inserts its 300 rows in one transaction; per-row
  implicit transactions fsynced 300 times and timed out on CI runners.
2026-07-12 03:13:15 +02:00
can1357 d54dcc2224 feat(coding-agent): allowed model fallback after retry budget exhaustion
- Permit model fallback even if the retry budget is exhausted when the current provider is locked by credential rotation or usage limits.
- Reset the retry budget when successfully switching to a fallback model to ensure the new model has a full allowance of retries.
- Added a regression test to verify that credential rotation failure triggers model fallback.
2026-07-12 01:35:50 +02:00
Christian Stewart bdad9ca8c0 fix(session): recover interrupted switched sessions
Close pending tool turns recorded during normal shutdown and apply interrupted-tail recovery when switching or reloading sessions.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-11 16:18:39 -07:00
Christian Stewart a420dd11e4 fix(session): preserve terminal failed tool turns
Skip interrupted-turn recovery when an error or aborted assistant already closed its tool calls with synthetic results.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-11 16:18:39 -07:00
Christian Stewart de236903d4 fix(session): close all interrupted transcript tails
Detect pending tool calls from assistant content and use the selected model metadata to terminate first-turn user tails after abnormal process exits.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-11 16:18:39 -07:00
Christian Stewart e3e5bf8775 fix(session): recovered interrupted session turns
Appended one terminal aborted assistant record when a persisted abnormal-exit diagnostic follows a non-terminal transcript tail, preserving partial history while making resumed context valid.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-11 16:18:39 -07:00
can1357 c4fa0ebaae feat(storage): implemented persistent model performance tracking and migration
- Added `model_perf` table to store persistent, recency-weighted TPS and TTFT statistics per model.
- Introduced `AgentStorage.recordModelPerf` to asynchronously capture and aggregate turn timing metrics.
- Implemented a background backfill mechanism to migrate historical request data from `stats.db` into new performance aggregates.
- Updated `AgentSession` to record performance metrics upon turn completion.
- Bumped `SCHEMA_VERSION` to 6 and cleared existing performance data to account for improved duration-based TPS calculation.
2026-07-12 01:06:19 +02:00
can1357 a3117c2842 feat(agent): migrated async-drain utility to shared package for reuse
- Relocated `AsyncDrain` class from `coding-agent` to `utils` package.
- Updated `HistoryStorage` to import `AsyncDrain` from shared utilities.
- Centralized the utility to allow reuse across the codebase.
2026-07-12 01:06:17 +02:00
can1357 54bafa1cce feat(coding-agent): implemented interactive fallback chain configuration
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.
2026-07-11 22:03:35 +02:00
can1357 7957d0b88c Merge remote-tracking branch 'origin/farm/3b0815b6/fix-btw-gpt-luna-esc' 2026-07-11 20:38:37 +02:00
can1357 6c292b97c3 refactor(coding-agent): preserved completed and abandoned tasks in session
- Updated task synchronization to include all todo phases regardless of completion status.
- Removed logic that filtered out completed or abandoned tasks during session synchronization.
2026-07-11 20:35:43 +02:00
roboomp c36f64d4f0 fix(coding-agent): pruned noop autolearn captures
Removed accepted empty Auto-Learn capture stops from live and persisted context, including the hidden capture prompt when it directly parents the no-op stop.

Extended the regression to prove the next real prompt does not replay the stale Auto-Learn nudge.

Fixes #5211
2026-07-11 17:46:44 +00:00
roboomp 9269823d16 fix(coding-agent): shared codex state for btw websockets
- Forwarded the session provider state map to ephemeral side-channel turns so Codex can allocate websocket transport state for the side request session id.
- Extended the /btw side-channel regression to assert providerSessionState reaches the stream options with the websocket preference.

Fixes #5213
2026-07-11 17:46:28 +00:00
roboomp 2c161d2a8f fix(coding-agent): preserved btw codex websocket routing
- Preserved the session websocket preference for ephemeral /btw side-channel turns so Codex websocket-only models do not fall back to SSE.
- Prioritized active /btw and /omfg panels in Esc handling before loop, maintenance, and main-turn interrupts.
- Added regression coverage for side-channel websocket options and /btw Escape priority.

Fixes #5213
2026-07-11 17:31:37 +00:00
roboomp 221196704a fix(coding-agent): cleared autolearn empty-stop override
Cleared the terminal empty-stop acceptance flag for every agent-initiated prompt so non-opt-in custom turns cannot inherit Auto-Learn capture behavior.

Added a regression covering a non-opt-in custom turn after a non-empty Auto-Learn capture response.

Fixes #5211
2026-07-11 17:20:46 +00:00
roboomp e2afc87a40 fix(coding-agent): accepted autolearn empty stops
Marked Auto-Learn capture turns as accepting terminal empty assistant stops so the empty-stop guard does not retry the expected no-reply completion.

Added a focused AgentSession regression covering the auto-continue capture turn.

Fixes #5211
2026-07-11 17:04:16 +00:00
Miroslav Drbal 59017f2616 fix(advisor): address final review: wip return type, import order, cap edge case, cap test
Blockers from final review:
- runtime.ts:457 TS2741: #collectAndMaintainBatch now returns wip in its
  result type; #drain destructures it and passes it to the retry-requeue
  unshift — a WIP batch that fails retry no longer silently loses its
  [in progress] heading.
- agent-session.ts import order: annotateForStaleness moved before
  formatAdvisorBatchContent (biome enforces case-insensitive alpha order).

Cap edge case (advisor nit + review CONCERN): final round of the
  coalescing for-loop now breaks BEFORE the late-item splice so any
  deltas that arrived during round MAX_COALESCE_ROUNDS-1's
  maintainContext call stay in #pending for the next drain iteration
  instead of being merged into an unbudgeted batch. Doc comment updated
  to match ('left for the next drain iteration' is now accurate).

Cap test: bounded version that only pushes new turns for the first 3
  maintainContext calls so the drain while-loop terminates cleanly after
  a second iteration. The previous unbounded version created an infinite
  drain loop (each maintainContext call unconditionally pushed another
  turn) and timed out.
2026-07-11 18:59:11 +02:00
Miroslav Drbal 441197b462 fix(advisor): address review: wip on PendingDelta, MAX_COALESCE_ROUNDS cap, annotateForStaleness extraction
Blocker: wip field added to PendingDelta and threaded into the reprime
  path. #collectAndMaintainBatch now captures the most-recent WIP state
  from each delta batch and forwards it to #renderDelta in both the
  normal and reprime branches, so a willContinue:true turn never loses
  its [in progress] heading through a reprime.

Safety cap: MAX_COALESCE_ROUNDS=3 constant defined and used in the
  coalescing for-loop, preventing indefinite dispatch stall under
  pathological fast-primary + slow-maintainContext conditions.

Testability: annotateForStaleness extracted as an exported pure function
  in advise-tool.ts and used in AgentSession#routeAdvice. Three unit
  tests added in advisor.test.ts covering the no-staleness, staleness,
  and note-preservation contracts. This addresses the ReviewSession
  regression-test concern without requiring a full AgentSession harness.

Reprime turn-tally coverage: new test 'backlog stays accurate when a
  delta arrives during the reprime-triggering maintainContext' asserts
  runtime.backlog === 0 after all three turns, catching a deleted
  turns += reduce(...) line.

Fragile double-await tests converted: sends-batch-when-maintenance-fails
  and expands-plan-mode-context now use Promise.withResolvers signals
  instead of counted await Promise.resolve() hops.

Docs: onTurnEnd JSDoc added; hasFreshBacklog comment broadened to cover
  all drain-busy phases (not just agent.prompt).
2026-07-11 18:59:11 +02:00
Miroslav Drbal 74715f8cca fix(advisor): reduce stale advisories via delta coalescing, WIP markers, and delivery annotation
Three related changes that address the pattern of the advisor flagging things
the primary already fixed:

Fix 1 — coalesce late-arriving deltas before agent.prompt (runtime.ts)
  Refactored #drain into a reusable #collectAndMaintainBatch helper that loops
  until the pending queue is stable (no new deltas arrive during a maintenance
  check) before calling agent.prompt. Previously, any turn queued during the
  maintainContext await was deferred a full extra model-call cycle; now it is
  merged into the current batch after re-checking the token budget for the
  expanded payload. Every await in the loop has an epoch guard so a
  reset/dispose mid-await cannot leak a stale batch. finalTurns always counts
  all merged turns so #backlog decrements correctly.

Fix 2 — hasFreshBacklog + delivery-time staleness annotation (runtime.ts, agent-session.ts)
  Added AdvisorRuntime.hasFreshBacklog getter (true when #pending.length > 0
  while agent.prompt is running — i.e., newer primary turns arrived after the
  reviewed window). #routeAdvice checks it at delivery time and appends a
  lightweight caveat to the note so the primary agent knows to verify before
  acting. Uses #pending.length not #backlog, which is always > 0 mid-call.

Fix 3 — willContinue WIP marker in rendered delta + system prompt (runtime.ts, agent-session.ts, system.md)
  onTurnEnd now accepts { willContinue } and passes it through to #renderDelta,
  which tags the heading '[in progress — more steps follow]' for intermediate
  turns. The agent-session.ts call site passes context.willContinue. The advisor
  system prompt instructs the model to withhold critique on WIP updates.

Also fixed pre-existing inline casts in #renderDelta and #dedupContextMessage
that suppressed the type checker instead of using the narrowing already provided
by the role discriminant.

All 75 advisor tests pass; pre-existing type errors in cursor.ts are unrelated.
2026-07-11 18:56:04 +02:00
can1357 5b20a7dea4 feat(coding-agent): implemented persistence for tool calls during rebuilds
- Enabled persistence of dangling tool calls during transcript rebuilds by adding `keepDanglingToolCalls` configuration.
- Integrated `seal()` logic for pending tool blocks to properly finalize history during idle session states.
- Enhanced UI helpers and session context to preserve assistant turns during streaming or mid-turn rebuilds.
- Added comprehensive unit and integration tests to verify correct tool call tracking and transcript integrity.
2026-07-11 18:19:48 +02:00
roboomp ea5324fb10 fix(advisor): trusted advisor tool result provenance
Included advisor tool-result text in the quarantine source check so legitimate findings from granted read/grep tools are not treated as model-generated contamination.

Kept assistant text out of the source set to avoid laundering prior advisor hallucinations.

Fixes #5181
2026-07-11 13:17:46 +00:00
roboomp 708eafaf8d fix(coding-agent): preserved late advisor terminal notes
Prevented late interrupting advisor findings from waking the primary after a terminal text answer when no queued work remains.

Added regression coverage for the advisor-confirmation path so duplicate primary turns are caught.

Fixes #4840
2026-07-11 12:59:22 +00:00
roboomp a58e8faa09 fix(advisor): quarantined unknown tool responses
Quarantined Advisor assistant turns that request tools outside the granted tool pool before they can enter the Advisor context.

Reset the Advisor runtime after quarantine so the next update re-primes from the primary transcript instead of replaying contaminated private context.

Fixes #5181
2026-07-11 12:47:26 +00:00