Commit Graph
1646 Commits
Author SHA1 Message Date
can1357 ac6cd57bd4 Merge PR #6820: fix(coding-agent): preserve parent todos in vibe mode (@Iron-Ham) 2026-07-28 10:59:37 +02:00
can1357 be0e593428 Merge PR #6880: fix(session): preserve auto-thinking level on classifier failure (@roboomp) 2026-07-28 10:59:37 +02:00
can1357 2b759b89e3 Merge PR #6838: fix(session): attribute user bang executions in advisor transcripts (@roboomp) 2026-07-28 10:59:36 +02:00
can1357 f464c9b168 Merge PR #6829: fix(cli): purge completed async jobs on /new (@roboomp) 2026-07-28 10:59:36 +02:00
roboomp 5d9cbf19e6 fix(session): preserved auto-thinking level on classifier failure
- Kept the last successfully classified effort when a later classification fails.
- Added regression coverage for the success-then-failure transition.

Fixes #6877
2026-07-28 08:06:17 +00:00
can1357 641d20fa5c refactor(coding-agent/advisor): removed stale-review-window warning from advisor notes
- Remove `annotateForStaleness` and `hasFreshBacklog` from the advisor runtime.
- Stop appending staleness warnings to delivered advisor notes when newer primary turns queue.
2026-07-28 04:25:41 +02:00
can1357 21b3764b08 refactor(coding-agent): replaced xdevregistry with state interface and helpers
- Replaced the `XdevRegistry` class with the `XdevState` interface and pure helper functions across core and session tools.
- Updated session configurations, tool execution, and renderers to utilize canonical tool map initialization and sharing.
- Adapted unit tests and mocks to use `XdevState` and associated helper functions for permission and dispatch verification.
2026-07-28 03:34:36 +02:00
can1357 bbe0236a0f fix(coding-agent): prevented duplicate artifact saves and mismatched ids in bash output 2026-07-28 02:05:55 +02:00
roboomp 550894de9a fix(session): attribute user bang executions in advisor transcripts
User-initiated `!`/`$` (and RPC bash) executions persist as
`bashExecution`/`pythonExecution` roles, which are always user-run —
the model's bash goes through `toolCall`. The history serializer emitted
them as a bare `→ bash! …` line and cleared the watched-mode role label,
so in advisor deltas the command trailed the `**agent**:` block with no
user attribution. The advisor then misread user-run commands as agent
actions (e.g. a false blocker over a user-run worktree cleanup).

Render these lines under a `**user**:` label in watched mode and prefix
them `→ user-bash!` / `→ user-python!` so provenance is explicit in every
render path.

Fixes #6837
2026-07-27 23:39:31 +00:00
can1357 cfd335d2b1 fix(tui): restored image attachments on /tree and esc-esc branch
- branch() and navigateTree() now return the selected user message's image
  parts (selectedImages/editorImages) alongside the text, extracted in marker
  order by #extractUserMessageImages.
- CustomEditor.setDraft() replaces the composer draft with text plus its
  pending images, so restored [Image #N] markers resolve on resubmit instead
  of degrading to literal text.
- Wired all six restore call sites (selector-controller, extension-ui-controller)
  through setDraft; updated rpc-subagents mocks for the new branch shape.
- Added offline regression tests for branch/navigateTree image restitution,
  multi-image marker order, and text-only prompts.
2026-07-28 00:57:34 +02:00
can1357 392aac4e49 fix(coding-agent): resolved third review pass on inspect_image vision mode
- read now treats an xd://-mounted inspect_image as available (top-level
  predicate OR mounted device gated by the effective mode), so default
  xdev sessions with a text-only model keep metadata-guidance reads
  instead of inlining images the provider boundary would scrub
- advisor tool session stops inheriting the primary's isToolActive and
  xdevRegistry: advisors cannot execute xd:// devices, so their reads
  inline images again
- setModelWithProviderSessionReset is now async and awaited at every
  callsite, so retry-fallback model switches cannot race the
  inspect_image tool-slate reconcile
- regression tests for both xd:// availability directions
2026-07-27 23:07:26 +02:00
can1357 1bba24f191 Merge PR #6830: feat(coding-agent): capability-aware inspect_image with tri-state mode and /vision toggle (@epsilver) 2026-07-27 23:07:26 +02:00
can1357 da6d11de0e feat(agent): introduced prepareToolCall phase supporting argument replacement
- Added a prepareToolCall phase to the agent loop running before tool scheduling for validation and hooks.
- Updated BeforeToolCallContext and result types to support argument replacement instead of in-place mutation.
- Updated coding-agent extension handling and runner to track emitted tool calls and re-evaluate approvals on input revisions.
- Added comprehensive test coverage for argument replacement, concurrency resolution, and schema validation.
2026-07-27 22:55:20 +02:00
alexis@epsilver.xyz b58a943a8e fix(coding-agent): address second codex pass on vision mode
- read now derives its image behavior from actual tool availability
  (session.isToolActive) with the mode computation as fallback, so
  restricted sessions whose explicit slate omits inspect_image (e.g.
  subagents) never get metadata-only reads pointing at an absent tool
- reconcile passes the post-change availability into the read
  description sync, keeping the advertised prompt correct across flips
  in both directions and when tool construction fails
- flat quoted-dotted inspect_image.mode is normalized into the nested
  target during migration instead of being silently dropped when a
  legacy flat enabled key is present
- regression tests for all three: availability-driven read behavior,
  flat+flat migration, description advertising
2026-07-27 16:24:02 -04:00
alexis@epsilver.xyz 3854c1c3b1 fix(coding-agent): address codex review on vision mode
- Reconcile inspect_image centrally from setModelWithProviderSessionReset
  so retry-fallback model changes (turn-recovery.ts) that bypass
  syncAfterModelChange cannot leave a stale tool set
- Apply persisted inspect_image.mode changes immediately from the
  settings selector via a new handleSettingChange branch
- Refresh the read tool's advertised description during reconciliation,
  before applyActiveToolsByName rebuilds the prompt, instead of only
  lazily on the next image read
- Fix the flat (quoted-dotted) enabled->mode migration to write the
  nested target form the resolver actually reads
- Add committed regression tests: tri-state x capability matrix,
  override precedence, and enabled->mode migration (nested, flat, and
  explicit-mode-wins)
2026-07-27 16:24:02 -04:00
alexis@epsilver.xyz c33b98e260 feat(coding-agent): capability-aware inspect_image with tri-state mode and /vision toggle
Replace the inspect_image.enabled boolean with inspect_image.mode
(auto|on|off, default auto). In auto the tool is registered only when
the active model lacks native image input, so vision-capable models
(e.g. kimi-code/k3) read images inline with their own capabilities
instead of delegating to a separate vision model. on/off force
registration regardless of model capability.

- New utils/inspect-image-mode.ts resolves the effective state from the
  /vision session override, the persisted setting, and model capability
- read tool re-evaluates the effective state per image read and
  re-renders its description, so it returns decoded image blocks again
  whenever inspect_image is hidden
- /vision [on|off|auto|status] slash command (modeled on /computer)
  overrides the mode for the current session only
- Tool set is reconciled on model switch with a status notice when
  inspect_image appears/disappears
- Legacy inspect_image.enabled true/false migrates to mode on/off
2026-07-27 16:24:02 -04:00
roboomp e4a586ac08 fix(cli): invalidated async deliveries by generation
- Stamped a per-owner delivery generation on each async-result follow-up.
- Bumped the generation on session transitions and dropped stale-generation
  deliveries at format-time and flush-time, closing the job-id-reuse race.
- Reverted the fragile persistent-suppression eviction to full row removal.
- Covered the reused-id late-delivery contract at the transcript level.

Fixes #6828
2026-07-27 20:09:26 +00:00
roboomp 3488b92508 fix(cli): dropped pending async deliveries on new session
- Evicted jobs now drop queued/in-flight deliveries and stay suppressed.
- /new clears already-queued async-result yield follow-ups.
- Added a transcript-level test asserting no prior-session result leaks.

Fixes #6828
2026-07-27 19:53:58 +00:00
roboomp 6bc98eb3b9 fix(cli): evicted failed jobs on new sessions
- Included failed terminal jobs in owner-scoped session cleanup.
- Extended the /new regression test to cover failed prior-session jobs.

Fixes #6828
2026-07-27 19:44:28 +00:00
roboomp fd24dcc551 fix(cli): purged prior session async jobs
- Added owner-scoped eviction for retained completed jobs.
- Cleared completed async rows during fresh-session cleanup.
- Covered /new isolation while preserving other owners' jobs.

Fixes #6828
2026-07-27 19:38:29 +00:00
can1357 6b42097d12 perf(agent): implemented caching for session file scans and actions
- Add an LRU cache to `scanSessionFile` in `session-listing.ts` keyed by file path, stat identity, and scan mode.
- Add a match key union probe in `CustomEditor` in `custom-editor.ts` to bypass per-action lookups on plain text input.
- Add tests covering cache hits, size and mtime invalidations, and negative result caching.
2026-07-27 20:33:38 +02:00
Hesham Salman 7f7e742db6 fix(coding-agent): preserve parent todos in vibe mode 2026-07-27 13:52:25 -04:00
can1357 d16a251777 chore: reorg tests 2026-07-27 16:43:53 +02:00
can1357 137bad2c8b style: applied biome formatting to review follow-up changes 2026-07-27 16:17:04 +02:00
can1357 8c5dc16344 fix(task): enforced per-spawn effort ceiling across retry fallbacks
- task.maxEffort only clamped the initial thinking level; a retry
  fallback candidate could clamp back up to its model floor and run a
  low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
  ModelControls (constructor, setThinkingLevel, auto classifier,
  restore) and in applyRetryFallbackCandidate; fallback candidates whose
  floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
  attribution added.
- Review follow-up for PR #6794.
2026-07-27 16:09:28 +02:00
can1357 3681faec41 Merge PR #6789: feat(coding-agent): show advisor cost separately in the status line (@paolomazzitti) 2026-07-27 15:57:46 +02:00
can1357 b0063dd180 Merge PR #6787: fix(mcp): deduplicate aliased server connections (@roboomp) 2026-07-27 15:57:44 +02:00
Paolo Mazzitti 9d240ea0ad feat(coding-agent): show advisor cost separately in the status line
Render the Advisor spend next to the primary-model cost as `$2.67 (sub) + $0.41 (adv)`, leaving the status line unchanged until an Advisor cost exists.

Record the cost from finalized advisor `message_end` events in a per-session ledger instead of deriving it from the live advisor transcript, so an in-session compaction or any other history rewrite no longer resets the reported spend. The ledger is cleared for a new session and once a different-session switch commits, and survives a switch that rolls back.
2026-07-27 13:52:22 +00:00
can1357 ae01a76136 fix(agent): hardened pre-model-call gate state cleanup and API surface
- Cleared the retained soft-requirement lifecycle alongside the deferred
  hard choice: clearDeferredToolDirectives() owns both, is called from
  clearAllQueues/reset and session-scoped tool-state cleanup, with a
  regression covering reminder re-injection after a queue clear.
- Allowed void-returning pre-model gates via the named AgentBeforeModelCall
  type and normalized gate results in the loop and Agent dispatcher.
- Documented that the first gate installed mid-run applies from the next
  run; corrected the onToolChoiceRejected contract docs; documented the
  cross-run lifetime of ToolChoiceQueue's in-flight claim.
- Removed the unused addBeforeModelContextBuild hook.
- Relocated both packages' changelog entries out of the released 17.1.4
  sections into Unreleased with PR attribution, folding the never-shipped
  Fixed bullet into Added and noting the input-event timing change.
2026-07-27 14:08:53 +02:00
can1357 6bbfc1110a Merge PR #6543: feat(agent): add a pre-model-call gate that can stop the turn (@paralin) 2026-07-27 14:01:11 +02:00
roboomp da11d906ff fix(mcp): unified tool collision handling
Moved first-wins MCP tool-name deduplication and origin-aware warnings into one shared helper used by startup extension registration, SDK custom-tool assembly, and deferred refreshes.

Added an SDK startup regression proving colliding MCP proxy tools keep the first origin instead of silently overwriting it.

Fixes #6786
2026-07-27 10:49:59 +00:00
roboomp ab8fb13eea fix(mcp): deduplicated aliased server connections
Deduplicated semantically identical MCP endpoints across provider-specific names while preserving provider priority and canonical direct names.

Kept the first registration on sanitized tool-name collisions and logged both origins.

Fixes #6786
2026-07-27 10:22:39 +00:00
can1357 733f20d963 style: formatted scheduleAgentContinue signature per biome 2026-07-27 05:00:25 +02:00
can1357 12a2b13e93 fix(session): unwedged auto-retry after assistant-tail removal miss
A context rebuild that recreated the failed turn's message object made the
identity-keyed active-context removal miss, so the scheduled retry
continuation rejected the terminal assistant error message locally
("Cannot continue from message role: assistant") before any provider
request. auto_retry_end never fired, retryPromise stayed pending, and the
in-flight prompt() plus the TUI retry indicator hung until a manual
follow-up.

The retry path now strips a still-failed assistant tail positionally after
the backoff (generation-guarded, never in preserveFailedTurn mode), and a
continuation that still fails locally closes the retry saga with a failed
auto_retry_end via the new scheduleAgentContinue onError hook.

Fixes #5382
2026-07-27 04:59:31 +02:00
can1357 e07db86f2d Merge PR #6670: fix(mcp): map mounted tools to xd routes (@jeffscottward) 2026-07-27 04:58:27 +02:00
can1357 9d76e168ba Merge PR #6706: fix(ai): preserve custom Anthropic web-search history (@roboomp) 2026-07-27 04:58:24 +02:00
Christian Stewart 01ecab6df7 fix(agent): clear deferred choices on branches
A pre-model gate can defer a claimed hard tool choice for the next call. Branch transitions cleared the coding-agent queue but left that agent-owned value alive, allowing an obsolete forced tool to cross into the replacement transcript.

Expose the narrow deferred-choice reset at the Agent owner and invoke it from the shared session-scoped tool-state cleanup used by both branch paths. Failed session switches retain their existing rollback behavior.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:05:30 -07:00
Christian Stewart e79eabc3b6 feat(agent): add a pre-model-call gate that can stop the turn
The agent loop had no place to refuse a provider request. A host that needs to
act on the assembled context before it is billed, checking that the prompt still
fits the window, that a budget boundary has not been crossed, or that the
session should hand off instead of spending, could only observe the request
after the fact, when the tokens were already committed.

Add `AgentLoopConfig.beforeModelCall`, asked once per turn beside the deadline
check and before `turn_start` is emitted. A `stop` result ends the stream with
no turn open, so nothing has to synthesise a cancellation event and no consumer
is left holding a half-open turn. Placing it there also keeps `turn_end`'s
contract intact: that event carries the assistant message for a completed turn,
and a gated stop has no assistant message to report.

`syncContextBeforeModelCall` keeps its existing void contract and its job of
refreshing prompt and tool state, so implementations typed as returning void are
unaffected.

`Agent.setBeforeModelCall` installs the host's callback, and `addBeforeModelCall`
registers an additional callback without displacing the host's, returning a
disposer so an extension can attach and detach independently. A supplied
`reason` is logged where the loop stops.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:02:18 -07:00
can1357 b646234c41 Merge PR #6646: fix(secrets): avoid hashline placeholder collisions (@roboomp) 2026-07-26 15:45:09 +02:00
can1357 60b0968e1f Merge PR #6484: fix(coding-agent): resume agent after /tree ask re-answer (@roboomp) 2026-07-26 15:41:31 +02:00
can1357 ca4f0d6c56 Merge PR #6494: fix(session): cancel in-flight session_stop handlers (@roboomp) 2026-07-26 15:41:28 +02:00
can1357 e758f10a96 Merge PR #6660: fix(prewalk): apply same-model effort downgrades instead of skipping (@roboomp) 2026-07-26 15:41:28 +02:00
can1357 c5a819b162 Merge PR #6669: fix(session): retry reasonless tool-call aborts (@roboomp) 2026-07-26 15:41:28 +02:00
can1357 7bbe140914 Merge PR #6626: fix(advisor): propagate advisor provider session id via metadata (@roboomp) 2026-07-26 15:38:52 +02:00
can1357 2cc6b28d98 Merge PR #6537: fix(advisor): bound Codex SSE retries (@jeffscottward) 2026-07-26 15:38:52 +02:00
roboomp 2fb9f888ad fix(ai): preserved custom anthropic web-search history
Retained complete native web-search call/result pairs through leaked-thinking projection at their original source anchors.

Kept opaque server-tool blocks atomic during persistence, stripped them on reparent, and covered custom continuations plus compaction retention.

Fixes #6703
2026-07-26 13:32:08 +00:00
Jeff Scott Ward fcdd33d2fe fix(mcp): map mounted tools to xd routes 2026-07-26 01:02:11 -04:00
roboomp 0988efbce6 fix(session): recover cursor reasonless tool aborts
Scoped the Cursor server-execution marker gate to the stream-stall path so an unmarked client-side Cursor tool call (todo/MCP) followed by a reasonless abort recovers via its synthetic executed:false result instead of settling the turn.

Fixes #6668
2026-07-26 04:33:34 +00:00
roboomp 464eba46b3 fix(session): retried reasonless tool-call aborts
Continued from synthetic unexecuted tool results when a reasonless request abort arrives after a complete streamed tool call. Preserved deliberate user, lifecycle, and streaming-edit guard abort behavior.

Fixes #6668
2026-07-26 04:26:53 +00:00
roboomp 583ff590f2 fix(prewalk): apply same-model effort downgrades instead of skipping
The prewalk arm/switch guard compared model identity only (modelsAreEqual /
provider+id), discarding the resolved thinkingLevel. A legal same-model target
at a cheaper effort (e.g. prewalk: "@task" resolving to the active model at a
lower level) was dropped as a no-op, so the session ran the expensive effort for
the whole run while still paying the plan/continue nudges — silently on the
session path, logger.debug only on the subagent path.

Compare (provider, id, effective thinking level) via a shared prewalkWouldBeNoop
helper. Effort-only deltas on the same model now switch; a genuine no-op emits a
user-visible notice on the session path and never arms on the subagent path.

Fixes #6659
2026-07-26 02:20:03 +00:00