Commit Graph
38 Commits
Author SHA1 Message Date
can1357 a5434d7341 Merge PR #8076: fix(advisor): preserve advisor fallback role ownership (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 3b66177abd Merge PR #8226: fix(advisor): accept silent Gemini reviews (@roboomp) 2026-08-11 15:15:44 +02:00
can1357 7e8be71ead Merge PR #7247: fix(advisor): prevent false full replays and preserve cache growth (@cuipengfei) 2026-08-11 15:07:43 +02:00
roboomp 86b8f510c8 fix(advisor): accepted silent Gemini reviews
Allowed advisor streams to treat Google STOP responses without visible content as successful silence while preserving the default retry behavior for interactive agents.

Added provider and advisor-path regressions covering retry counts, system instructions, and the advise declaration.

Fixes #8223
2026-08-11 08:22:54 +00:00
roboomp 121bcb3663 fix(advisor): preserved advisor role fallback ownership
- Passed the known advisor role through retry fallback resolution so model and wildcard keys retain precedence while ambiguous role matches cannot select another role.

- Covered shared-model roles with distinct thinking levels and asserted advisor fallback lifecycle ownership.

Fixes #8075
2026-08-09 13:48:46 +00:00
can1357 7ca140f66e fix(ai): routed policy-rejected accounts through sibling rotation
- Added account-scoped policy error detection to correctly identify Codex cyber-policy rejections.
- Updated credential storage and retry logic to route denied accounts through sibling rotation instead of bypassing it.
- Ensured coding-agent sessions exhaust all sibling accounts before falling back on cyber denials.
- Added comprehensive test coverage for credential rotation and retry behavior on policy errors.
2026-08-08 19:29:49 +02:00
can1357 a11eacc9d6 fix(advisor): drop unreachable status guard 2026-08-05 21:50:23 +02:00
Mantas Vidutis 6e32705e71 fix(advisor): fail a refusal over to the model fallback chain
A classifier refusal returned before the `onTurnError` hook that owns
model fallback, so `AdvisorRuntime` treated one provider's policy verdict
as terminal: `Refusal (cyber)` on the advisor model disabled the advisor
outright even with a fallback chain configured. Its only recovery was
stripping echoed primary reasoning and resending once, which does nothing
for a refusal about the content itself.

Route a refusal that outlives the strip through the same hook the generic
failure path uses, mirroring its epoch guard, session-transition requeue,
and requeue-on-recovery. The cascade walks the chain to exhaustion and
only reports the advisor unavailable once the host runs out of candidates,
matching what turn-recovery already allows for the primary.

Each cascade visits a model at most once. A switch re-arms
`#includeThinking` through `#syncModelIdentity`, so chain keys that point
back at each other (A to B, B to A) would otherwise strip-and-resend
against the same pair forever. A successful turn or a reset starts a
fresh walk.

Also stop `/advisor status` throwing when a live advisor has no roster
entry: `formatAdvisorStatus` guarded only the inactive case before
dereferencing `stats.advisors[0]`, and `#ensureAdvisors` clears
`#advisorStatuses` before repopulating it, so a status call landing in
that window hit `undefined.contextWindow`.
2026-08-04 11:47:54 -07:00
cuipengfei 21d0db72cf feat(advisor): log reset reason on every advisor context re-prime (issue #7226) 2026-08-04 22:36:41 +08:00
cuipengfei 5b25030720 fix(advisor): address review — dedupe obfuscation, restore dedup map on rollback, complete fingerprint fields 2026-08-04 22:36:41 +08:00
cuipengfei b506c8ee65 fix(advisor): split Session update into per-message user messages to keep prompt cache growing
The advisor sends its whole Session update as a single ever-growing user
message. Provider prompt caches are prefix-based: a single user message whose
text keeps growing invalidates the entire message on every turn, so cache_read
stays pinned at the instructions/tools boundary (observed 14491 tokens in
production, 11066 in tests) instead of growing with the session.

Split the update into multiple user messages — one per source message —
delivered via a single Agent.prompt(AgentMessage[]) call, so the provider
caches each appended message incrementally. Verified end-to-end: cache_read
grows 0 -> 11126 -> 11457 -> 11583 across turns with the split, versus pinned
11066 on the old single-message behavior.

- delta-split.ts: pure renderAdvisorDeltaChunks using chunked
  formatSessionHistoryMarkdown (shared toolResultIndex/consumedToolCallIds/
  watchedRoleState) so toolCall/result pairing and role collapsing stay
  byte-identical to the old single-block render (equivalence-tested).
- session-history-format.ts: add HistoryFormatOptions.watchedRoleState so
  chunked renders collapse consecutive same-role messages exactly like the
  single-block render.
- runtime.ts: #prepareBatch does a single dedup+render pass; #drain delivers
  agent.prompt(preparedMessages) (array), falling back to the string.
- Keep field-selective fingerprint (candidate 1) + wip-marker-at-tail
  (candidate 3) as complementary wins.

Tests: advisor suite 209 pass / 0 fail; type check clean; lint clean.
Affected subsets (342 tests) green; full suite hits WSL EMFILE fd limit.
2026-08-04 22:36:41 +08:00
can1357 cc2265f681 feat: use simple form replace when replace is the edit mode 2026-08-03 01:00:51 +02:00
can1357 61d066bd2a Merge PR #6756: fix(coding-agent): suppress WIP advisor non-blockers (@wolfiesch) 2026-08-01 20:14:39 +02:00
can1357 93a9f0d991 Merge PR #7262: fix(coding-agent): refresh context files on plugin reload (@roboomp) 2026-08-01 20:13:38 +02:00
roboomp 55115f4bfd fix(coding-agent): refresh advisor context prompt on reload
Rebuilt advisor runtimes with the rediscovered context files so advisor turns stop evaluating against stale AGENTS.md instructions after /reload-plugins.

Fixes #7258
2026-08-01 11:46:55 +00:00
roboomp 5283c4c99b fix(agent): routed codex v2 compaction through websockets
- Reused the live Codex provider session for WebSocket-first V2 compaction.
- Fell back to SSE V2 on WebSocket transport failure before the existing V1 fallback.
- Propagated the configured WebSocket preference through manual, automatic, and advisor compaction paths.
- Added transport reuse and fallback regression coverage.

Fixes #7198
2026-07-31 21:00:15 +00:00
Wolfgang Schoenberger 26e422a00a fix(coding-agent): suppress WIP advisor non-blockers 2026-07-30 15:18:06 -07:00
can1357 c9a5605b5b fix(advisor): degraded reasoning on provider refusals
(cherry picked from commit 776f244d810ac152eb0bd130b36c8474201d0b85)
2026-07-30 02:01:02 +02:00
can1357 43e5df012e Merge PR #6791: feat(ai): handle Cursor's modern exec wire protocol (@quantmind-br) 2026-07-30 01:48:52 +02:00
Diogo Soares Rodriguesandcan1357 5ab98cded8 fix(cursor): persist resource listings, wire advisor MCP resources, reject unavailable pi edit/write
Three remaining review findings:

- `list_mcp_resources` frames a handler answered now synthesize a
  `list_mcp_resources` block and pair a result derived from the same
  answer sent on the wire; the streamed `ListMcpResourcesToolCall` /
  `ReadMcpResourceToolCall` announcements join the exec-owned set so
  they cannot double-render. No-handler frames still synthesize
  nothing, since nothing ran.

- Advisors receive the same `MCPManager`-backed resource adapter as the
  primary bridge, so their `list_mcp_resources` no longer reports every
  server as empty and `read_mcp_resource` no longer answers `not_found`
  against live connections the advisor shares.

- An unavailable `pi_edit`/`pi_write` answers with the protocol's
  `rejected` variant instead of `error`: refusal and failure are
  separate oneof cases, and a denial reported as an execution error
  invites a retry of an operation that was never permitted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SSWZTe6YA2PX1cqtukZvYi
(cherry picked from commit 47ce936c8df05d6970504af19e5ef7d2e8c38d7b)
2026-07-30 01:43:07 +02:00
usr-bin-roygbivandcan1357 7fff8869a0 fix: skip unauthenticated compaction candidates
(cherry picked from commit fa5f7d73ec1f169f9b9952648195057caa24657e)
2026-07-30 01:42:53 +02:00
usr-bin-roygbivandcan1357 c83e506954 fix(compaction): require native-capable fallbacks
(cherry picked from commit 6ac4efd9d2fe37fb50be6e9ba6664e8924bce602)
2026-07-30 01:42:33 +02:00
usr-bin-roygbivandcan1357 28e738830d fix(compaction): align advisor provider boundary
(cherry picked from commit 771a9917caa2630b96e8fd8f6f8d73d5462242d3)
2026-07-30 01:42:33 +02:00
Diogo Soares Rodriguesandcan1357 59434149d1 fix(cursor): gate resource downloads, fix pi_grep cap and MCP transcript
Download-mode resource reads created and overwrote workspace files
without running a registry tool - the same hole the native `delete`
frame had - so a session that withheld `write`/`edit`, or whose `write`
tier is `deny`/`always-ask`, still had files written. Both frames now
share one grant and one policy check, and the download refuses before
the read so a blocked call never fetches the resource.

`allowNativeDelete` is renamed `allowDirectFileMutation`: it now gates
more than deletion. The primary session derives it from the registry
BEFORE its own rewriting (Cursor moves `edit` out of the tool map and
`write` may be auto-registered later, so reading the map at bridge
construction would misjudge both) and unconditionally, since the bridge
is installed for every session and one that starts on another provider
can switch to Cursor later.

`pi_grep` with a match cap: the local tool windows to 20 files and
suggests `skip`, which `PiGrepExecArgs` cannot express - 100 matches
requested over 25 one-match files returned 20, with the cap reported
unreached. A capped search now reads cap+1 files, so a result landing
exactly on the cap is distinguishable from a clipped one, and
`match_limit_reached` is truthful either way.

`read_mcp_resource` synthesized no transcript block and paired no
result, so a read - including a download that mutates the workspace -
was invisible in the UI and stripped from every rebuilt history. It now
synthesizes a `read_mcp_resource` block (not `read`: the name drives
rendering and prune semantics) and pairs success, not-found and error.

(cherry picked from commit 5ff27a3efe8bec522d9d5dbd7763055eb03eae3b)
2026-07-30 01:42:30 +02:00
Diogo Soares Rodriguesandcan1357 a363f95009 fix(cursor): pair interrupted calls and gate advisor bridge approval
Two independent bugs found in review.

A stream that dies mid-turn takes the terminal-error path: `settleH2`
rejects when the transport closes without `turnEnded`, so the flush on
the success path never runs. `connect_scm` and native todo blocks are
stamped `kCursorExecResolved` at start, so `agent-loop.ts` synthesizes
no placeholder and only their completion frame pairs a result - the call
was left unpaired and its card animating, and `buildSessionContext`
strips a dangling call from every rebuilt transcript. The catch path now
closes open blocks and pairs those server-owned calls with an
interrupted result. Exec-settled MCP blocks are excluded: the dispatch
that ran them owns their result and `drainInFlightDispatches` awaits it,
so pairing here would duplicate against the same id.

Separately, the advisor bridge supplied no `getToolContext`.
`ExtensionToolWrapper` reads the approval mode, per-tool policies and
`autoApprove` only from that execute-time context, so every wrapped
advisor bridge tool resolved as `yolo` with empty policies - a
configured `ask` or `deny` on `edit`/`grep` did not apply to native
frames. Advisors now get the same `ToolContextStore` as the primary
bridge.

Both are covered against the real paths: the interrupted call through
the HTTP/2 fixture server (a helper-level test passes even with the
catch-path flush removed), and approval through real deny policies.

(cherry picked from commit 5ace682578af96708caf88db2d44b9077a1c0e74)
2026-07-30 01:42:28 +02:00
Diogo Soares Rodriguesandcan1357 5779c762de fix(cursor): give advisors the replace-mode edit pi_edit needs
The primary bridge builds a `replace`-mode `EditTool` because
`PiEditExecArgs` carries `old_text`/`new_text` pairs that no other mode
accepts. The advisor roster passed its own instances straight through,
and those follow the session's configured `edit.mode` - `hashline` by
default, whose schema is a single `input` string - so every native
advisor edit failed validation instead of touching the file.

Both bridge-only tools now come from `cursor-bridge-tools.ts`:
`createBridgeEditTool` builds the wrapped `replace` instance, and
`bridgeToolMap` substitutes it into a granted map. The substitution is
gated on `edit` actually having been granted, since the tool is
constructed rather than looked up - handing one to a read-only roster is
the #5680 escalation. The advisor's own loop keeps its instance; only
the exec map is swapped.

(cherry picked from commit e6cf9f8046c595cab9793b4b2a5795d8488d8d22)
2026-07-30 01:42:28 +02:00
Diogo Soares Rodriguesandcan1357 6eaf090ccc fix(cursor): repair pi_edit and close the scoped-grep approval bypass
Three defects the exec bridge shipped with, all found by review.

`pi_edit` never worked. The session removes `edit` from the tool
registry for Cursor so the model is steered to full-file `write`
(8ba0498eb), but that same registry is the bridge's tool source, so the
native frame — which the server sends regardless of the advertised
catalog — resolved nothing and answered `Tool "edit" not available`.
Retaining the instance is not enough either: `PiEditExecArgs` carries
`old_text`/`new_text` pairs, which only `replace` accepts, while the
default mode is `hashline` (`{ input: string }`). `EditTool` now takes
an optional mode, and the bridge resolves a pinned `replace` instance
through its fallback resolver.

A `pi_grep` frame carrying `context` or `limit` escaped the approval
gate. Honoring those needs a per-call tool, and the per-call instance
was built raw while every registry tool is wrapped — so exactly those
calls skipped `tools.approval.grep` and the exec-tier SSH check. Both
callsites now go through one `createBridgeGrepFactory`.

Advisors ignored the same two fields: only the primary session supplied
the factory. They now get it too, gated on the advisor actually holding
`grep` so the factory cannot grant a denied tool.

Also moves the pure Pi arg translation to `providers/cursor-pi-args`.
The legacy shim shares it and is compiled into the bundled virtual
registry, where `./providers/*` cannot match a nested specifier — it
fell through to `Bun.resolveSync`, unsatisfiable under bunfs (#3442) —
and the exec module would have dragged the protobuf graph along.

Verified against real files and the real module graph: `pi_edit` mutates
a temp file, the bundled probe executes the shim's shared module in a
subprocess, and the grep test drives the shared factory. Mutation-
checked: returning a raw tool from the factory, ignoring the pinned edit
mode, dropping the `getTool` fallback, or moving the helpers back to a
nested path each fails a test.

(cherry picked from commit e46ba22b634e449005f7c22b6d0efd19a45ce1f8)
2026-07-30 01:42:23 +02:00
usr-bin-roygbivandcan1357 eeb542b2ef fix(advisor): stop provider fallback on native compaction failure
(cherry picked from commit c636aa6be3f12eed9475bb2e1aa6a03749049523)
2026-07-30 01:42:19 +02:00
Paolo Mazzitti 3c7233af44 fix(coding-agent): scope advisor cost to active session 2026-07-29 07:18:48 +00:00
can1357 641d20fa5c refactor(coding-agent/advisor): removed stale-review-window warning from advisor notes
- Remove `annotateForStaleness` and `hasFreshBacklog` from the advisor runtime.
- Stop appending staleness warnings to delivered advisor notes when newer primary turns queue.
2026-07-28 04:25:41 +02:00
Paolo Mazzitti 9d240ea0ad feat(coding-agent): show advisor cost separately in the status line
Render the Advisor spend next to the primary-model cost as `$2.67 (sub) + $0.41 (adv)`, leaving the status line unchanged until an Advisor cost exists.

Record the cost from finalized advisor `message_end` events in a per-session ledger instead of deriving it from the live advisor transcript, so an in-session compaction or any other history rewrite no longer resets the reported spend. The ledger is cleared for a new session and once a different-session switch commits, and survives a switch that rolls back.
2026-07-27 13:52:22 +00:00
can1357 7bbe140914 Merge PR #6626: fix(advisor): propagate advisor provider session id via metadata (@roboomp) 2026-07-26 15:38:52 +02:00
roboomp f1f118acd0 fix(advisor): rebind identity on every provider-session change
Advisor provider-identity refresh was only invoked from resetSessionState(), so transitions that update the primary identity without re-priming the advisor (branch with skipConversationRestore, fork) left the advisor emitting the previous session id/metadata/telemetry. Move the refresh into #syncAgentSessionId() via a new SessionAdvisors.refreshProviderIdentity() so it fires on every provider-session change regardless of conversation restore. Add a fork regression test.
2026-07-25 17:55:48 +00:00
roboomp f6a93a34b4 fix(advisor): forwarded session metadata to overflow compaction
The advisor overflow-compaction one-shot calls compact() directly, bypassing the advisor Agent and its metadata resolver, so it omitted metadata.user_id. Resolve the advisor's session metadata per candidate provider (after credential selection) and pass it into the compact() options alongside sessionId/promptCacheKey. Add a regression test asserting the direct compaction request carries the advisor session id.
2026-07-25 17:49:21 +00:00
roboomp 8a998ca429 fix(advisor): refreshed provider identity on session switch
Rebind the live advisor Agent's provider session id, prompt cache key, credential resolver, metadata resolver, and telemetry identity whenever AgentSession crosses a conversation boundary. Add a regression test proving /new preserves the advisor Agent while assigning and emitting a distinct new provider session UUID.
2026-07-25 17:39:12 +00:00
roboomp c5af78d403 fix(advisor): propagate advisor provider session id via metadata
The advisor Agent is constructed separately in session-advisors.ts with
its own advisorProviderSessionId, but unlike AgentSession it never had a
metadata resolver installed. Its outbound requests therefore omitted the
metadata.user_id session identity that main and subagent requests carry,
so custom Anthropic-compatible proxies saw advisor traffic with no stable
session id to route or attribute on.

Extract buildSessionMetadata into session/session-metadata.ts and install
it as the advisor agent's metadata resolver, scoped to the advisor's own
provider session id and resolved live so token refreshes surface the
current account_uuid.

Fixes #6625
2026-07-25 17:30:18 +00:00
Jeff Scott Ward 050388b486 fix(advisor): bound Codex SSE retries 2026-07-24 15:24:14 -04:00
can1357 7eeaba0471 refactor(coding-agent/session): restructured monolithic agent session
- Extracted internal handlers and logic from AgentSession into dedicated runner, guard, and coordinator modules.
- Created standalone modules for bash execution, evaluation runners, IRC bridging, and prewalk coordination.
- Established dedicated session components for tracking stats, todos, streams, and retry fallback chains.
- Preserved existing session behavior while significantly reducing monolithic class size and complexity.
2026-07-24 01:24:44 +02:00