Commit Graph

297 Commits

Author SHA1 Message Date
usr-bin-roygbiv 17e678fa2d fix(agent): honor explicit compaction endpoint
(cherry picked from commit 3bc5c578ddd2e626eda2049a660de7bffb674b52)
2026-07-30 01:42:53 +02:00
usr-bin-roygbiv 0c31095d34 fix(agent): skip V2 retries for compaction auth errors
(cherry picked from commit 07a54771834435b7897208ab9d6d42401f0203e1)
2026-07-30 01:42:06 +02:00
usr-bin-roygbiv b927da95e6 fix(agent): preserve native compaction error precedence
(cherry picked from commit 3b1c1a420132485ac4f5a4803615a70d1e05c794)
2026-07-30 01:42:06 +02:00
usr-bin-roygbiv 4c2b40f8b3 fix: preserve streaming compaction auth errors
(cherry picked from commit 6fe25c6a08e169a15fca057a56e1ce36f46bc641)
2026-07-30 01:42:06 +02:00
Roy 5ce80fdedc fix: preserve provider-native compaction semantics
(cherry picked from commit 426ac1e147c08092c7d0b6c4a7af5f2e09eaf941)
2026-07-30 01:42:06 +02:00
can1357 daeb683528 feat(agent): restructured tool call dispatch to validate arguments earlier
- Added `prepareToolCallDispatch` and `PreparedToolCall` to handle argument validation and `beforeToolCall` before message snapshotting.
- Implemented `preparedDispatchByMessage` WeakMap to store pre-dispatch results for streamed messages.
- Updated `executeToolCalls` to consume pre-computed dispatch preparation results.
- Updated documentation and changelog to specify `beforeToolCall` timing on the streamed path.
2026-07-27 23:01:50 +02:00
can1357 da6d11de0e feat(agent): introduced prepareToolCall phase supporting argument replacement
- Added a prepareToolCall phase to the agent loop running before tool scheduling for validation and hooks.
- Updated BeforeToolCallContext and result types to support argument replacement instead of in-place mutation.
- Updated coding-agent extension handling and runner to track emitted tool calls and re-evaluate approvals on input revisions.
- Added comprehensive test coverage for argument replacement, concurrency resolution, and schema validation.
2026-07-27 22:55:20 +02:00
can1357 ae01a76136 fix(agent): hardened pre-model-call gate state cleanup and API surface
- Cleared the retained soft-requirement lifecycle alongside the deferred
  hard choice: clearDeferredToolDirectives() owns both, is called from
  clearAllQueues/reset and session-scoped tool-state cleanup, with a
  regression covering reminder re-injection after a queue clear.
- Allowed void-returning pre-model gates via the named AgentBeforeModelCall
  type and normalized gate results in the loop and Agent dispatcher.
- Documented that the first gate installed mid-run applies from the next
  run; corrected the onToolChoiceRejected contract docs; documented the
  cross-run lifetime of ToolChoiceQueue's in-flight claim.
- Removed the unused addBeforeModelContextBuild hook.
- Relocated both packages' changelog entries out of the released 17.1.4
  sections into Unreleased with PR attribution, folding the never-shipped
  Fixed bullet into Added and noting the input-event timing change.
2026-07-27 14:08:53 +02:00
can1357 6bbfc1110a Merge PR #6543: feat(agent): add a pre-model-call gate that can stop the turn (@paralin) 2026-07-27 14:01:11 +02:00
can1357 9d76e168ba Merge PR #6706: fix(ai): preserve custom Anthropic web-search history (@roboomp) 2026-07-27 04:58:24 +02:00
Christian Stewart 01ecab6df7 fix(agent): clear deferred choices on branches
A pre-model gate can defer a claimed hard tool choice for the next call. Branch transitions cleared the coding-agent queue but left that agent-owned value alive, allowing an obsolete forced tool to cross into the replacement transcript.

Expose the narrow deferred-choice reset at the Agent owner and invoke it from the shared session-scoped tool-state cleanup used by both branch paths. Failed session switches retain their existing rollback behavior.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:05:30 -07:00
Christian Stewart c5e5f6dbc3 fix(agent): close aborted harmony retry turns
A Harmony retry keeps its logical turn open while the next provider call is prepared and gated. An abort during that gate previously ended the agent stream directly, leaving observers with an unmatched turn_start event.

Route the aborted gate through the existing pre-model stop owner so it emits the synthetic aborted message and closes the open turn before ending the stream. Fresh turns retain their existing no-provider-call cancellation behavior.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:02:18 -07:00
Christian Stewart e79eabc3b6 feat(agent): add a pre-model-call gate that can stop the turn
The agent loop had no place to refuse a provider request. A host that needs to
act on the assembled context before it is billed, checking that the prompt still
fits the window, that a budget boundary has not been crossed, or that the
session should hand off instead of spending, could only observe the request
after the fact, when the tokens were already committed.

Add `AgentLoopConfig.beforeModelCall`, asked once per turn beside the deadline
check and before `turn_start` is emitted. A `stop` result ends the stream with
no turn open, so nothing has to synthesise a cancellation event and no consumer
is left holding a half-open turn. Placing it there also keeps `turn_end`'s
contract intact: that event carries the assistant message for a completed turn,
and a gated stop has no assistant message to report.

`syncContextBeforeModelCall` keeps its existing void contract and its job of
refreshing prompt and tool state, so implementations typed as returning void are
unaffected.

`Agent.setBeforeModelCall` installs the host's callback, and `addBeforeModelCall`
registers an additional callback without displacing the host's, returning a
disposer so an extension can attach and detach independently. A supplied
`reason` is logged where the loop stops.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:02:18 -07:00
can1357 0d3fd6197c Merge PR #6544: feat(agent): wake steering on an event instead of polling for it (@paralin) 2026-07-26 15:55:40 +02:00
roboomp f26de4fad3 fix(agent): preserved legacy proxy terminal streams
Made finalized terminal content optional and retained the delta-reconstructed assistant content when older proxy servers omit it.

Covered both legacy done and error events so terminal cleanup remains safe and the stream resolves.

Fixes #6703
2026-07-26 13:50:34 +00:00
roboomp e26c3034f0 fix(agent): carried terminal content through proxy streams
Made proxy done and error events carry finalized assistant content so provider-only blocks survive event reconstruction.

Covered Anthropic native web-search history restored after live events omitted the opaque blocks.

Fixes #6703
2026-07-26 13:44:14 +00:00
can1357 26f89bd897 Merge PR #6556: fix(omp): fit native compaction after large tool output (@oleksoleksoleks) 2026-07-26 15:41:27 +02:00
Diogo Soares Rodrigues 9088fe821b fix(cursor): await error-drain transforms and sanitize mirrored todo labels
Two boundary defects on the mirrored-todo path.

1. The Agent error drain snapshotted #cursorToolResultBuffer without
   awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
   async cursorOnToolResult still running when the provider errored
   patched an entry the catch path had already detached, so the
   pre-transform payload was persisted. A provider error is exactly when
   a transform is most likely to be in flight.

2. The todo renderer interpolated mirrored provider text straight into
   terminal output. A Cursor snapshot carries model-authored task
   content, phase names and summary text verbatim, so a label holding
   ANSI/C0 sequences rewrote the terminal on every render and replay.
   sanitizeText alone is not enough - it preserves tabs, which punch
   holes in bordered output - so every display path now funnels through
   one forDisplay() helper: task labels, blocker notes, phase headers,
   the zero-task fallback, and the streaming renderCall preview. Raw
   values are untouched; content and phase name are the identity keys
   the local list is looked up by and what gets persisted.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues eaf7a3d605 fix(cursor): pair a result for every server-resolved call
Three orphan paths, same failure mode: the assistant block is marked
kCursorExecResolved before the work runs, so agent-loop.ts emits no
placeholder for it, and any path that produces no toolResult leaves the
call unpaired — buildSessionContext then strips the whole interaction
from every rebuilt transcript.

1. resolveExecHandler returned no toolResult on three exits (no handler
   installed, handler produced nothing, handler threw). Each now pairs a
   result carrying the same text the server sees in execResult, routed
   through onToolResult like a real one. `pairing` is a required
   parameter so a new callsite cannot silently recreate the orphan.

2. Agent only installed its result-buffer sink when cursorExecHandlers
   or cursorOnToolResult was set. Both are optional, so a bare SDK host
   dropped the provider result on the floor. Installed unconditionally;
   a non-Cursor provider never calls it.

3. A todo completion frame with no tool_call (the field is optional)
   skipped settlement entirely. It now settles as "nothing to mirror".

Also fixes an empty update_todos with a nonzero total_count being
mirrored as an authoritative clear: the length guard added earlier
skipped the mismatch check for empty responses, so a partial or
size-limited merge response deleted every local task at once.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues a9f1e3a257 fix(agent): await pending cursor result transformers before draining
An async cursorOnToolResult that resolved after the message_end drain
had its rewrite silently discarded: the reservation kept the call from
dangling, but the late patch mutated a buffer entry the drain had
already detached, so the persisted message kept the pre-transform
payload.

Each entry now records the in-flight transformer promise, and
#emitCursorSplitAssistantMessage awaits any that are still pending
before appending and emitting. This matches the exec-channel paths,
which already await onToolResult (cursor.ts:1461).

A rejecting transformer is swallowed per-entry, so a failing hook can
neither take the turn down nor cost the reserved result.

The previous test asserted the old limitation (late rewrite NOT
persisted) and its premise is now unreachable, so it is replaced by the
rejection contract. Stale limitation notes in agent.ts and types.ts are
updated.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues cc210094ef fix(cursor): refuse todo dependency graphs and sanitize failure text
Two review findings on the native todo sync.

- TodoItem.dependencies is a graph the local model cannot store: rows
  are keyed by content, carry no id, and hold no edges. An imported
  dependent row files as plain pending and nextActionableTask then
  offers work the server considers blocked. Refuse snapshots with an
  edge pointing at an unfinished row; edges whose blockers already
  finished constrain nothing and still mirror.

- The todo failure warning interpolated the provider error verbatim.
  Collapse and truncate it at the render boundary.

Also documents two known, unfixed defects: an async cursorOnToolResult
transformer resolving after the buffer drain, and the todo card
lifecycle race. Emitting a synthetic tool_execution_start for the
latter was measured and rejected -- the completion deletes the entry it
creates, so the late streamed block adds a second card.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues ea023c380a fix(cursor): persist the phase-bearing todo result and close a buffer race
The previous commit paired every server-resolved todo block with a
result, but built that result in the provider from the flat snapshot.
`todoToolRenderer.renderResult` reconstructs the list exclusively from
`details.phases`, so the block survived the dangling-strip only to replay
as `Todo 0 tasks`.

Only the host computes that grouping -- the provider sees a flat list --
so `todoSync` now returns the result it already assembled and the
provider persists it verbatim. A refused snapshot never reaches the host,
so the provider's summary-only fallback still covers that path, and
exactly one result is emitted either way.

Separately, `Agent`'s Cursor buffering wrapper pushed its entry only
after awaiting the optional `cursorOnToolResult` transformer. The
provider dispatches decoded messages with `void handleServerMessage(...)`,
so a `message_end` from the same chunk could drain the buffer while a
transformer was still pending, dropping the result. The entry is now
reserved synchronously and patched in place when the transformer
resolves, keeping buffer order and still applying the customization.
Production is unaffected -- `sdk.ts` sets no transformer -- but the
option is supported and its contract returns a Promise.

Tests: a delayed-transformer case that loses the result without the
buffering change, and a replay case driving the persisted result through
`buildSessionContext` and asserting `details.phases` rebuilds a non-empty
list -- the id-pair assertion alone did not catch the empty render.
2026-07-26 09:29:13 -03:00
Christian Stewart d363ca0968 feat(agent): wake steering on an event instead of polling for it
While a turn is running, the loop watched for out-of-band steering by
re-checking the queue on a fixed interval. The interval sets a floor on how long
a message waits and burns a wakeup on every tick that finds nothing, and the
delay is worst exactly when it matters: a cancellation or a correction issued
during a long tool loop sits until the next tick.

Wait on the queue instead. `Agent` keeps a set of steering waiters and notifies
them whenever a message is enqueued, and the loop awaits
`waitForSteeringMessages`, re-checking only when woken. The timer path remains
for the IRC interrupt queue, which is session-owned and has no wake callback, so
nothing regresses where no event source exists.

Every wait races against local abort. The callback contract does not require an
implementation to observe the signal, and one that resolves only on the next
queue event would otherwise never settle once a batch finished, so awaiting it
during teardown would hang a batch that simply had no steer.

Which controllers a steer raises is unchanged: this is a timing change, not a
change to what an interrupt does.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-25 03:54:46 -07:00
Alexander Kirilin 46067950d6 fix(omp): reserve maximum image compaction budget 2026-07-24 21:29:43 -04:00
Alexander Kirilin 1a0a7c3e74 fix(omp): harden native compaction fit checks 2026-07-24 21:20:47 -04:00
Alexander Kirilin 51195b2300 fix(omp): preserve compaction input when trimming fails 2026-07-24 21:07:09 -04:00
Alexander Kirilin cc336acec4 fix(omp): fit native compaction after large tool output 2026-07-24 20:16:48 -04:00
can1357 b0f1b5c0c5 Merge PR #6496: fix(ai): preserve Anthropic server-tool history (@roboomp) 2026-07-24 16:26:52 +02:00
can1357 9f8aa87dbf feat(task): removed per-call model override from task tool
- Removes `model` field from task item/schema, TaskParams, and TaskItem types.
- Removes model selector validation, formatting, and approval display logic.
- Updates task tool priority docs to reflect that model is no longer per-call overridable.
- Updates eval agent() helper docs and prompt templates to remove model parameter.
- Updates tests to reflect removal of model override capability.
2026-07-24 14:51:48 +02:00
roboomp b461b8e766 fix(agent): counted Anthropic server-tool blocks in token estimate
estimateTokens now charges for serialized anthropicServerTool blocks so context maintenance sees the server-tool payload replayed on the wire; excluded from the compaction floor like other encrypted reasoning.
2026-07-24 10:43:58 +00:00
can1357 c780662881 feat(agent): added resolveFallbackTool option for routing unadvertised tool calls
- Add `resolveFallbackTool` callback to `AgentOptions` and `AgentLoopConfig` that resolves tool calls not found in the advertised set.
- Use the callback as a third lookup step after `name` and `customWireName` match, enabling side transports like `xd://` device mounts.
- Add test coverage verifying the fallback resolves known devices and preserves "not found" errors for unknown names.
- Wire the coding agent's device registry as `resolveFallbackTool` in both `createAgentSession` and `streamAgentSession` paths.
2026-07-24 12:18:17 +02:00
usr-bin-roygbiv b9504f65e7 feat: add native Codex computer use 2026-07-24 01:40:04 +00:00
can1357 2d5f8d52d0 Merge PR #6411: feat(agent): classify Cloudflare response cache status (@riverpilot) 2026-07-24 02:24:30 +02:00
can1357 46d2b7d864 Merge PR #6439: fix(compaction): stop feeding reasoning back to Claude summarizer (@roboomp) 2026-07-23 23:05:47 +02:00
roboomp 5a70eda7d7 fix(compaction): stop feeding reasoning back to Claude summarizer
Both compaction serializers reproduced prior assistant reasoning as text
bound for a Claude target, tripping Anthropic's reasoning_extraction
refusal and wedging Fable 5 sessions:

- context-full: serializeConversation rendered thinking verbatim inside
  <thinking> tags via the anthropic dialect renderer. Now drops thinking
  blocks when the summary target dialect is anthropic; other dialects
  (e.g. Harmony) keep native reasoning.
- snapcompact: emitted ¶think sections baked into replayed archive
  frames. Added an includeThinking serialize option (default true) and
  wired the agent session to disable it for Anthropic-dialect models.

Fixes #6093
2026-07-23 20:38:27 +00:00
can1357 29c2bd2595 Merge PR #6415: fix(compaction): judge remote-preserve reuse against the active model (@roboomp)
# Conflicts:
#	packages/coding-agent/src/session/agent-session.ts
2026-07-23 22:16:11 +02:00
Alexander Kirilin b225b27e92 feat(agent): classify Cloudflare response cache status 2026-07-23 15:18:37 -04:00
roboomp 116b8f4597 fix(compaction): judged remote-preserve reuse against the active model
After an OpenAI remote compaction, prepareCompaction decided whether to
keep the provider-native replay boundary or re-expand its originals by
asking whether *any* compaction candidate (every role model plus the
largest-context available model) shared the payload's provider. In a
multi-role setup where a role such as modelRoles.smol stays on OpenAI,
the check passed forever, so a session switched to a non-OpenAI active
model kept a placeholder-only summary and never recovered the compacted
span for the rest of the session.

Judge reusability against the active model — the one that assembles the
request context every turn — instead of the candidate set. When the
active model cannot replay the payload, re-expand the originals into a
portable local summary, matching the self-healing already present for
single-provider migrations.

Fixes #6343
2026-07-23 19:12:50 +00:00
roboomp 0640b7e3c7 fix(agent): deliver queued steer on empty transcript instead of looping
A steer queued on a session with an empty transcript was undeliverable:
Agent.continue() threw "No messages to continue from" before dequeuing any
steering message, so the RPC idle-drain re-armed continue() on every microtask
(gated only on hasQueuedMessages(), which never cleared) — an unbounded
allocation loop that OOM-killed the process.

continue() now consumes a queued steer/follow-up as the opening turn when the
transcript is empty, mirroring the assistant-tail branch. The queue drains, the
re-arm's hasQueuedMessages() gate goes false, and the loop terminates.

Fixes #6344
2026-07-23 19:02:47 +00:00
can1357 c522eceff2 Merge PR #6119: feat: lift subagent async/auto-background limits via owner-routed delivery and quiescence (@korri123)
# Conflicts:
#	packages/coding-agent/src/task/executor.ts
2026-07-23 17:52:52 +02:00
can1357 f741936c56 test(agent): signal tool boundary from the gate itself
The Promise.withResolvers signal resolved inside the scripted model
response fires before the loop can possibly dispatch the tool, so the
parked-state assertions passed even with parking disabled (verified by
simulation: 20/20 green with waitUntilResumed stubbed out). Resolve the
readiness signal from a test-local wrap of agentPauseGate.waitUntilResumed
instead: deterministic (no wall-clock race with the cold yieldIfDue
timer), immune to sibling restoreAllMocks (manual patch, restored in
finally), and a non-parking regression now hangs the await and fails
the test.
2026-07-22 21:12:22 +02:00
can1357 206d812244 Merge PR #6185: test(agent): synchronize pause gate tool boundary (@any-victor) 2026-07-22 21:12:21 +02:00
Victor Araújo 3179a524be test(agent): synchronize pause gate tool boundary 2026-07-22 00:31:36 -03:00
Victor Araújo eee940c1c3 test(agent): avoid global pause gate spy 2026-07-21 23:29:13 -03:00
Victor Araújo b927c4d0c5 test(agent): synchronize pause gate tool boundary 2026-07-21 23:22:01 -03:00
usr_bin_roygbiv 7c6d691c10 fix(agent): recover tools after stream parse errors 2026-07-20 17:22:25 -05:00
Kormákur 21a7725a09 feat: lift subagent async/auto-background limits via owner-routed delivery and quiescence
Three-piece architecture so subagents inherit async.enabled and
bash.autoBackground.enabled instead of having both force-disabled:

- Owner-routed delivery: AsyncJobManager gains registerDeliverySink /
  waitForOwnerJobs; every AgentSession registers a sink for its own agent
  id, so background job results inject into the owning agent's run.
  Owned deliveries with no live sink dead-letter (result retained on the
  job row) instead of misrouting into the first top-level session.

- Quiescence barrier: a subagent's final yield with owner jobs still
  running/undelivered is a scheduling pause, not completion. The run
  driver notifies the model once (hub wait/cancel), settles owner work,
  and folds results in as async-result follow-ups; teardown cancels and
  awaits surviving jobs before isolation worktree capture/cleanup.

- Steering soft channel: queued steering no longer hard-aborts
  non-interruptible tools; it aborts interruptible waits and raises a
  cooperative ToolCallContext.steeringSignal. The mid-batch watch runs
  for every batch, and auto-backgroundable bash backgrounds itself on
  steer so incoming messages inject promptly with no work lost.
2026-07-20 15:00:37 +00:00
can1357 65534c5c07 Merge PR #5947: perf(session): memoize convertToLlm and estimateTokens over settled history (@roboomp) 2026-07-18 19:42:52 +02:00
roboomp a713b941dc fix(agent): preserved side-effecting hub outcomes
Resolved tool interruptibility from each call's raw arguments so mixed-operation tools can keep side-effecting calls non-interruptible.

Restricted the unified hub to interrupt passive waits and followed logs while preserving start, send, and lifecycle operation results.

Fixes #5995
2026-07-18 14:41:19 +00:00
roboomp a28eb0f470 perf(session): memoized convertToLlm and estimateTokens over settled history
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.

- Added a per-message estimate cache in agent-core keyed by identity, with a
  settle gate (assistants cache only with real usage + terminal non-error
  stopReason; streaming partials bypass) and dual option-split WeakMaps for the
  default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
  with an exact-repeat outer-array reuse and slice-on-growth for append-only
  turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
  the prewalk plan-nudge scrub, via invalidateMessageCache /
  registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
  convert and repeat estimate are all >10x faster with noise under 20%.

Fixes #5934
2026-07-18 02:04:09 +00:00