Commit Graph
443 Commits
Author SHA1 Message Date
can1357 29a94ac49c Merge PR #6200: fix(session): retry past synthetic tool results after mid-tool-call stall (@roboomp) 2026-07-22 21:13:15 +02:00
roboomp 31e0c8a9ee fix(session): retry past synthetic tool results after mid-tool-call stall
A stream that stalls or aborts mid-tool-call ends the assistant turn with
stopReason error/aborted, then appends a synthetic tool_result per un-run
tool call to keep the provider's tool_use/tool_result pairing intact. That
placeholder trailed the failed turn, so AgentSession.retry() — which only
inspected the last message and required role assistant — short-circuited to
false and /retry printed 'Nothing to retry'.

retry() now walks back over trailing synthetic tool results (details
__synthetic true) before the assistant + stopReason check, stripping both
the placeholders and the failed turn. Only synthetic results are skipped, so
a turn whose tools actually ran stays non-retryable. Adds an exported
isSyntheticToolResultMessage guard in agent-loop.ts.

Fixes #6056
2026-07-21 20:30:44 +00:00
usr_bin_roygbiv 7c6d691c10 fix(agent): recover tools after stream parse errors 2026-07-20 17:22:25 -05:00
can1357 65534c5c07 Merge PR #5947: perf(session): memoize convertToLlm and estimateTokens over settled history (@roboomp) 2026-07-18 19:42:52 +02:00
roboomp a713b941dc fix(agent): preserved side-effecting hub outcomes
Resolved tool interruptibility from each call's raw arguments so mixed-operation tools can keep side-effecting calls non-interruptible.

Restricted the unified hub to interrupt passive waits and followed logs while preserving start, send, and lifecycle operation results.

Fixes #5995
2026-07-18 14:41:19 +00:00
roboomp a28eb0f470 perf(session): memoized convertToLlm and estimateTokens over settled history
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.

- Added a per-message estimate cache in agent-core keyed by identity, with a
  settle gate (assistants cache only with real usage + terminal non-error
  stopReason; streaming partials bypass) and dual option-split WeakMaps for the
  default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
  with an exact-repeat outer-array reuse and slice-on-growth for append-only
  turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
  the prewalk plan-nudge scrub, via invalidateMessageCache /
  registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
  convert and repeat estimate are all >10x faster with noise under 20%.

Fixes #5934
2026-07-18 02:04:09 +00:00
can1357 bcc7d34a9a merge PR #5651 via eval/pr-5651: fix(cursor): gated mounted device execution through approval 2026-07-17 04:37:09 +02:00
can1357 c0d0ad7629 merged PR #5538: fix(agent): surface provider stream failures 2026-07-16 18:38:53 +02:00
roboomp d8aaffa814 fix(cursor): resolved tools against the live model per call
The context refresh captured the run-start model, so mid-run switches into or out of cursor-agent (retry fallback, prewalk, plan-yolo) sent the wrong tool set. Resolve supplemental tools against this.#state.model each call.

Fixes #5650
2026-07-16 03:21:23 +00:00
roboomp 8386ab2c0b fix(cursor): exposed mounted xd devices to cursor-agent
Forwarded the session xd registry into Cursor provider tool contexts.

Routed Cursor MCP execution through the mounted registry fallback and added regression coverage for built-in devices and external MCP tools.

Fixes #5650
2026-07-16 03:15:32 +00:00
can1357 e28197c694 Revert "merged PR #5602: fix(agent-loop): reclassify empty toolUse stop as retryable error"
This reverts commit 2baaea8782, reversing
changes made to 6783c3a473.
2026-07-16 04:09:44 +02:00
can1357 1df790a5d2 Revert "fix(agent-loop): discard incomplete sibling tool calls"
This reverts commit 7a2e34988b.
2026-07-16 04:08:59 +02:00
can1357 93cc1fed1c merged PR #5417: fix(openai): render native response images
# Conflicts:
#	packages/coding-agent/src/modes/components/chat-transcript-builder.ts
#	packages/coding-agent/test/agent-session-skill-keywords.test.ts
2026-07-16 03:48:21 +02:00
can1357 7a2e34988b fix(agent-loop): discard incomplete sibling tool calls 2026-07-16 03:31:57 +02:00
DarkPhilosophy a0b1541cb8 Merge remote-tracking branch 'can1357/main' into fix/visible-provider-errors 2026-07-16 02:13:42 +03:00
roboomp 4200dec047 fix(agent-loop): strip incomplete tool calls on empty toolUse stop
A dropped stream that emitted toolcall_start/delta but never toolcall_end
leaves an incomplete toolCall block in content. The prior guard bailed on
any toolCall block, so the outer loop dispatched empty or partially
parsed arguments instead of retrying the transport failure.

Track streamed tool-call ids (toolcall_start/delta) alongside completed
ones (toolcall_end): a call streamed but never completed is incomplete.
reclassifyEmptyToolUseStop now reclassifies unless a usable (atomic or
completed) tool call remains, stripping incomplete blocks first. Atomic
deliveries (single done/end(result) message, e.g. Cursor) emit no
granular events and stay usable.

Fixes #5600
2026-07-15 18:06:13 +00:00
roboomp 94fc54859f fix(agent-loop): reclassify empty toolUse in result-only completion
Streams finalized via end(result) with no terminal done/error event fall
through to the trailing-result branch, which returned response.result()
unchanged. An empty toolUse turn completing that way stayed a silent
success and never retried. Apply the same
retainCompletedToolCalls/recoverTransientErrorToolTurn/
reclassifyEmptyToolUseStop chain to the trailing branch.

Fixes #5600
2026-07-15 17:52:25 +00:00
roboomp f64a17c529 fix(agent-loop): reclassified empty toolUse stop as retryable error
A provider stream that closes after the thinking block but before the
tool call JSON is emitted finalizes with stopReason=toolUse and zero
toolCall content blocks. The loop treated this as a successful turn:
dispatched no tools, rendered an empty tool widget, and never retried.

reclassifyEmptyToolUseStop stamps such a turn as stopReason=error with
the Transient classifier bit set explicitly, so AgentSession's standard
retry-with-backoff path fires regardless of message-text matching.

Fixes #5600
2026-07-15 17:44:53 +00:00
can1357 5ff277349c refactor(coding-agent): consolidated tool surface onto xd:// devices and hub
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
2026-07-15 15:16:29 +02:00
DarkPhilosophy d00e5548e2 fix(agent): drop incomplete failed tool calls 2026-07-15 03:11:25 +03:00
DarkPhilosophy 5a7f107802 fix(agent): preserve Cursor results on stream failure 2026-07-15 03:01:09 +03:00
DarkPhilosophy 4086418227 fix(agent): pair tools on failed partial streams 2026-07-15 02:47:15 +03:00
DarkPhilosophy b3145170ab fix(agent): surface provider stream failures 2026-07-15 02:32:37 +03:00
can1357 b58fa6e1d3 Merge PR #5513: fix(agent): keep completed tool results from false "skipped" placeholder (@roboomp) 2026-07-14 23:11:08 +02:00
can1357 6b381eb765 Merge PR #5451: fix(eval): honor timeout zero and classify session deadlines (@roboomp) 2026-07-14 22:58:49 +02:00
roboomp b591f3d054 fix(agent): kept completed tool results from false skipped placeholder
The interrupt-clobber branch in executeToolCalls replaced a tool's real
result with the "Skipped due to queued user message" placeholder whenever
an interrupt fired, the tool's own signal aborted, and the result was an
error. It ignored whether tool.execute() actually completed, so steering
while a tool was in flight could discard a genuine error result (e.g. a
command exiting non-zero) that the tool had already produced.

Gate the clobber on !completedToolExecution so a tool that ran to
completion keeps its real result; only tools cut off before returning are
reported as skipped. Align the aborted telemetry status the same way.

Fixes #4752
2026-07-14 19:57:38 +00:00
roboomp 8e6d26b1e8 fix(eval): honored unlimited cell timeouts
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.

Fixes #5250
2026-07-14 17:33:56 +00:00
can1357 3239b4e7d3 Merge PR #5190: fix(agent): escape Harmony compaction markers (@roboomp) 2026-07-14 18:34:14 +02:00
roboomp b6b947bdbe fix(openai): rendered native response images
- Normalized completed image_generation_call results into assistant image blocks.
- Persisted image bytes through the session blob store and rendered them in live, replay, ACP, proxy, telemetry, and HTML paths.
- Added response normalization, persistence, and TUI rendering regressions.

Fixes #4768
2026-07-14 16:15:53 +00:00
can1357 4903a13511 feat(session): implemented automated recovery and status reporting
- Added `warning` field to `CompactionEntry` and `CompactionSummaryMessage` to persist dead-end status in session history.
- Introduced a multi-tier rescue mechanism in `AgentSession` that automatically performs `elide` and `dropImages` passes when maintenance fails to recover sufficient headroom.
- Integrated visual indicators for dead-ends into the `CompactionSummaryMessageComponent`, surfacing warnings directly on the compaction divider and detail block.
- Updated `AgentSession` logic to re-evaluate progress after each rescue tier and emit recovery notices, ensuring transparent reporting of automated history rewrites.
2026-07-13 01:40:14 +02:00
can1357 9a868d2e7a feat: introduced agent suspension mechanism with pause command and ui
- Introduced an `AgentPauseGate` mechanism to suspend and resume agent loops and tool executions safely.
- Added a `/pause` slash command to trigger a new fullscreen UI that manages agent suspension and lifecycle.
- Integrated pause checks into the core agent loop and tool execution pipeline to ensure responsive state handling.
- Provided a new pause screen component to facilitate user interaction and resume control during suspension.
2026-07-11 17:15:00 +02:00
roboomp 9831386deb fix(agent): escaped harmony compaction markers
Escaped Harmony control-token markers when compaction serializes transcripts into plain summary prompts so Copilot gpt-5.6 models do not reject analysis-channel text.

Added regression coverage for summary-bound Harmony serialization while keeping native transcript rendering unchanged.

Fixes #5184
2026-07-11 13:36:32 +00:00
can1357 0f8fed88e1 Merge remote-tracking branch 'origin/farm/6d5a35ab/advisor-steering-skip-message' 2026-07-11 07:39:30 +02:00
can1357 531880c620 feat: improved json serialization for bigint values
- Introduced `stringifyJson` helper to preserve bigint precision by serializing them as decimal strings.
- Replaced native `JSON.stringify` across compaction and session management modules to prevent serialization errors when handling bigint values in tool arguments.
- Added regression tests in `agent` and `coding-agent` packages to ensure bigint tool arguments remain intact through compaction and persistence flows.
2026-07-11 00:12:29 +02:00
can1357 c30bdc54ce feat(ai): enforced all_turns reasoning context for responses lite
- Force `reasoning.context` to `all_turns` in OpenAI Codex requests when `responsesLite` is enabled.
- Include `reasoning.encrypted_content` in the `include` header for lite responses.
- Update request transformation logic to ensure these fields are populated even when reasoning effort is not explicitly set.
2026-07-10 20:09:31 +02:00
roboomp 74c63fa6c5 fix(agent): labeled system steering skips accurately
- Carried steering queue origin through mid-batch interrupt polling.
- Preserved queued-user skip wording while labeling advisor/system steering as system advisory skips.
- Added regression coverage for advisor steering skip wording.

Fixes #5074
2026-07-10 14:00:59 +00:00
can1357 d435385ab1 feat: introduced max reasoning effort tier across model and rpc systems
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
2026-07-10 13:39:42 +02:00
can1357 17c7c6d0f8 fix(agent): preserve external abort boundaries 2026-07-10 12:37:46 +02:00
can1357 a3c0d536cc Merge PR #4971: fix(agent): stop terminal yield loops after IRC wake 2026-07-10 12:37:38 +02:00
can1357 2dafa7ac79 feat: further codex metadata 2026-07-10 11:35:09 +02:00
can1357 29deeef876 feat: enabled codex responses lite for gpt-5.6 models and remote compaction
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
2026-07-10 09:46:27 +02:00
roboomp 2c838e9622 Merge remote-tracking branch 'origin/main' into farm/8134d5cd/fix-revived-yield-loop 2026-07-10 03:38:00 +00:00
can1357 fde4a19c62 feat: added prompt-cache affinity support for grok models
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
2026-07-09 22:28:58 +02:00
roboomp 0f3cf73a85 fix(agent): stopped terminal yield wake loops
- Aborted the active agent loop synchronously when a terminal yield tool result finishes, so IRC-wake turns stop before another provider call.
- Added a regression covering idle IRC wake handling after a terminal yield.

Fixes #4963
2026-07-09 19:13:29 +00:00
can1357 53df3c82b7 style: applied biome formatting to merged sources 2026-07-08 15:37:43 +02:00
can1357 1bb29873ea fix(agent): adapted scoped TTSR abort labels to completed-call retention
- derive per-tool abort labels from a tool-scoped abort signal for provider-built aborted messages
- restore main's single-call TTSR label test dropped by the merge
- complete the innocent read in the sibling-label test; incomplete matched calls mint no placeholder under the retention policy
2026-07-08 15:34:37 +02:00
can1357 32158b74ea merge PR #4542: fix(coding-agent): scoped TTSR abort reason to matching tool call 2026-07-08 15:27:25 +02:00
can1357 9547ff6f55 merge PR #4633: fix(agent): support chat completions remote compaction endpoints 2026-07-08 15:19:35 +02:00
roboomp 389515baad fix(agent): retried handoff auto-only tool choice errors
Retried handoff generation with toolChoice auto when a provider rejects the cache-preserving toolChoice none request as auto-only.
Kept unrelated provider 400s terminal so bad request failures still surface without masking the cause.

Fixes #4715
2026-07-06 14:13:02 +00:00
roboomp da6f0ebc38 fix(agent): used wire model id for chat compaction
- Sent remoteCompaction.model or requestModelId in chat-completions remote compaction requests instead of the local catalog id.
- Covered both direct requestRemoteCompaction formatting and end-to-end openai-completions compaction with wire model ids.

Fixes #4630
2026-07-05 21:12:08 +00:00