Commit Graph
219 Commits
Author SHA1 Message Date
can1357 2dafa7ac79 feat: further codex metadata 2026-07-10 11:35:09 +02:00
can1357 29deeef876 feat: enabled codex responses lite for gpt-5.6 models and remote compaction
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
2026-07-10 09:46:27 +02:00
can1357 9547ff6f55 merge PR #4633: fix(agent): support chat completions remote compaction endpoints 2026-07-08 15:19:35 +02:00
roboomp 389515baad fix(agent): retried handoff auto-only tool choice errors
Retried handoff generation with toolChoice auto when a provider rejects the cache-preserving toolChoice none request as auto-only.
Kept unrelated provider 400s terminal so bad request failures still surface without masking the cause.

Fixes #4715
2026-07-06 14:13:02 +00:00
roboomp da6f0ebc38 fix(agent): used wire model id for chat compaction
- Sent remoteCompaction.model or requestModelId in chat-completions remote compaction requests instead of the local catalog id.
- Covered both direct requestRemoteCompaction formatting and end-to-end openai-completions compaction with wire model ids.

Fixes #4630
2026-07-05 21:12:08 +00:00
roboomp 241beb9b3d fix(agent): supported chat completions remote compaction
- Sent OpenAI-compatible chat messages when compaction.remoteEndpoint targets /chat/completions while preserving the existing custom summarizer payload elsewhere.
- Added regressions for direct wire formatting and end-to-end openai-completions compaction against a configured chat endpoint.

Fixes #4630
2026-07-05 20:53:06 +00:00
roboomp 38454d3321 fix(agent): excluded orchestration tokens from context sizing
calculateContextTokens returned usage.totalTokens which, with the new
Usage.orchestration sidecar, folds provider-side orchestration back into the
context size used by auto-compaction/context promotion thresholds. Subtract
the orchestration sidecar so context sizing stays conversation-only while
cost and totalTokens keep the orchestration spend visible.

Refs #4469
2026-07-03 16:53:11 +00:00
can1357 82f3706410 Merge remote-tracking branch 'origin/farm/6f26b11c/distinguish-synthetic-tool-failure' 2026-07-02 23:42:38 +02:00
roboomp 107503d201 test(agent-loop): align tool result details type with parameter shape 2026-07-02 21:27:32 +00:00
roboomp eb64e7e5d1 fix(agent-loop): skip cursor exec-resolved toolCall blocks to avoid double-execution
Codex review on PR #4351: synthesizing toolCall content blocks for
Cursor's exec-channel native tools made the shared agent loop treat the
finalized assistant message as a fresh runnable tool turn. Because
executeToolCalls filters message.content for any toolCall block on
stop/toolUse, bash/write/delete/etc. ran a second time after Cursor
already executed them server-side via the bridge, duplicating side
effects and appending conflicting toolResults.

- packages/ai/src/utils/block-symbols.ts: add `kCursorExecResolved`
  symbol and `CursorExecResolvedCarrier` carrier type. Symbol-keyed so
  the marker never leaks into JSONL; rebuild pairs blocks with toolResult
  messages by id.
- packages/ai/src/providers/cursor.ts: stamp the marker onto every
  block `synthesizeCursorExecToolCall` emits and extend `ToolCallState`.
- packages/agent/src/agent-loop.ts: filter marked blocks out of the
  runnable-toolCall extraction in both the main runnable path and the
  error/aborted placeholder path, plus defense-in-depth inside
  `executeToolCalls`. Marked blocks stay in `assistantMessage.content`
  for persistence + rebuild rendering; they just never re-execute.
- packages/agent/test/agent-loop.test.ts: two regression tests — one
  proves a marked block yields zero `tool.execute` calls and no
  `tool_execution_*` events from the loop, the other verifies mixed
  batches still run the unmarked blocks unchanged.

Fixes #4348
2026-07-02 21:26:37 +00:00
roboomp 3d86935a5e fix(agent): distinguish synthetic placeholder tool results from real tool failures
When an assistant turn ends with stopReason="error" after a tool call
was already streamed, agent-loop synthesizes a placeholder tool result
via createAbortedToolResult() to preserve the tool_use / tool_result
pairing the provider API requires. The previous wording ("Tool
execution failed due to an error: <upstream>") and event shape
(normal tool_execution_start / tool_execution_end with empty details)
were indistinguishable from a real local tool failure — a Codex
websocket close mid-turn showed up in the CLI as a broken Edit panel,
misattributing provider-transport faults to the local tool.

Reword the "error" placeholder to state explicitly that the tool
never ran ("Tool call was not executed because the provider stream
ended with an error before the tool could run: <upstream>") and thread
a SyntheticToolResultDetails discriminator ({ __synthetic: true,
source: "assistant_stop_error" | "assistant_stop_aborted" |
"assistant_stop_skipped" | "assistant_stop_length", executed: false,
upstreamError }) through both the ToolResultMessage.details and the
tool_execution_end event's result.details, so downstream UI/telemetry/
ACP consumers can render "provider transport failed, tool not
executed" without string-matching content.

Fixes #4321
2026-07-02 15:19:30 +00:00
can1357 21f5728ef8 fix(agent): prevented consuming legacy steering queue during mid-batch interrupts
- Stopped calling the consuming `getSteeringMessages` getter during mid-batch interrupt polls to prevent stranding or dropping messages before they reach the injection boundary.
- Skip subsequent steering checks in the poll loop once an interrupt has already triggered.
- Added a regression test to ensure legacy steering remains queued until the injection boundary when no non-consuming peek exists.
2026-07-02 03:34:14 +02:00
can1357 fc91aadccd fix(agent): honored explicit compaction reserve equal to default
- Made CompactionSettings.reserveTokens optional so field presence carries provenance; the proportional small-window fallback only applies to genuinely defaulted reserves.
- Clamped the fallback reserve to >= 1 and the derived threshold strictly below the context window.
- Changed the coding-agent settings-schema default from 16384 to unset so Settings.get() no longer materializes a default that masks provenance.
2026-07-02 00:32:48 +02:00
can1357 021d4fc1e3 Merge PR #4161: fix(agent): interrupt waits for IRC delivery (@roboomp) 2026-07-01 21:53:19 +02:00
can1357 6ce3f686b2 fix(agent): budget branch tool results after truncation 2026-07-01 21:53:15 +02:00
can1357 2a632d41c0 Merge PR #4112: fix(agent): preserve tool results in branch summaries (@roboomp) 2026-07-01 21:53:15 +02:00
can1357 23ea5e0808 test(agent): require queued skip text content 2026-07-01 21:47:55 +02:00
can1357 8665bd5da6 Merge PR #3855: fix(agent): clarify queued skipped tool results (@wolfiesch) 2026-07-01 21:47:55 +02:00
can1357 12120a1cd4 Merge PR #3647: fix(tool): require browser run code in schema (@roboomp) 2026-07-01 21:42:23 +02:00
roboomp 1754c108df fix(agent): scoped irc-only aborts to interruptible tools
Previously an IRC-only interrupt shared the batch-wide abort controller with
user steering, so a peer message that landed while an interruptible wait ran
alongside a foreground non-interruptible tool (e.g. bash) killed the foreground
tool too. Split the batch signal into a shared steering/external channel and an
interruptible-only IRC channel; each record picks its per-tool signal based on
the tool's interruptible flag, and only that signal is used for validation,
before/after hooks, and execute. User steering still upgrades an in-flight IRC
interrupt to a full batch abort.
2026-07-01 17:14:04 +00:00
roboomp 619bfda3eb fix(agent): interrupted irc waits
Fixes #4160
2026-07-01 16:56:08 +00:00
roboomp a1a75e91b7 fix(agent): skipped useless tool results before branch-summary budget
Useless non-error toolResult entries are dropped by serializeConversation() anyway. Skip them in prepareBranchEntries() too so a large discardable payload at the branch tip cannot exhaust the token budget and starve older useful context.

Fixes review comment on #4112
2026-07-01 07:11:18 +00:00
roboomp 644a20638b fix(agent): preserved tool results in branch summaries
Included informative tool result messages in branch summary serialization so abandoned-branch observations survive tree navigation. Added regression coverage for informative and useless tool outputs.

Fixes #4076
2026-07-01 07:05:51 +00:00
Wolfgang Schoenberger 0d9ef89549 fix(agent): clarify queued skipped tool results 2026-06-29 19:46:13 -07:00
ben dc17aead5f fix(agent): narrow stream recovery review fixes 2026-06-28 16:18:46 +08:00
ben cb4e68888f fix(agent): recover completed tools after stream read errors 2026-06-28 16:18:46 +08:00
can1357 1190aade82 test(agent): strengthened consistency and persistence logic in integration tests
- Added regression test in `remote-compaction` to verify that concurrent v2 compaction preparation correctly reuses preserved history and avoids redundant re-expansion.
- Added mock-backend verification in `session-storage` to ensure that failed atomic title updates do not rollback newer optimistic state.
- Updated `sql-session-storage` expectations to account for the preserved fixed-width title slot header in session files.
- Expanded `remote-compaction` fetch header validation to include `x-client-request-id` assertion.
2026-06-28 09:52:44 +02:00
can1357 102d6d54ad feat: implemented v2 streaming remote compaction for model history state
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
2026-06-28 07:27:02 +02:00
roboomp a67506f938 fix(agent): scoped intent injection inside union variants
When a tool schema is a pure anyOf/oneOf (no own properties), push the intent field into each closed branch and skip the root sibling. The prior pass added properties: { i } / required: [i] next to the alternation; OpenAI strict sanitization then promoted that to a closed root that rejected every input. allOf members are sub-constraints, not alternatives, so they are no longer recursed.
Added normalizeTools tests for the union-shape path and a post-normalize strict-mode satisfiability test for the browser tool.

Fixes #3645
2026-06-27 11:31:11 +00:00
can1357 cfa0cd84ba feat(agent): resolved partial json leakage by enforcing streaming cleanup
- Update scrubPartialJson to utilize clearStreamingPartialJson for consistent tool-call cleanup.
- Adjust execution order in streamProxy to ensure partial error messages are finalized before scrubbing.
- Remove redundant test expectation comment regarding partialJson leakage.
2026-06-27 12:26:11 +02:00
can1357 357c29224d feat: implemented symbol-based streaming state for isolated metadata
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
2026-06-27 12:07:08 +02:00
can1357 0b8206a46d fix(compaction): preserve provider defaults for remote compaction 2026-06-27 01:39:34 +02:00
can1357 14cc9cba0f Merge PR #3106: fix(compaction): enable custom provider remote compaction (@roboomp) 2026-06-27 01:39:34 +02:00
can1357 3e5360ca4c Merge PR #3060: feat(provider): add GitLab Duo Agent provider (@jiwangyihao) 2026-06-27 01:39:31 +02:00
can1357 1ed343d9a4 Merge PR #3595: fix(agent): stop replaying provider refusals (@roboomp) 2026-06-26 23:27:40 +02:00
can1357 2899ce1bc2 Merge PR #3328: fix(proxy): scrub transient partialJson from final tool calls (@roboomp) 2026-06-26 23:27:38 +02:00
can1357 894d338ce8 docs(ai/dialect): clarified escaping and halting rules for tool dialects
- Explicitly prohibited HTML escaping in tool arguments for all dialects to ensure raw data transmission.
- Clarified that tool call bodies are delimiter-parsed rather than XML-parsed where applicable.
- Enforced strict requirements to complete tool call output before emitting stop sequences and halting.
2026-06-26 22:35:20 +02:00
roboomp d68980fe58 fix(agent): stopped replaying provider refusals
API-level refusals now stay visible as terminal errors without being sent back as assistant dialogue on the next provider request. Added core and coding-agent conversion coverage for Anthropic refusal metadata.

Fixes #3592
2026-06-26 17:56:07 +00:00
jiwangyihao af32c22ffe fix(agent): /move 后按会话实时 cwd 重新作用域 Duo 发现
机器人指出 agent.ts 的 #cwd 在构造时固定,/move 更新 SessionManager 与
进程 cwd 后不会重建 Agent,导致 GitLab Duo Agent 的 namespace/project 发现
持续读取旧仓库的 git remote。

按既有 resolver 模式(getReasoning/getServiceTier)修复:

- Agent 新增可选 cwdResolver;构造时存入 #cwdResolver。
- AgentLoopConfig 新增 getCwd 每调用解析器,config 同时携带静态 cwd 与
  getCwd。
- agent-loop 在 streamFunction 调用点计算 effectiveCwd = getCwd?.() ?? cwd,
  每次 LLM 调用读取一次,因此运行中途的 /move 也能被工作区级 provider 发现
  感知。
- sdk.ts 主 Agent 传入 cwdResolver: () => sessionManager.getCwd(),该值在
  /move 时由 SessionManager.#cwd 更新。

新增针对可观测契约的回归测试(mock streamFn 记录 options.cwd):resolver
覆盖静态 cwd、resolver 返回 undefined 时回退静态 cwd、以及运行中途变更可被
逐次调用读取(模拟 /move)。
2026-06-26 16:13:24 +08:00
roboomp 2b68785e82 fix(agent): include provider payloads in append-only digests
Responses-style providers serialize providerPayload history items instead of the
visible message blocks when replaying native history. Include providerPayload in
the append-only per-message digest so payload-only history rewrites stop the
stable-prefix walk and re-sync the changed message before any later divergent
tail.

Add a regression where an assistant message keeps identical visible content and
id but changes its openaiResponsesHistory providerPayload while a later message
also diverges; syncMessages must preserve the prefix before the assistant and
refresh the assistant payload.

Fixes #3406
2026-06-24 23:11:12 +00:00
roboomp 5b9e009a8b fix(agent): bound append-only stable prefix by log length
Direct callers can clear AppendOnlyContextManager.log without resetting the
private sync cursor. The advisor reset path does this when recycling its helper
agent, leaving lastSyncCount and messageDigests describing the old transcript
while the physical log is empty.

Clamp the stable-prefix reuse count to the current log length before truncating
and appending. A direct log clear now forces the next sync to replay from index 0
instead of starting from a stale private cursor and dropping prefix messages from
the provider context.

Add a regression that clears the public log after syncing two messages, then
resyncs a context with the same first message and a rewritten second; both
messages must be present in the rebuilt append-only log.

Fixes #3406
2026-06-24 23:07:11 +00:00
roboomp cce627ee4b fix(agent): include tool result metadata in append-only digests
Track internal tool-result metadata in append-only per-message digests so
metadata-only rewrites of toolCallId, toolName, or isError stop the stable-prefix
walk and re-sync the changed tool result before any later divergent tail.

This prevents stale tool-result pairing or error state from being preserved when
the text content stays unchanged but provider-serialized metadata changes.

Fixes #3406
2026-06-24 22:42:51 +00:00
roboomp c803a3e30d style: bun run fix 2026-06-24 22:35:03 +00:00
roboomp 65a974aa90 fix(agent): preserve append-only log prefix on in-place message rewrites
`AppendOnlyContextManager.syncMessages` hashed a single rolling digest
over the entire synced prefix, so any in-place rewrite of an already-
synced message — per-turn `pruneSupersededToolResults` / `pruneToolOutputs`
collapsing a tool result, image stripping, or a `transformContext` re-render
— triggered `log.clear()` and re-appended the full conversation from
the current (mutated) view. The provider's cached bytes still matched
the prefix, but every position past the divergence had to be re-prefilled.
On llama.cpp / Ollama / LM Studio this re-prefilled tens of thousands
of tokens every few turns (`n_past \u2248 end-of-system-prompt` collapse,
~40k-token full re-prefill, GPU pinned >400W).

Replace the rolling digest with per-message digests in `#messageDigests`,
walk the new sync against them to find the longest byte-stable prefix,
truncate the log down to that prefix via a new `AppendOnlyLog.truncate(count)`,
and only re-append the diverged tail. Genuine compaction (`length <
lastSyncCount`) still clears the log.

- Tail-only rewrite: prefix stays byte-stable; only the trailing message
  re-syncs.
- Deep rewrite: prefix up to the divergence stays byte-stable; the
  provider re-prefills from the divergent message onward (architectural
  minimum).
- True compaction: unchanged, full replay.

Replace the now-misleading `detects in-place rewrite of already-synced
messages` / `detects in-place rewrite via digest mismatch` tests with
`preserves the byte-stable prefix when a deep message is rewritten (#3406)`,
`preserves the prefix when the tail is rewritten (#3406)`, `appended
new messages keep the prefix stable even when the prior tail also
diverged (#3406)`, and `rewriting the first message still re-syncs
from scratch` so each invariant is asserted directly.

Fixes #3406
2026-06-24 22:33:22 +00:00
can1357 c053afa087 test(agent): migrated yield tests to use YieldGate for stability
- Replaced global timer mocks with local `YieldGate` instances to avoid test flakiness from concurrent environment interference.
- Introduced an injected clock and counting sleep pattern to test gating logic without relying on `process` globals.
- Added a test case to ensure the gate correctly handles negative clock jumps without stalling.
2026-06-24 15:30:29 +02:00
oldschoola c210bdd391 fix(proxy): scrub partialJson in catch-block error path on stream disconnect
Address review feedback: when the SSE stream disconnects after a
toolcall_delta but before toolcall_end/done/error, the catch-block
at lines 183-192 pushes the partial message as the error result
without calling scrubPartialJson. This leaked the internal partialJson
field into the final error message.

Added scrubPartialJson(partial) call in the catch block, before
pushing the error event.

Added test verifying partialJson does not leak when server disconnects
mid-tool-call (toolcall_start + partial toolcall_delta, no terminal event).
2026-06-23 10:16:06 -07:00
oldschoola 2a2aa9fe59 fix(proxy): preserve partialJson on content during streaming, scrub at terminal events
Address review feedback: downstream renderers (event-controller.ts:535)
read content.partialJson during toolcall_delta to pace streaming
previews (bash env assignments, write/edit smooth streaming).

Revised approach:
- toolcall_start: initialize partialJson on content via typed
  ToolCall & { partialJson: string } intersection (not as any)
- toolcall_delta: accumulate in side-channel Map, write onto content
  via typed intersection cast
- toolcall_end: delete partialJson from content + side-channel map
- done/error: scrubPartialJson() cleans any remaining blocks that
  never got toolcall_end (the original leak bug, now fixed for all
  terminal paths)

Added test verifying partialJson IS present during streaming and
IS absent after completion.
2026-06-23 09:40:32 -07:00
oldschoola 9c795f886e fix(proxy): use side-channel Map for partialJson, eliminating as-any casts
streamProxy stored internal partialJson streaming state directly on typed
ToolCall objects via 4 'as any' casts. If toolcall_end was skipped (stream
error, early done), the field leaked into the final AssistantMessage content,
corrupting downstream serialization.

Replace with a side-channel Map<number, string> keyed by contentIndex:
- toolcall_start initializes the map entry
- toolcall_delta accumulates into it
- toolcall_end cleans it up

The typed ToolCall object never carries non-spec fields. All 4 'as any'
casts are eliminated.

Added 4 contract tests covering argument parsing, partialJson isolation on
normal completion, partialJson isolation when toolcall_end is missing, and
multiple concurrent tool calls with interleaved deltas.
2026-06-23 08:17:29 -07:00
can1357 5c21b28786 feat: optimized handoff generation and harden request safety
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
2026-06-22 20:05:40 +02:00
can1357 0b5e2d8276 fix(ai-providers): normalized Anthropic tool call IDs
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
2026-06-21 06:04:49 +02:00