Commit Graph
1584 Commits
Author SHA1 Message Date
can1357 da794ebb24 Merge PR #6938: feat(coding-agent): allow checkpoint/rewind/learn/manage_skill in subagents when explicitly requested (@szavadsky) 2026-07-30 01:26:49 +02:00
can1357 adad262ba9 fix(xdev): include truncation marker in summary byte cap
(cherry picked from commit aa2067bf7952191beab85b71002aafe812f544bc)
2026-07-29 23:09:01 +02:00
can1357 4666b1ae41 Merge PR #7010: fix(xdev): bound device summaries in UTF-8 bytes and flag untrusted metadata (@terrxo) 2026-07-29 23:09:01 +02:00
Nik Divjakandcan1357 630f9e5324 fix(xdev): bound device summaries in UTF-8 bytes and flag untrusted metadata
Catalog summaries of mounted xd:// devices are inlined verbatim into the
system prompt. External devices (MCP servers, plugins) supply that text, and
it was bounded only by character count: a summary of multi-byte script passed
roughly three times the intended budget, and control characters survived into
the prompt where they can forge structure.

Summaries now go through a single sanitize-and-bound step that strips C0/C1
control characters and bounds the result in UTF-8 bytes via the central
truncateHeadBytes helper, so a cut lands on a code point boundary and never
renders a partial code point. The built-in/external distinction is derived
once per entry, and that same boolean both selects the description cap and is
exposed as `dynamic`, so the cap and the flag cannot disagree. The prompt uses
the flag to state that dynamic summaries are untrusted metadata, and the mount
notice says the same for newly appeared devices.

(cherry picked from commit 5989da6235d820bc687779a791e655e6f1b2df0f)
2026-07-29 23:09:00 +02:00
can1357 4cc259f6cf Merge PR #6934: fix(hashline,coding-agent): key snapshot tag on actually-persisted content after ACP bridge writes (@marton78) 2026-07-29 23:08:39 +02:00
Márton Danóczyandcan1357 7708f372b5 fix(hashline,coding-agent): key snapshot tag on actually-persisted content after ACP bridge writes
Root cause of the reported "edit tool silently reformats the whole
file" corruption: fs/write_text_file has no verbatim guarantee. When
an ACP client (e.g. Zed with format_on_save: on) reformats a buffer
on save, routeWriteThroughBridge reported the pre-write content as
successfully written, and Patcher.commit keyed the returned snapshot
tag on that same pre-write text instead of what actually landed on
disk. The next edit anchored on that tag then resolved hunks against
a baseline the file had already drifted away from, which is what
produced whole-file "corruption" from single-line hunks -- reproduced
live in this session against real Swift/JSON/TypeScript files with
Zed as the ACP client.

- routeWriteThroughBridge reads the file back after the bridge write
  and returns the verified content plus a drift flag (best-effort:
  ACP defines no ordering between the client acking the write and its
  own async format-on-save settling, so this degrades gracefully to
  the old stale-tag-on-next-read failure mode, never to corruption).
- HashlineFilesystem.writeText propagates that verified content in
  view-space (the same space readText returns -- e.g. a notebook's
  editable cell text, not its raw JSON), not storage-space, so tag
  validation on the next edit compares like with like.
- Patcher.commit keys fileHash/header/snapshot on the verified
  post-write content (normalized, so BOM/line-ending restoration never
  produces a false "drift") when it diverges from what was sent, and
  appends a warning naming the drift -- but deliberately leaves the
  returned `after` (and therefore the model-visible diff) scoped to
  the intended hunk. Diffing against the full drifted file would
  balloon the tool response to span every reformatted line (measured
  ~6.8x inflation on a 245-line file with one touched line); the
  warning is the correct O(1) channel for "your editor reformatted
  this," not an O(file-size) diff.
- write.ts keys its own snapshot header on the verified bridge content
  too (no diff-size concern there since write always replaces the
  whole file).

Caught via code review (dispatched against the first pass of this
fix): a naive "just use the verified content everywhere" fix broke
.ipynb editing outright (write-space vs read-space content mismatch,
tag invalid on every notebook edit) and would have inflated every
drifted edit response by ~6.8x. Both are now covered by regression
tests that fail against the pre-fix code and pass against this one.

(cherry picked from commit 35ab80e43be5800b2f48728e4400eb9fd7f7f7d2)
2026-07-29 23:08:38 +02:00
usr-bin-roygbivandcan1357 8b81b1c0a0 fix(launch): preserve logs on replay failure
(cherry picked from commit 2bebc32a05553e75edef16f71218fa2bfbfc1f77)
2026-07-29 23:08:21 +02:00
usr-bin-roygbivandcan1357 4436038127 fix(launch): isolate legacy xterm replay
(cherry picked from commit 47baab894a224f3354085b4794337a211d3cf5c8)
2026-07-29 23:08:21 +02:00
Slava Zavadsky fdd46bd971 fix(coding-agent): pair checkpoint/rewind for restricted sessions too
The !restrictToolNames guard on the pairing blocks was wrong: a restricted
session with tools:[checkpoint] passes isToolAllowed (requestedTools is
defined) but the pairing is skipped, stranding the agent without rewind.
Remove the guard — this is a safety pairing, not a convenience widening.
Added restricted-session tests in both createTools and SDK active-set paths.
2026-07-28 21:18:43 -04:00
Slava Zavadsky 4364cfa5b3 fix(coding-agent): mirror checkpoint/rewind pairing in SDK active-tool path
Address Codex review: createTools auto-includes the sister tool in the
registry, but createAgentSession rebuilds the active set from the original
toolNames — so a one-sided tools: entry left the sister tool registered
but inactive. Mirror the pairing into explicitlyRequestedToolNames, gated
to !restrictToolNames for consistency with the manage_skill/learn mirror.

Also gate the index.ts pairing block with !restrictToolNames to match its
AST/auto-learn siblings, and fix the prompt to say 'or' not 'and'.
2026-07-28 20:27:14 -04:00
Slava Zavadsky abef08116e fix(coding-agent): auto-pair checkpoint/rewind and add changelog
Address review feedback on PR #6938:
- One-sided tools: list checkpoint without rewind (or vice versa) now
  auto-includes the sister tool, preventing a stuck subagent
- Add changelog entry under [Unreleased]
2026-07-28 18:12:01 -04:00
Slava Zavadsky afa76546a6 feat(coding-agent): allow checkpoint/rewind/learn/manage_skill in subagents when explicitly requested
Closes #3762

When an agent definition's frontmatter  list explicitly includes
checkpoint, rewind, learn, or manage_skill, allow them in subagents.
Previously all four were hard-gated to top-level sessions.

- Relax taskDepth gates in isToolAllowed using the already-captured
  requestedTools variable (no signature change needed)
- Remove isTopLevelSession function and its 4 guard sites from checkpoint.ts
- Update checkpoint prompt with enablement docs
- Add tests for subagent explicit-request, no-request, disabled-setting,
  and top-level paths
2026-07-28 17:57:02 -04:00
can1357 84a7937325 docs(tools): note the literal-path escape in the selector-list guard
The guard now probes the full target before refusing, so an existing file
named like a selector list stays writable. Say so in the CHANGELOG entry
and the readSelectorListMisfire doc comment.
2026-07-28 10:59:36 +02:00
can1357 c66fac2ddd fix(tools): let an existing literal selector-list filename stay writable
Probe the full target with probeLiteralPathExists before classifying it as a
mis-dispatched read-selector list, matching the single-selector guard, so an
existing POSIX file like 'report:1-2;archive:3-4' can still be overwritten.
2026-07-28 10:59:36 +02:00
can1357 fefccc4acb Merge PR #6811: fix(tools): refuse write targets shaped as a read-selector list (@roboomp) 2026-07-28 10:59:36 +02:00
can1357 a4f4b3766f fix(coding-agent): budget Ask markdown against padded block width
The symmetric-padding change shrank the framed-block content width to
outputBlockContentWidth(width); the Ask renderer still pre-rendered its
question/result Markdown at the old width-2 budget, so maximal-width rows
re-wrapped inside the block and spilled a trailing fragment row.
2026-07-28 10:59:33 +02:00
can1357 05e98891f0 Merge PR #6833: fix(tui): wrapped Markdown list and output block layout (@usr-bin-roygbiv) 2026-07-28 10:59:33 +02:00
usr-bin-roygbiv a28ca5a268 fix(tui): honor padded markdown width budgets 2026-07-28 07:41:59 +00:00
can1357 169a1b81c7 feat(coding-agent/goals): replaced guided goal modal workflow with interview brief
- Reworked the `/guided-goal` command to send a hidden interview brief instead of a modal popup flow.
- Removed the deprecated `guided-setup.ts` module and system prompt template.
- Updated goal tool availability and activation logic to support goal creation during the interview.
- Replaced existing tests and added new verification for the updated guided-goal workflow.
2026-07-28 08:55:49 +02:00
can1357 e7009452b4 feat(coding-agent/tools): simplified browser screenshot persistence and return paths
- Remove the per-call `save` option from `tab.screenshot()` to simplify usage.
- Update `tab.screenshot()` to return the saved file path as a promise string.
- Configure screenshot persistence to use daemon path or custom `browser.screenshotDir`.
- Add comprehensive tests verifying temp path return and custom directory saving.
2026-07-28 06:40:31 +02:00
can1357 21b3764b08 refactor(coding-agent): replaced xdevregistry with state interface and helpers
- Replaced the `XdevRegistry` class with the `XdevState` interface and pure helper functions across core and session tools.
- Updated session configurations, tool execution, and renderers to utilize canonical tool map initialization and sharing.
- Adapted unit tests and mocks to use `XdevState` and associated helper functions for permission and dispatch verification.
2026-07-28 03:34:36 +02:00
can1357 bbe0236a0f fix(coding-agent): prevented duplicate artifact saves and mismatched ids in bash output 2026-07-28 02:05:55 +02:00
can1357 392aac4e49 fix(coding-agent): resolved third review pass on inspect_image vision mode
- read now treats an xd://-mounted inspect_image as available (top-level
  predicate OR mounted device gated by the effective mode), so default
  xdev sessions with a text-only model keep metadata-guidance reads
  instead of inlining images the provider boundary would scrub
- advisor tool session stops inheriting the primary's isToolActive and
  xdevRegistry: advisors cannot execute xd:// devices, so their reads
  inline images again
- setModelWithProviderSessionReset is now async and awaited at every
  callsite, so retry-fallback model switches cannot race the
  inspect_image tool-slate reconcile
- regression tests for both xd:// availability directions
2026-07-27 23:07:26 +02:00
alexis@epsilver.xyz b58a943a8e fix(coding-agent): address second codex pass on vision mode
- read now derives its image behavior from actual tool availability
  (session.isToolActive) with the mode computation as fallback, so
  restricted sessions whose explicit slate omits inspect_image (e.g.
  subagents) never get metadata-only reads pointing at an absent tool
- reconcile passes the post-change availability into the read
  description sync, keeping the advertised prompt correct across flips
  in both directions and when tool construction fails
- flat quoted-dotted inspect_image.mode is normalized into the nested
  target during migration instead of being silently dropped when a
  legacy flat enabled key is present
- regression tests for all three: availability-driven read behavior,
  flat+flat migration, description advertising
2026-07-27 16:24:02 -04:00
alexis@epsilver.xyz 3854c1c3b1 fix(coding-agent): address codex review on vision mode
- Reconcile inspect_image centrally from setModelWithProviderSessionReset
  so retry-fallback model changes (turn-recovery.ts) that bypass
  syncAfterModelChange cannot leave a stale tool set
- Apply persisted inspect_image.mode changes immediately from the
  settings selector via a new handleSettingChange branch
- Refresh the read tool's advertised description during reconciliation,
  before applyActiveToolsByName rebuilds the prompt, instead of only
  lazily on the next image read
- Fix the flat (quoted-dotted) enabled->mode migration to write the
  nested target form the resolver actually reads
- Add committed regression tests: tri-state x capability matrix,
  override precedence, and enabled->mode migration (nested, flat, and
  explicit-mode-wins)
2026-07-27 16:24:02 -04:00
alexis@epsilver.xyz c33b98e260 feat(coding-agent): capability-aware inspect_image with tri-state mode and /vision toggle
Replace the inspect_image.enabled boolean with inspect_image.mode
(auto|on|off, default auto). In auto the tool is registered only when
the active model lacks native image input, so vision-capable models
(e.g. kimi-code/k3) read images inline with their own capabilities
instead of delegating to a separate vision model. on/off force
registration regardless of model capability.

- New utils/inspect-image-mode.ts resolves the effective state from the
  /vision session override, the persisted setting, and model capability
- read tool re-evaluates the effective state per image read and
  re-renders its description, so it returns decoded image blocks again
  whenever inspect_image is hidden
- /vision [on|off|auto|status] slash command (modeled on /computer)
  overrides the mode for the current session only
- Tool set is reconciled on model switch with a status notice when
  inspect_image appears/disappears
- Legacy inspect_image.enabled true/false migrates to mode on/off
2026-07-27 16:24:02 -04:00
can1357 222e3569bd perf: optimized markdown rendering speed and device doc schemas
- Render device doc parameter schemas as TypeScript types instead of raw JSON schema dumps.
- Optimize marked streaming block rules and add pre-gates to reduce CPU overhead.
- Increase the markdown render cache entry budget from 32 KiB to 256 KiB.
2026-07-27 22:19:51 +02:00
can1357 4210727227 feat(coding-agent): integrated v8 cpuprofile parsing into read tool
- Added V8 `.cpuprofile` parser and bottleneck summary generation utilities.
- Integrated profile summary rendering into the read tool execution.
- Refactored profile rendering machinery into shared tree utilities.
- Added comprehensive unit and integration tests for cpuprofile parsing and read tool dispatch.
2026-07-27 20:37:46 +02:00
can1357 7db4463da6 feat: introduced parser for macos sample reports and updated models
- Added a new parser and bottleneck summary renderer for macOS `/usr/bin/sample` reports with symbol demangling.
- Integrated automated summary parsing for sample reports into the ReadTool.
- Updated model configurations and pricing parameters across multiple providers.
- Added comprehensive unit and integration tests for sample profile parsing and ReadTool integration.
2026-07-27 20:08:30 +02:00
can1357 d16a251777 chore: reorg tests 2026-07-27 16:43:53 +02:00
roboomp 4ac322db93 fix(tools): refuse write targets shaped as a read-selector list
The read-selector-misfire guard (#6123/#6387) short-circuited whenever
`content` was non-empty, so a semicolon-joined list of read selectors
(`a.txt:1-2;b/c.txt:3-4`) passed as a write path with content fell through
to ordinary filesystem creation and silently built a nested directory tree
in the workspace. `read` accepts no such list, so this shape is always a
mis-dispatched multi-file read.

Refuse any target that splits on `;` into 2+ segments each carrying its own
read selector, regardless of `content` — the non-empty-content escape hatch
covers a lone selector-shaped filename, never a `;`-list.

Fixes #6809
2026-07-27 14:20:33 +00:00
Dongmen Laohu 2c77c8535a fix(mcp): resolve native resource URIs 2026-07-27 18:59:12 +08:00
can1357 e2ed58d4d5 Merge PR #4168: fix(inspect-image): bounded per-request timeout on the vision-model call (@roboomp)
# Conflicts:
#	packages/coding-agent/src/config/settings-schema.ts
2026-07-27 04:59:19 +02:00
can1357 dccf0d3c8a Merge PR #6748: perf(launch): isolate PTY replay in broker (@usr-bin-roygbiv) 2026-07-27 04:58:28 +02:00
can1357 48a4f1b1e6 Merge PR #6742: perf(coding-agent): lazily construct computer schema (@usr-bin-roygbiv) 2026-07-27 04:58:27 +02:00
can1357 bf8e07ae1b Merge PR #6733: fix(coding-agent): safely glob memory directories (@usr-bin-roygbiv) 2026-07-27 04:58:26 +02:00
usr-bin-roygbiv 2755226e59 style(coding-agent): format schema parity type 2026-07-27 01:21:27 +00:00
usr-bin-roygbiv 403f90783e perf(launch): isolate PTY replay in broker 2026-07-27 01:17:19 +00:00
usr-bin-roygbiv 76425c45bb refactor(coding-agent): assert computer schema type parity 2026-07-27 00:51:35 +00:00
usr-bin-roygbiv 73add3f98c perf(coding-agent): lazily construct computer schema 2026-07-27 00:39:09 +00:00
usr-bin-roygbiv 49d5145fff fix(coding-agent): preserve encoded memory glob bases 2026-07-26 21:33:13 +00:00
can1357 f8dbb3669f feat(utils): implemented windows shell resolution for bash execution
- Added resolveWindowsShell to locate Git Bash, scoop installs, and path binaries with a fallback to cmd.exe.
- Updated bash-executor to prevent wrapping user commands in cmd.exe when using fallback shell paths.
- Updated installation script to report optional shell status rather than failing when bash is absent.
2026-07-26 20:27:16 +02:00
Roy 90cb91489e fix(coding-agent): safely glob memory directories 2026-07-26 16:36:47 +00:00
can1357 d1dc436d1e Merge PR #6596: fix(coding-agent): preserve Claude computer coordinates (@wolfiesch)
# Conflicts:
#	docs/computer-use.md
#	packages/coding-agent/src/tools/computer.ts
#	packages/coding-agent/test/tools/computer.test.ts
2026-07-26 15:44:51 +02:00
can1357 56a8ab1413 Merge PR #6697: fix(coding-agent): match bash deny/prompt patterns per command segment (@roboomp) 2026-07-26 15:41:28 +02:00
can1357 9b651f2869 Merge PR #6616: fix(cursor): sync native todo list from server-resolved tool calls (@quantmind-br) 2026-07-26 15:41:27 +02:00
Diogo Soares Rodrigues 9088fe821b fix(cursor): await error-drain transforms and sanitize mirrored todo labels
Two boundary defects on the mirrored-todo path.

1. The Agent error drain snapshotted #cursorToolResultBuffer without
   awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
   async cursorOnToolResult still running when the provider errored
   patched an entry the catch path had already detached, so the
   pre-transform payload was persisted. A provider error is exactly when
   a transform is most likely to be in flight.

2. The todo renderer interpolated mirrored provider text straight into
   terminal output. A Cursor snapshot carries model-authored task
   content, phase names and summary text verbatim, so a label holding
   ANSI/C0 sequences rewrote the terminal on every render and replay.
   sanitizeText alone is not enough - it preserves tabs, which punch
   holes in bordered output - so every display path now funnels through
   one forDisplay() helper: task labels, blocker notes, phase headers,
   the zero-task fallback, and the streaming renderCall preview. Raw
   values are untouched; content and phase name are the identity keys
   the local list is looked up by and what gets persisted.
2026-07-26 09:29:13 -03:00
roboomp e650607cad fix(coding-agent): segment bash approval commands with shell-aware tokenizer
The regex splitter only recognized `&&`, `||`, `;`, `|` and newlines, so a
single `&` (background operator) — also a command terminator — slipped a
dangerous command past a deny rule (`sleep 1 & rm -rf /tmp/x`), which under
approvalMode: yolo executed with no prompt.

Extract the shell-aware tokenizer from gh-cache-invalidation into a shared
tools/shell-tokenize.ts and reuse it for deny/prompt segmentation. It honors
every command boundary (`&&`, `||`, `;`, `|`, single `&`, subshells,
newlines) plus quoting and escapes, so both callers share one implementation.

Fixes #6695
2026-07-26 11:36:19 +00:00
roboomp 85110c366c fix(coding-agent): match bash deny/prompt patterns per command segment
bashApprovalPatternToRegExp anchors globs with ^...$ against the whole
normalized command, so a bash.patterns deny rule only fired when the
dangerous command was first in the line. A compound command such as
`cd /tmp && rm -rf /tmp/x` bypassed the rule and, under approvalMode:
yolo, executed with no prompt -- deny is the guard that outranks yolo.

deny/prompt rules now match the whole command or any single segment
(split on &&, ||, ;, |, newlines). allow rules still require the entire
command to match and never apply to compound lines, so a narrow allow
cannot vouch for a smuggled unsafe segment.

Fixes #6695
2026-07-26 11:28:48 +00:00
Wolfgang Schoenberger 0c5fa60c77 fix(coding-agent): cap captures for resizing transports
Use the resolved supportsImageDetailOriginal capability to constrain native computer frames whenever a Responses transport clamps screenshot detail to auto. Preserve the established Claude-family fallback for transports that do not expose this capability.

Cover Copilot GPT-5 Responses through real catalog model resolution and the controller's observable capture options.
2026-07-26 02:53:01 -07:00