Catalog summaries of mounted xd:// devices are inlined verbatim into the
system prompt. External devices (MCP servers, plugins) supply that text, and
it was bounded only by character count: a summary of multi-byte script passed
roughly three times the intended budget, and control characters survived into
the prompt where they can forge structure.
Summaries now go through a single sanitize-and-bound step that strips C0/C1
control characters and bounds the result in UTF-8 bytes via the central
truncateHeadBytes helper, so a cut lands on a code point boundary and never
renders a partial code point. The built-in/external distinction is derived
once per entry, and that same boolean both selects the description cap and is
exposed as `dynamic`, so the cap and the flag cannot disagree. The prompt uses
the flag to state that dynamic summaries are untrusted metadata, and the mount
notice says the same for newly appeared devices.
(cherry picked from commit 5989da6235d820bc687779a791e655e6f1b2df0f)
Root cause of the reported "edit tool silently reformats the whole
file" corruption: fs/write_text_file has no verbatim guarantee. When
an ACP client (e.g. Zed with format_on_save: on) reformats a buffer
on save, routeWriteThroughBridge reported the pre-write content as
successfully written, and Patcher.commit keyed the returned snapshot
tag on that same pre-write text instead of what actually landed on
disk. The next edit anchored on that tag then resolved hunks against
a baseline the file had already drifted away from, which is what
produced whole-file "corruption" from single-line hunks -- reproduced
live in this session against real Swift/JSON/TypeScript files with
Zed as the ACP client.
- routeWriteThroughBridge reads the file back after the bridge write
and returns the verified content plus a drift flag (best-effort:
ACP defines no ordering between the client acking the write and its
own async format-on-save settling, so this degrades gracefully to
the old stale-tag-on-next-read failure mode, never to corruption).
- HashlineFilesystem.writeText propagates that verified content in
view-space (the same space readText returns -- e.g. a notebook's
editable cell text, not its raw JSON), not storage-space, so tag
validation on the next edit compares like with like.
- Patcher.commit keys fileHash/header/snapshot on the verified
post-write content (normalized, so BOM/line-ending restoration never
produces a false "drift") when it diverges from what was sent, and
appends a warning naming the drift -- but deliberately leaves the
returned `after` (and therefore the model-visible diff) scoped to
the intended hunk. Diffing against the full drifted file would
balloon the tool response to span every reformatted line (measured
~6.8x inflation on a 245-line file with one touched line); the
warning is the correct O(1) channel for "your editor reformatted
this," not an O(file-size) diff.
- write.ts keys its own snapshot header on the verified bridge content
too (no diff-size concern there since write always replaces the
whole file).
Caught via code review (dispatched against the first pass of this
fix): a naive "just use the verified content everywhere" fix broke
.ipynb editing outright (write-space vs read-space content mismatch,
tag invalid on every notebook edit) and would have inflated every
drifted edit response by ~6.8x. Both are now covered by regression
tests that fail against the pre-fix code and pass against this one.
(cherry picked from commit 35ab80e43be5800b2f48728e4400eb9fd7f7f7d2)
The !restrictToolNames guard on the pairing blocks was wrong: a restricted
session with tools:[checkpoint] passes isToolAllowed (requestedTools is
defined) but the pairing is skipped, stranding the agent without rewind.
Remove the guard — this is a safety pairing, not a convenience widening.
Added restricted-session tests in both createTools and SDK active-set paths.
Address Codex review: createTools auto-includes the sister tool in the
registry, but createAgentSession rebuilds the active set from the original
toolNames — so a one-sided tools: entry left the sister tool registered
but inactive. Mirror the pairing into explicitlyRequestedToolNames, gated
to !restrictToolNames for consistency with the manage_skill/learn mirror.
Also gate the index.ts pairing block with !restrictToolNames to match its
AST/auto-learn siblings, and fix the prompt to say 'or' not 'and'.
Address review feedback on PR #6938:
- One-sided tools: list checkpoint without rewind (or vice versa) now
auto-includes the sister tool, preventing a stuck subagent
- Add changelog entry under [Unreleased]
Closes#3762
When an agent definition's frontmatter list explicitly includes
checkpoint, rewind, learn, or manage_skill, allow them in subagents.
Previously all four were hard-gated to top-level sessions.
- Relax taskDepth gates in isToolAllowed using the already-captured
requestedTools variable (no signature change needed)
- Remove isTopLevelSession function and its 4 guard sites from checkpoint.ts
- Update checkpoint prompt with enablement docs
- Add tests for subagent explicit-request, no-request, disabled-setting,
and top-level paths
The guard now probes the full target before refusing, so an existing file
named like a selector list stays writable. Say so in the CHANGELOG entry
and the readSelectorListMisfire doc comment.
Probe the full target with probeLiteralPathExists before classifying it as a
mis-dispatched read-selector list, matching the single-selector guard, so an
existing POSIX file like 'report:1-2;archive:3-4' can still be overwritten.
The symmetric-padding change shrank the framed-block content width to
outputBlockContentWidth(width); the Ask renderer still pre-rendered its
question/result Markdown at the old width-2 budget, so maximal-width rows
re-wrapped inside the block and spilled a trailing fragment row.
- Reworked the `/guided-goal` command to send a hidden interview brief instead of a modal popup flow.
- Removed the deprecated `guided-setup.ts` module and system prompt template.
- Updated goal tool availability and activation logic to support goal creation during the interview.
- Replaced existing tests and added new verification for the updated guided-goal workflow.
- Remove the per-call `save` option from `tab.screenshot()` to simplify usage.
- Update `tab.screenshot()` to return the saved file path as a promise string.
- Configure screenshot persistence to use daemon path or custom `browser.screenshotDir`.
- Add comprehensive tests verifying temp path return and custom directory saving.
- Replaced the `XdevRegistry` class with the `XdevState` interface and pure helper functions across core and session tools.
- Updated session configurations, tool execution, and renderers to utilize canonical tool map initialization and sharing.
- Adapted unit tests and mocks to use `XdevState` and associated helper functions for permission and dispatch verification.
- read now treats an xd://-mounted inspect_image as available (top-level
predicate OR mounted device gated by the effective mode), so default
xdev sessions with a text-only model keep metadata-guidance reads
instead of inlining images the provider boundary would scrub
- advisor tool session stops inheriting the primary's isToolActive and
xdevRegistry: advisors cannot execute xd:// devices, so their reads
inline images again
- setModelWithProviderSessionReset is now async and awaited at every
callsite, so retry-fallback model switches cannot race the
inspect_image tool-slate reconcile
- regression tests for both xd:// availability directions
- read now derives its image behavior from actual tool availability
(session.isToolActive) with the mode computation as fallback, so
restricted sessions whose explicit slate omits inspect_image (e.g.
subagents) never get metadata-only reads pointing at an absent tool
- reconcile passes the post-change availability into the read
description sync, keeping the advertised prompt correct across flips
in both directions and when tool construction fails
- flat quoted-dotted inspect_image.mode is normalized into the nested
target during migration instead of being silently dropped when a
legacy flat enabled key is present
- regression tests for all three: availability-driven read behavior,
flat+flat migration, description advertising
- Reconcile inspect_image centrally from setModelWithProviderSessionReset
so retry-fallback model changes (turn-recovery.ts) that bypass
syncAfterModelChange cannot leave a stale tool set
- Apply persisted inspect_image.mode changes immediately from the
settings selector via a new handleSettingChange branch
- Refresh the read tool's advertised description during reconciliation,
before applyActiveToolsByName rebuilds the prompt, instead of only
lazily on the next image read
- Fix the flat (quoted-dotted) enabled->mode migration to write the
nested target form the resolver actually reads
- Add committed regression tests: tri-state x capability matrix,
override precedence, and enabled->mode migration (nested, flat, and
explicit-mode-wins)
Replace the inspect_image.enabled boolean with inspect_image.mode
(auto|on|off, default auto). In auto the tool is registered only when
the active model lacks native image input, so vision-capable models
(e.g. kimi-code/k3) read images inline with their own capabilities
instead of delegating to a separate vision model. on/off force
registration regardless of model capability.
- New utils/inspect-image-mode.ts resolves the effective state from the
/vision session override, the persisted setting, and model capability
- read tool re-evaluates the effective state per image read and
re-renders its description, so it returns decoded image blocks again
whenever inspect_image is hidden
- /vision [on|off|auto|status] slash command (modeled on /computer)
overrides the mode for the current session only
- Tool set is reconciled on model switch with a status notice when
inspect_image appears/disappears
- Legacy inspect_image.enabled true/false migrates to mode on/off
- Render device doc parameter schemas as TypeScript types instead of raw JSON schema dumps.
- Optimize marked streaming block rules and add pre-gates to reduce CPU overhead.
- Increase the markdown render cache entry budget from 32 KiB to 256 KiB.
- Added V8 `.cpuprofile` parser and bottleneck summary generation utilities.
- Integrated profile summary rendering into the read tool execution.
- Refactored profile rendering machinery into shared tree utilities.
- Added comprehensive unit and integration tests for cpuprofile parsing and read tool dispatch.
- Added a new parser and bottleneck summary renderer for macOS `/usr/bin/sample` reports with symbol demangling.
- Integrated automated summary parsing for sample reports into the ReadTool.
- Updated model configurations and pricing parameters across multiple providers.
- Added comprehensive unit and integration tests for sample profile parsing and ReadTool integration.
The read-selector-misfire guard (#6123/#6387) short-circuited whenever
`content` was non-empty, so a semicolon-joined list of read selectors
(`a.txt:1-2;b/c.txt:3-4`) passed as a write path with content fell through
to ordinary filesystem creation and silently built a nested directory tree
in the workspace. `read` accepts no such list, so this shape is always a
mis-dispatched multi-file read.
Refuse any target that splits on `;` into 2+ segments each carrying its own
read selector, regardless of `content` — the non-empty-content escape hatch
covers a lone selector-shaped filename, never a `;`-list.
Fixes#6809
- Added resolveWindowsShell to locate Git Bash, scoop installs, and path binaries with a fallback to cmd.exe.
- Updated bash-executor to prevent wrapping user commands in cmd.exe when using fallback shell paths.
- Updated installation script to report optional shell status rather than failing when bash is absent.
Two boundary defects on the mirrored-todo path.
1. The Agent error drain snapshotted #cursorToolResultBuffer without
awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
async cursorOnToolResult still running when the provider errored
patched an entry the catch path had already detached, so the
pre-transform payload was persisted. A provider error is exactly when
a transform is most likely to be in flight.
2. The todo renderer interpolated mirrored provider text straight into
terminal output. A Cursor snapshot carries model-authored task
content, phase names and summary text verbatim, so a label holding
ANSI/C0 sequences rewrote the terminal on every render and replay.
sanitizeText alone is not enough - it preserves tabs, which punch
holes in bordered output - so every display path now funnels through
one forDisplay() helper: task labels, blocker notes, phase headers,
the zero-task fallback, and the streaming renderCall preview. Raw
values are untouched; content and phase name are the identity keys
the local list is looked up by and what gets persisted.
The regex splitter only recognized `&&`, `||`, `;`, `|` and newlines, so a
single `&` (background operator) — also a command terminator — slipped a
dangerous command past a deny rule (`sleep 1 & rm -rf /tmp/x`), which under
approvalMode: yolo executed with no prompt.
Extract the shell-aware tokenizer from gh-cache-invalidation into a shared
tools/shell-tokenize.ts and reuse it for deny/prompt segmentation. It honors
every command boundary (`&&`, `||`, `;`, `|`, single `&`, subshells,
newlines) plus quoting and escapes, so both callers share one implementation.
Fixes#6695
bashApprovalPatternToRegExp anchors globs with ^...$ against the whole
normalized command, so a bash.patterns deny rule only fired when the
dangerous command was first in the line. A compound command such as
`cd /tmp && rm -rf /tmp/x` bypassed the rule and, under approvalMode:
yolo, executed with no prompt -- deny is the guard that outranks yolo.
deny/prompt rules now match the whole command or any single segment
(split on &&, ||, ;, |, newlines). allow rules still require the entire
command to match and never apply to compound lines, so a narrow allow
cannot vouch for a smuggled unsafe segment.
Fixes#6695
Use the resolved supportsImageDetailOriginal capability to constrain native computer frames whenever a Responses transport clamps screenshot detail to auto. Preserve the established Claude-family fallback for transports that do not expose this capability.
Cover Copilot GPT-5 Responses through real catalog model resolution and the controller's observable capture options.