Commit Graph

15269 Commits

Author SHA1 Message Date
can1357 3d6b5b7ac8 chore: bump version to 17.2.0 2026-07-30 07:58:50 +02:00
can1357 a4773c0baa refactor(typescript-edit-benchmark): updated mutation plans
- Update benchmark fixtures archive file.
- Adjust mutation plan block sizes and counts in generator script.
- Remove postmortem quit test.
2026-07-30 07:57:55 +02:00
can1357 38ebd30337 chore: update stale tests 2026-07-30 07:49:43 +02:00
can1357 9856f904d7 feat(typescript-edit-benchmark): introduced empirical edit mutation planning
- Add new structural, multi-edit, and block-level mutation classes with updated category mappings.
- Introduce hunk extraction, placement, rendering, and solver utilities along with unit tests.
- Implement size-based mutation planning, prompt validation logic, and new prompt markdown templates.
- Update benchmark generation scripts and package configurations to support empirical edit shape statistics.
2026-07-30 07:45:01 +02:00
can1357 e05f229f43 refactor: standardized editing syntax by removing copy and delete operations
- Removed copy and delete operations across tokenizer, parser, grammar, and clipboard logic.
- Standardized line-editing operations and block resolvers to use cut exclusively.
- Updated documentation, prompts, and test suites to reflect the removal of copy and delete syntax.
2026-07-30 07:42:48 +02:00
can1357 b9ae2a3f9b docs(hashline): condensed and clarified instructions in prompt documentation
- Condensed instructions and rule descriptions in `packages/hashline/src/prompt.md`.
- Streamlined formatting examples and anti-patterns for clarity.
- Clarified block operation boundaries and markdown heading section rules.
2026-07-30 07:21:44 +02:00
can1357 aaa1ff4e3f refactor(hashline): restructured Lark grammar rules to reduce complexity
- Regrouped `grammar.lark` around shared `target` and `pos` rules, reducing hunk rules from twelve to seven.
- Maintained byte-identical language acceptance while simplifying internal grammar structure.
2026-07-30 07:21:44 +02:00
can1357 9aada058ee feat(hashline): implemented clipboard operations in hashline engine
- Implemented clipboard register management, parsing, and execution rules for CUT, COPY, and PASTE operations in the hashline engine.
- Added session-persistent clipboard state and integration across agent session execution, diff previews, and streaming tools.
- Added comprehensive validation, error messages, recovery handling, and test coverage for clipboard and block operations.
2026-07-30 07:21:43 +02:00
can1357 6b4efa896f feat(coding-agent): implemented oauth credential pin persistence and seeding
- Added hashing utilities and session entry definitions for OAuth credential pins.
- Added session manager methods to append and retrieve credential pins with backdated timestamp support.
- Added credential pin recording after assistant turns and seeding during session restoration.
- Added comprehensive unit tests covering credential pin recording, persistence, and seeding.
2026-07-30 07:21:08 +02:00
can1357 4e5f480f14 feat(metaharness/adapters): implemented interactive cli runner and reporting
- Added comprehensive CLI arguments for thinking levels, timeouts, concurrency, and report formatting.
- Implemented real-time interactive terminal progress rendering and runtime statistics summary.
- Enhanced runner logic with tokenizer-based hashline operation detection and aggregated reporting.
- Added bench:edit script to package configuration and updated default report output paths.
2026-07-30 07:21:08 +02:00
can1357 a6001c04a3 fix(coding-agent): prevented race condition in concurrent createAgentSession calls
- Export AgentRegistry from the SDK to allow passing a private registry instance.
- Provide a dedicated AgentRegistry per in-process client in the benchmark runner.
2026-07-30 06:03:32 +02:00
can1357 00fedcf6ac feat(ai): extended codex and stream timeout defaults to 300 seconds
- Increased codex websocket first event timeout default to 300 seconds.
- Increased default stream idle and first event timeout thresholds to 300 seconds.
- Updated stream timeout test expectations to match new 300s global default.
2026-07-30 05:33:06 +02:00
can1357 d562e53b83 refactor: extracted audio and voice engine into a standalone library crate
- Extracted audio capture and playback implementations, along with the WebRTC peer engine, from `pi-natives` into a new `pi-voice` library crate.
- Updated `pi-natives` bindings to consume the extracted `pi_voice` audio streams and live peer core.
- Added release validation gate jobs, parallelized Linux binary builds, and introduced a concurrent macOS release build job in the CI workflow.
- Updated Bazel workspace configurations, Cargo manifests, and documentation to include the new `pi-voice` crate and its dependencies.
2026-07-30 05:12:40 +02:00
can1357 ccd3bb9565 chore: reformat 2026-07-30 04:25:32 +02:00
can1357 2905a23fc1 fix(coding-agent): prevented duplicate rows in task execution scrollback
- Updated task tool execution to pin live regions and drop partial snapshots once rows commit.
- Tracked background task frozen styled rows and render timestamps to prevent clock drift on committed history.
- Added tests verifying detached and blocking task progress do not duplicate rows in scrollback.
2026-07-30 04:01:06 +02:00
can1357 09545697ee fix(advisor): capture model identity lazily and drop duplicated release notes 2026-07-30 02:01:03 +02:00
can1357 92014ab605 fix: align merged branches with current type contracts 2026-07-30 02:01:03 +02:00
can1357 1b25ff01a2 style: apply biome formatting to merged changes 2026-07-30 02:01:03 +02:00
can1357 5e901c84d2 chore: normalize changelogs after merging open fixes 2026-07-30 02:01:03 +02:00
can1357 c9a5605b5b fix(advisor): degraded reasoning on provider refusals
(cherry picked from commit 776f244d810ac152eb0bd130b36c8474201d0b85)
2026-07-30 02:01:02 +02:00
can1357 7ccbe51b2d fix(session): preserved Codex commentary on empty stops
(cherry picked from commit 0b4887a4098e6296ac72ea2837db22be18dc9894)
2026-07-30 02:01:02 +02:00
can1357 43e5df012e Merge PR #6791: feat(ai): handle Cursor's modern exec wire protocol (@quantmind-br) 2026-07-30 01:48:52 +02:00
can1357 0728a55b8a Merge PR #6731: fix: preserve provider-native compaction semantics (@usr-bin-roygbiv) 2026-07-30 01:48:51 +02:00
can1357 020af06099 Merge PR #6535: feat(mcp): expose server-initiated notifications to extensions via mcp_notification event (@asteriskSF) 2026-07-30 01:48:51 +02:00
can1357 c9c890c49a Merge PR #6858: feat: add opt-in Codex reset fireworks (@joshrzemien) 2026-07-30 01:48:51 +02:00
can1357 2395848f7c Merge PR #6857: feat(coding-agent): add startup changelog display modes (@wolfiesch) 2026-07-30 01:48:51 +02:00
can1357 70e6d2c7dc Merge PR #6680: feat(coding-agent): add opt-in max ceiling for auto thinking (@everton-dgn) 2026-07-30 01:48:51 +02:00
can1357 27cb968359 Merge PR #7007: feat(tools): add a browser.cdpUrl setting for the default automation target (@terrxo) 2026-07-30 01:48:50 +02:00
can1357 31e7e19aa3 test(cursor): updated extension runner fixture 2026-07-30 01:47:13 +02:00
can1357 1efa4e42ac fix(cursor): activated modern exec frames 2026-07-30 01:46:13 +02:00
can1357 34d7f2fb38 fix(cursor): reported full file size for ranged reads 2026-07-30 01:45:41 +02:00
can1357 28b33dd858 test(compaction): aligned provider isolation assertion 2026-07-30 01:45:03 +02:00
can1357 65ef740003 fix(coding-agent): preserved Codex quota identity 2026-07-30 01:43:54 +02:00
Diogo Soares Rodrigues be5292bffd fix(coding-agent): gave the primary Cursor bridge the session's live cwd
The bridge is constructed once, at session creation, and was handed the
startup `cwd` by value. The session's own cwd moves under it — `/cd`,
resume, branch restore all call `sessionManager.moveTo` — and the two
frames that confine a path themselves (the native `delete`, and a
`read_mcp_resource` carrying `download_path`) resolve against whichever cwd
the bridge holds. So after a move the primary deleted or overwrote the
relative path in the workspace the session had left, and reported success
for the path the server actually named.

The advisor bridge already passed a live resolver; this is the same
resolver on the path that was missed. Locked by a wiring test: the seam is
the session handing its handlers to the provider, so the test captures them
there, moves the session, and asserts the frame acts on the new workspace
and leaves the old file alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 079c7ac61104d017eecbf781aa1c58eebd39b0b1)
2026-07-30 01:43:28 +02:00
Diogo Soares Rodrigues 51712bbf4f fix(ai): recorded what the Cursor exec frames actually did
Three places where the answer and the record disagreed.

A windowed `read` reported the window's length as the file's: `total_lines`
and `file_size` were counted off the payload, which is the whole file only
for an unranged read. A 20-line page of a 100-line file went out as
`total_lines: 20` next to `range_applied: true`, which a paginating server
reads as the end of the file. The count now comes from the read's own
record of the file, `details.meta.truncation.totalLines` — deliberately not
the flat `details.truncation.totalLines`, which counts from the window's
start line. Counting the payload remains the answer for a read that
returned the file whole, where it is exact.

`pi_grep`'s `context` and `limit` left no trace. The bridge honors both by
building a scoped tool, and neither is expressible in the model-facing
`grep` schema, so the synthesized block recorded a plain pattern/path
search — replaying a context-widened or capped search as an ordinary grep
beside output no ordinary grep produces. Both are now on the block, the
same way `pi_read` renders its range into the displayed path.

An MCP listing shrank to a count. The full URI/name/mime catalog goes out
on the wire while the paired result recorded `Listed N MCP resource(s)`,
and rebuilt history is serialized from that result — so one reload later
the model knew it had seen N resources and could name none of the URIs a
follow-up read needs. The result now lists what the answer carried, still
derived from the same `execResult` so the block cannot drift from the wire.

`file_size` under a window is left as-is: no source available here records
the file's byte length, and inventing one would trade a visible
inconsistency for an invisible guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 64afea3f7197adc32e2099a7d0ee74408d3012a6)
2026-07-30 01:43:28 +02:00
Diogo Soares Rodrigues 7f97581d0a fix(ai,coding-agent): closed two ways an exec answer misdescribed its own work
A `download_path` naming a FIFO hung the turn outright. The target is
opened write-only, which on POSIX blocks until a reader attaches, so the
`isFile()` refusal sitting behind that open was unreachable — the open
never returned. The path comes from the server, so this needed no planted
file to reach, only a named pipe where a download was aimed. Opening
non-blocking turns a readerless pipe into an immediate refusal and leaves
the existing guard to reject one that has a reader; the flag is inert on
regular files, which is every legitimate target. (The repo already fixed
this shape once, for discovery context-file reads, by stat-gating; the
flag closes the same hole without the stat's TOCTOU window.)

A `pi_grep` that hit the native backend's own match ceiling answered as an
unqualified success. `GrepTool` folds that cap into the flat
`details.truncated` and sets neither `details.truncation` nor
`perFileLimitReached` — the two fields the Pi result reads — so the one
truncation a caller can neither detect nor page around was the one it was
never told about. The flat flag now translates into a `PiTruncation`, and
only once the specific counters came back empty, so a cap that already
reported itself is never restated.

Both regressions are locked: the FIFO test detects a relapse by timing out
rather than by a failed assertion, since a relapse never reaches the
assertion.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 20438ff68cf9c8aaaec30703f5b9972c7bda205e)
2026-07-30 01:43:28 +02:00
Diogo Soares Rodrigues 263573b8ae chore(changelog): restored blank line before the 17.1.8 heading
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit c3c3cc658c329de6a905f2a5dd3ad8beaea29a2b)
2026-07-30 01:43:27 +02:00
Diogo Soares Rodrigues d468942945 chore(changelog): merged duplicate Unreleased headings after upstream merge
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 023d13277cb6a5e2452bbb853978345572754ed9)
2026-07-30 01:43:27 +02:00
Diogo Soares Rodrigues 53408d0412 test(coding-agent): pinned the advisor approval-gate test to a resolvable model
`runs advisor tools through the approval gate` built its advisor from the
`advisor` role chain, which resolves against `modelRegistry.getAvailable()`
— the models the host holds auth for. On a developer box whose environment
carries provider keys the roster resolved and the test passed; in CI, where
the suite's isolated auth storage is empty, every advisor resolved to
`no_model` and `getAdvisorAgent()` returned undefined ("expected an advisor
agent").

The advisor now names `gpt-4o-mini` outright and runs inside the file's
`withProviderAuth` helper, so the roster resolves from the granted key
rather than from whatever the machine happens to have configured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit ef2054da5dbe3d3d1cca4025e12e5c350f47173d)
2026-07-30 01:43:27 +02:00
Diogo Soares Rodrigues 5ab98cded8 fix(cursor): persist resource listings, wire advisor MCP resources, reject unavailable pi edit/write
Three remaining review findings:

- `list_mcp_resources` frames a handler answered now synthesize a
  `list_mcp_resources` block and pair a result derived from the same
  answer sent on the wire; the streamed `ListMcpResourcesToolCall` /
  `ReadMcpResourceToolCall` announcements join the exec-owned set so
  they cannot double-render. No-handler frames still synthesize
  nothing, since nothing ran.

- Advisors receive the same `MCPManager`-backed resource adapter as the
  primary bridge, so their `list_mcp_resources` no longer reports every
  server as empty and `read_mcp_resource` no longer answers `not_found`
  against live connections the advisor shares.

- An unavailable `pi_edit`/`pi_write` answers with the protocol's
  `rejected` variant instead of `error`: refusal and failure are
  separate oneof cases, and a denial reported as an execution error
  invites a retry of an operation that was never permitted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SSWZTe6YA2PX1cqtukZvYi
(cherry picked from commit 47ce936c8df05d6970504af19e5ef7d2e8c38d7b)
2026-07-30 01:43:07 +02:00
Diogo Soares Rodrigues f6f2bab02e fix(cursor): mirror executed pagination in synthesized calls, resolve approval probes from policy
Forwarding the legacy read/grep frames' range and page fixed only the
execution: the transcript block was built from a second translation and
still showed a bare path and an unskipped search. That block is what a
reloaded session replays, so a slice read as the whole file and a later
window presented as page one. Both now come from the shared helpers,
`limit: 0` included -- recorded as the zero lines it returns.

The approval probe answered `approved` unconditionally, which laundered a
configured `deny` into a server-side blessing. It now resolves through a
bridge preflight against the same policy the wrapper applies at execution
time: approved only for a definite allow, refused for a deny, for a mode
demanding a prompt this frame cannot raise, and for an unknown tool.
Still never executes.

Comments and changelog no longer assert server-side semantics for
`range_applied`/`offset_applied`; they describe what the client did,
which is all the proto establishes.

(cherry picked from commit 5ac1870f6d67bf365c84b1affae63c89de32d1cb)
2026-07-30 01:43:06 +02:00
Diogo Soares Rodrigues f947ee3fb2 fix(cursor): do not execute MCP approval probes
A modern `mcpArgs` frame carrying `smart_mode_approval_only` asks only
whether a call would be permitted: the server resolves the smart-mode
permission decision ahead of the real invocation and expects the
dedicated `approved` variant back. `decodeMcpCall` dropped the flag, so
the frame ran a side-effecting MCP tool the user had not been asked
about - and ran it a second time when the real call followed.

The flag now rides on `CursorMcpCall` and short-circuits before both
execution and block synthesis: nothing ran, so a transcript entry would
claim work that never happened.

(cherry picked from commit a07dea15f5a6d376d8e70c528410f3cdd029fd30)
2026-07-30 01:43:06 +02:00
Diogo Soares Rodrigues 4a946ad8c0 fix(cursor): gate advisor tools, echo grep offset_applied
Advisor tools are built straight from the builtin table, outside the
loop that wraps every registry tool in `ExtensionToolWrapper` - which is
where the approval mode, per-tool `tools.approval.<tool>` policies and
`autoApprove` are enforced. Both the advisor's own agent loop and its
Cursor exec bridge (`pi_write`, `pi_bash`) run those instances directly,
so an advisor granted `write` or `bash` executed them regardless of a
configured `ask` or `deny`. Verified before the fix: a raw `write`
instance created the file under `tools.approval.write: deny`; the
wrapped one refuses. `bridgeToolMap` and the grep factory only ever
wrapped the two tools they build themselves.

The grep answer also now echoes `offset_applied`. Forwarding the offset
without acknowledging it leaves the server unable to distinguish a
honored page from a client that ignored the field, so it re-paginates
from the same place. Set on all three result variants (files, count,
content); absent when the frame requested no offset.

(cherry picked from commit 13c2ed565f30ff86e31c8dc9896f940faf3aa088)
2026-07-30 01:43:06 +02:00
Diogo Soares Rodrigues 4414ee4b0d fix(cursor): honor legacy read range and grep offset
Modern Cursor builds paginate the legacy `read` and `grep` frames with
fields this branch modeled in the proto but never wired.

`read` composed no range, so every page returned the whole file (or its
own truncation) and a model walking a large file never advanced past the
first window. It now goes through `piReadPath`, the same helper the Pi
frame uses, so both translate a range identically - including the
`limit: 0` case, which asks for zero lines and has no selector. The
answer reports `range_applied`, left false for an unranged read since
that is precisely the server's "this is the whole file".

`grep` dropped its `offset`. The local tool paginates by file through
`skip` and advertises exactly that unit in its own "use skip=N" advice,
so an unforwarded offset re-ran the identical search and answered page
one forever. A present `0` stays unset: it means "start at the
beginning", which is the un-skipped search.

(cherry picked from commit 4f647d86b6e4a507d57fb4d247f19db8240e51cf)
2026-07-30 01:43:06 +02:00
Diogo Soares Rodrigues d9bc1e80b6 fix(cursor): handle the resource-read refusal variant, unblock lint
`ReadMcpResourceExecResult` has a `rejected` variant carrying `reason`,
not `error`, so the pairing text's collapsed error branch did not
typecheck against the full union. Each variant is now switched
explicitly; the refusal is unreachable today (the handler answers
content or `null`) but a collapsed default would have read `undefined`
if the client ever builds one.

`bun check` type-checks the workspace projects; the union error only
surfaced through `ci:check:full`, which is what CI runs.

Also renames a loop variable that shadowed the global `escape`, which
was failing `biome check` on the full tree.

(cherry picked from commit f5dca418d91983a14492f51f2873514781a9e066)
2026-07-30 01:43:06 +02:00
Diogo Soares Rodrigues 6254b6e81b fix(cursor): build the pi_edit bridge independently of the session's provider
Every native `pi_edit` failed after a session switched onto Cursor. The
replace-mode `edit` instance the frame needs was built only for sessions
CREATED on Cursor, and the tool roster is built once, at creation - a
session that started elsewhere kept its configured-mode `edit` in the
registry, which `executeTool` resolves before its fallback, so the
frame's `old_text`/`new_text` pairs failed validation against a
`hashline` schema.

The instance is now built from the `edit` grant regardless of the
initial provider, lazily so a session that never reaches Cursor never
constructs one, and `pi_edit` asks for it through a dedicated
`getEditReplaceTool` accessor rather than relying on Cursor sessions
having deleted `edit` from the registry. A session that was never
granted `edit` is still refused.

That accessor also closes an escalation the previous wiring opened up.
The session's device resolver is handed to the bridge as `getTool` and
installed as the agent loop's `resolveFallbackTool`, which runs for ANY
call outside the advertised set - so serving `edit` from it let a
hallucinated call, or one naming a tool the session deselected after
startup, execute a replace-mode edit the model was never offered. It is
device-only again.

Regressions cover both directions at the SDK level, driving a real
unadvertised `edit` through the loop and asserting the surfaced
`Tool edit not found`: an unchanged file alone would also pass if the
fallback had resolved the tool and the edit then failed validation.

(cherry picked from commit 11a28dcf7b995a9e94913269733b3199d6f4790d)
2026-07-30 01:43:02 +02:00
usr-bin-roygbiv 17e678fa2d fix(agent): honor explicit compaction endpoint
(cherry picked from commit 3bc5c578ddd2e626eda2049a660de7bffb674b52)
2026-07-30 01:42:53 +02:00
usr-bin-roygbiv 64234e05c9 fix: align manual native compaction fallback
(cherry picked from commit 2d5397a52f6feaeee4136fe4e7d8271309043b75)
2026-07-30 01:42:53 +02:00
usr-bin-roygbiv 9dd8d3c6ce fix: preserve native compaction failures
(cherry picked from commit d13e9f30a06cad347226d2fa377cab8c086debc3)
2026-07-30 01:42:53 +02:00
usr-bin-roygbiv 7fff8869a0 fix: skip unauthenticated compaction candidates
(cherry picked from commit fa5f7d73ec1f169f9b9952648195057caa24657e)
2026-07-30 01:42:53 +02:00