- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
Matched fuzzy candidates against immutable source text while excluding ranges already selected for replacement. Inserted replacement content can no longer become a later exact or fuzzy candidate.
Added regression coverage for multiple fuzzy source matches when the replacement contains the original search text.
Fixes#7432
- Update hashline block resolution formatting to correctly incorporate anchor lines within operation labels.
- Fix and update test assertions and mock contexts across coding agent tests.
- Add the `app.live.toggle` keybinding defaulted to `Ctrl+L` to start or stop live voice mode.
- Remap the default display-reset action (`app.display.reset`) from `Ctrl+L` to `Alt+L`.
- Update the live visualizer to listen for stop keys so the toggle chord terminates active sessions.
- Replaced legacy `SWAP`, `INS`, and `PASTE` commands with unified `PUT` and `CUT` hunks across parser, grammar, tokenizer, and test suites.
- Added support for named registers and span paste operations in clipboard and block execution logic.
- Implemented indentation repair and enhanced gap locator formatting for improved patch resilience.
- Updated documentation, system prompts, and session analysis scripts to reflect the new syntax and header shapes.
- Add new structural, multi-edit, and block-level mutation classes with updated category mappings.
- Introduce hunk extraction, placement, rendering, and solver utilities along with unit tests.
- Implement size-based mutation planning, prompt validation logic, and new prompt markdown templates.
- Update benchmark generation scripts and package configurations to support empirical edit shape statistics.
- Removed copy and delete operations across tokenizer, parser, grammar, and clipboard logic.
- Standardized line-editing operations and block resolvers to use cut exclusively.
- Updated documentation, prompts, and test suites to reflect the removal of copy and delete syntax.
- Implemented clipboard register management, parsing, and execution rules for CUT, COPY, and PASTE operations in the hashline engine.
- Added session-persistent clipboard state and integration across agent session execution, diff previews, and streaming tools.
- Added comprehensive validation, error messages, recovery handling, and test coverage for clipboard and block operations.
Three defects the exec bridge shipped with, all found by review.
`pi_edit` never worked. The session removes `edit` from the tool
registry for Cursor so the model is steered to full-file `write`
(8ba0498eb), but that same registry is the bridge's tool source, so the
native frame — which the server sends regardless of the advertised
catalog — resolved nothing and answered `Tool "edit" not available`.
Retaining the instance is not enough either: `PiEditExecArgs` carries
`old_text`/`new_text` pairs, which only `replace` accepts, while the
default mode is `hashline` (`{ input: string }`). `EditTool` now takes
an optional mode, and the bridge resolves a pinned `replace` instance
through its fallback resolver.
A `pi_grep` frame carrying `context` or `limit` escaped the approval
gate. Honoring those needs a per-call tool, and the per-call instance
was built raw while every registry tool is wrapped — so exactly those
calls skipped `tools.approval.grep` and the exec-tier SSH check. Both
callsites now go through one `createBridgeGrepFactory`.
Advisors ignored the same two fields: only the primary session supplied
the factory. They now get it too, gated on the advisor actually holding
`grep` so the factory cannot grant a denied tool.
Also moves the pure Pi arg translation to `providers/cursor-pi-args`.
The legacy shim shares it and is compiled into the bundled virtual
registry, where `./providers/*` cannot match a nested specifier — it
fell through to `Bun.resolveSync`, unsatisfiable under bunfs (#3442) —
and the exec module would have dragged the protobuf graph along.
Verified against real files and the real module graph: `pi_edit` mutates
a temp file, the bundled probe executes the shim's shared module in a
subprocess, and the grep test drives the shared factory. Mutation-
checked: returning a raw tool from the factory, ignoring the pinned edit
mode, dropping the `getTool` fallback, or moving the helpers back to a
nested path each fails a test.
(cherry picked from commit e46ba22b634e449005f7c22b6d0efd19a45ce1f8)
Root cause of the reported "edit tool silently reformats the whole
file" corruption: fs/write_text_file has no verbatim guarantee. When
an ACP client (e.g. Zed with format_on_save: on) reformats a buffer
on save, routeWriteThroughBridge reported the pre-write content as
successfully written, and Patcher.commit keyed the returned snapshot
tag on that same pre-write text instead of what actually landed on
disk. The next edit anchored on that tag then resolved hunks against
a baseline the file had already drifted away from, which is what
produced whole-file "corruption" from single-line hunks -- reproduced
live in this session against real Swift/JSON/TypeScript files with
Zed as the ACP client.
- routeWriteThroughBridge reads the file back after the bridge write
and returns the verified content plus a drift flag (best-effort:
ACP defines no ordering between the client acking the write and its
own async format-on-save settling, so this degrades gracefully to
the old stale-tag-on-next-read failure mode, never to corruption).
- HashlineFilesystem.writeText propagates that verified content in
view-space (the same space readText returns -- e.g. a notebook's
editable cell text, not its raw JSON), not storage-space, so tag
validation on the next edit compares like with like.
- Patcher.commit keys fileHash/header/snapshot on the verified
post-write content (normalized, so BOM/line-ending restoration never
produces a false "drift") when it diverges from what was sent, and
appends a warning naming the drift -- but deliberately leaves the
returned `after` (and therefore the model-visible diff) scoped to
the intended hunk. Diffing against the full drifted file would
balloon the tool response to span every reformatted line (measured
~6.8x inflation on a 245-line file with one touched line); the
warning is the correct O(1) channel for "your editor reformatted
this," not an O(file-size) diff.
- write.ts keys its own snapshot header on the verified bridge content
too (no diff-size concern there since write always replaces the
whole file).
Caught via code review (dispatched against the first pass of this
fix): a naive "just use the verified content everywhere" fix broke
.ipynb editing outright (write-space vs read-space content mismatch,
tag invalid on every notebook edit) and would have inflated every
drifted edit response by ~6.8x. Both are now covered by regression
tests that fail against the pre-fix code and pass against this one.
(cherry picked from commit 35ab80e43be5800b2f48728e4400eb9fd7f7f7d2)
Passed session-scoped settings through Edit and Write generated-file checks and fell back to schema defaults when no global singleton exists.
Guarded inline image sizing against an uninitialized global settings proxy and added isolated-session regression coverage.
Fixes#6549
- Shared unique workspace suffix resolution between read and direct edit modes.
- Preserved create destinations and ambiguous-path failures while resolving existing update targets.
- Added direct replace, patch, and apply_patch regression coverage.
Fixes#6359
- Implemented native UTF-16 text processing in Rust diff module with support for unpaired surrogates.
- Removed `similar` crate from Rust workspace and `diff` npm package from coding-agent, hashline, and natives.
- Removed jsdiff fallback wrappers and `isWellFormed()` guards from TypeScript diff implementations.
- Added comprehensive test suite for native diff functions covering random inputs and edge cases including surrogates and emoji.
- Renamed model `codex-auto-review` to `gpt-5.3-codex-spark` with updated pricing and context window.
- Added the `enforceSeenLines` option to hashline `PatcherOptions` (defaults `true`); the seen-line guard in `Patcher` now runs only when enabled.
- Added the `edit.enforceSeenLines` coding-agent setting (default off) and wired it through `edit/hashline/execute.ts` into the `Patcher`.
- Stopped `file-snapshot-store` excluding column-clipped (>512-char) lines from a snapshot's seen set, so single-line edits on long lines apply without a full-width re-read.
- Updated `seen-line-guard` tests and the hashline/coding-agent changelogs.
- Bounded expanded partial edit diff rendering in `formatStreamingDiff` to `previewWindowRows()` instead of an unbounded budget, preventing runaway preview growth during live updates.
- Updated streaming diff tests to simulate terminal height and verify expanded previews stay full only within the viewport, then switch to a truncated tail with the `more lines above` marker when too tall.
- Reinitialized in-memory `Settings` before each initial-messages test since the test suite reads global display configuration and needs isolation.
- Offloaded slow LSP diagnostics to a deferred channel in the `write` tool to prevent blocking agent execution for the full 3-second poll window.
- Abstracted deferred diagnostic logic into a reusable `DeferredDiagnostics` class to standardize tracking and deduplication across tools.
- Updated `WriteTool` to support `beginDeferredDiagnosticsForPath` callbacks, surfacing diagnostics as an aside instead of stalling the tool result.
- Added a regression test validating that `write` completes immediately while diagnostics arrive asynchronously.
Two wrap artifacts in the Edit result card, both in wrapEditRendererLine:
- A row that broke inside an intra-line diff highlight ended with inverse
video still active (only the foreground was reset), so the frame's
right-edge padding painted as a default-foreground block. Every wrapped
diff row now closes inverse alongside the foreground reset; the next row
re-opens its own state, so highlights spanning the break render the same.
- The gutter matcher required a marker at column 0 immediately followed by
digits, a shape only produced when marker and number exactly fill the
gutter. Left-padded gutters (" -42│", any line number narrower than the
widest in the diff) and dedup-blanked gutters (" +│" on the added row
of a single-line replacement) fell back to generic wrapping, so their
continuation rows escaped into the line-number column. │-separated gutters
now accept padded and blank line numbers; ASCII "|" gutters still require
the canonical marker+number shape emitted by the plain fallback, so body
lines that merely start with "|", " |", or "123|" keep wrapping
generically.
Regression tests cover continuation-gutter containment, net-inverse-off at
every row end (with a precondition proving a highlight actually crossed a
break), phantom-gutter rejection for pipe- and digit-leading body lines, and
the plain-fallback canonical-row path.
Validated path-like renderer inputs before calling path helpers so provider-supplied arrays or objects cannot crash TUI rendering before schema validation reports the bad tool call.
Added renderer regression coverage for read, write, and edit call/result components with array and object path arguments.
Fixes#4525
- Replaced commit-based stability checks with a unified `isTranscriptBlockFinalized` tracking mechanism.
- Removed deprecated provisional rendering configuration and flags across tool and renderer interfaces.
- Standardized native scrollback boundary logic to pin at the first unfinalized block using settled row verification.
- Updated and refactored test suites to validate block finalization and settled row boundaries instead of deprecated commit stability methods.
- Introduced an automated retry recovery system to track, manage, and persist recovered error states within agent sessions.
- Enabled compact transcript rendering for recovered auto-retry errors by removing heuristic commit machinery.
- Improved raw read tracking and provenance in the ReadTool to support refined file snapshot recording and hashline editing.
- Excluded recovered assistant messages from default model context and updated event controllers to handle retry recovery life cycles.
Bridge-backed writes now respect tool timeout/cancel: routeWriteThroughBridge takes an optional AbortSignal and forwards it to notifyWorkspaceWatchedFiles, so a wedged LSP server no longer hangs the bridge path.
Updated write, replace, patch, and hashline callers to pass the tool signal.
Refs #4459
Announced harness-authored create, change, and delete operations to active LSP clients with workspace/didChangeWatchedFiles before edit-time diagnostics are read.
Added regression coverage for non-LSP sibling files in a batched write so diagnostics see the workspace state the harness just produced.
Fixes#4459
- Added `normalizeSingleStringField` to dynamically map misplaced string inputs to required schema fields for single-argument tools.
- Integrated argument normalization into `validateToolArguments` to handle model-specific variations in JSON payloads during validation passes.
- Updated `coding-agent` streaming and rendering components to recognize `_input` as a legacy alias for `input` across various UI paths and logic flows.
- Refactored `hashlineEditParamsSchema` to strictly enforce the `input` field while maintaining support for legacy aliases via runtime coercion rather than schema definition.
- Corrected unit tests to reflect that `_input` is rejected by the strict schema but handled gracefully by the validation layer.
- Simplified match logic to rely exclusively on content hash equality.
- Removed strict validation that rejected colliding snapshot tags.
- Updated recovery behavior to resolve collisions to the most-recently recorded snapshot.
- Refactored tests to expect successful preview and patching despite tag ambiguity.
Codex reviewer flagged that the ranged-read fallback recommended by
the patcher's over-cap reveal path (`path:N-M`) itself applies the
512-column cap and still records the displayed line numbers via
`recordSeenLinesFromBody`. On a minified wide anchor line, `read
file:N` shows only the clipped prefix but the line lands in the
tag's seenLines — a subsequent edit anchored at N then slips past
the seen-line guard.
- `packages/coding-agent/src/edit/file-snapshot-store.ts`:
`recordSeenLinesFromBody` grows an optional `excludedLines` set;
parsed line numbers matching it are filtered before recording.
- `packages/coding-agent/src/tools/read.ts`:
`#readLocalFileMultiRange` and the single-range disk path build a
`clippedLines` set alongside `columnTruncated` for both direct-range
lines and `buildLineEntriesWithBlockContext` context lines, then
pass it into `recordSeenLinesFromBody`.
- `packages/coding-agent/src/tools/grep.ts`:
same wiring for the match line via `match.truncated`, plus a
conservative length+`...`-marker heuristic for context lines
(native `crates/pi-natives/src/grep.rs` `truncate_line` doesn't
propagate a per-line flag on `contextBefore`/`contextAfter`; a
proper native-side flag is a follow-up).
- `packages/coding-agent/test/edit/seen-line-guard.test.ts`:
new case reads a 4KB single line and asserts the clipped line
number stays out of `seenLines` and the edit against it still
rejects with the seen-line guard.
- Reverted to flushing only on the last file write or explicitly on early failure paths within `apply_patch` multi-file operations.
- Refactored error counting logic within single path entries to use clean booleans instead of numeric counters.
- Replaced custom preview capping logic in task progress rendering with `capPreviewLines` and added an option to hide the expand hint.
- Introduced the `allowCreateOverwrite` option to permit `op: "create"` to replace existing files.
- Enabled `allowCreateOverwrite` specifically for the JSON-based `patch` edit mode to support full-file restructures.
- Maintained the strict non-overwriting behavior for Codex `apply_patch` envelope-based file additions.
- Configured patch diff previews to respect the configured overwrite permission during streaming.
- Fixed an issue where stopping a multi-file patch application early skipped flushing the active LSP writethrough batch.
- Guard against 16-bit snapshot tag collisions by requiring live text to be byte-identical to the retained snapshot.
- Transition base text resolution to query exact matches via `snapshots.byHashExact`.
- Prevent applying incorrect preview edits when live file contents drift to a colliding state.
The apply_patch language documents `*** Add File` and `*** Move to` as
strictly non-overwriting (create / rename), but the fs-level create and
rename paths in applyNormalizedPatch wrote through to the resolved target
without checking whether it already existed. Existing destinations were
silently replaced, and in the rename case the source was also deleted.
The multi-file executeApplyPatchPerFile aggregator caught each per-file
exception, appended an error entry, and kept iterating. Later files
still ran against an inconsistent post-state, and the aggregate result
had no top-level isError — so a mixed partial application looked like a
successful edit to the agent loop.
Changes:
- Add fs.exists guards before the create write and before the rename
write/delete in packages/coding-agent/src/edit/modes/patch.ts. Both
reject with ApplyPatchError before any side effect.
- Make executeApplyPatchPerFile in packages/coding-agent/src/edit/index.ts
stop at the first per-file failure, list applied vs. skipped files in
the aggregate text, and propagate isError, matching executeSinglePathEntries.
- Rename the two apply-patch scenario fixtures (010_move_..., 011_add_...)
that pinned the buggy overwrite behavior to _rejects_ variants, and
flip their expected/ trees so source and pre-existing destination
remain byte-identical after the rejected apply.
- Cover both failure modes with new regressions in
packages/coding-agent/test/core/apply-patch.test.ts and a new
packages/coding-agent/test/core/apply-patch-multi-file.test.ts.
Fixes#4074
Four non-overlapping algorithmic complexity reductions in hot paths.
(Streaming-reveal throughput is owned separately by #3843.)
1. session/session-manager.ts pathTo: O(n^2) branch.unshift() leaf->root
walk -> O(n) push + single reverse(). Hot path (5-10x/turn via getBranch).
2. edit/streaming.ts extractAddedLines: O(n^2) progressive string
concat per streaming tick -> array push + single join.
3. edit/modes/patch.ts: collapseConsecutiveSharedLines O(n*m) filter
+ includes -> Set (O(n+m)); collapseRepeatedBlocks O(n^3) with
per-iteration slice allocations + every() -> index arithmetic with
a single shared.has() guard. Semantics preserved.
Honorable: tui/src/utils.ts replaceTabs reallocated
" ".repeat(DEFAULT_TAB_WIDTH) every call -> hoisted TAB_SPACES const.
The hashline multi-section path in `executeHashlineSingle` returned
`perFileResults: rendered.map(r => r.perFileResult)` directly. Each
per-section result had already been individually pruned by
`renderSection`, but the whole array bypassed the shared aggregate
budget added in 3987969 — a single hashline payload touching many files
with sub-32 KB snapshots each could still serialize unbounded snapshot
bytes into one session JSONL line.
Wrap the multi-section return in `pruneOversizedEditSnapshots`, which
delegates to `capPerFileSnapshots` and enforces the shared cap walking
left-to-right; early sections keep their ACP diff visualization, later
sections in a many-file batch degrade to text-only.
End-to-end regression test seeds five real on-disk files (~10 KB
combined snapshots each), runs a multi-section hashline SWAP via
`executeHashlineSingle`, and asserts the aggregate result holds the
cumulative kept snapshot bytes under `MAX_EDIT_SNAPSHOT_TEXT_CHARS`
with at least one section pruned.
The per-entry 32 KB cap let a many-file batch (apply_patch / hashline
touching N files) accumulate unbounded snapshot bytes because each
`perFileResults` entry was checked independently. 100 files × 30 KB
each kept everything (~3 MB) even though the whole array still
serializes into one session JSONL line.
`capPerFileSnapshots` walks entries left-to-right with one shared
`MAX_EDIT_SNAPSHOT_TEXT_CHARS` budget. Each per-entry payload is still
capped individually by `pruneSnapshot`; if an entry's surviving bytes
would push the running aggregate past the cap, the entry is stripped
and stamped with `snapshotsPruned: true`. Early entries keep their ACP
diff visualization; later entries in a large batch degrade to text-only
exactly like over-sized single edits.
Regression test exercises five equal-size entries that each fit the
per-entry budget but bust it cumulatively, asserting only the first two
keep snapshots and the trailing three carry the pruned marker.
When a multi-entry single-path edit prunes the first entry's snapshots
(large pre-image) and keeps a later entry's snapshots (file shrunk
between entries), the aggregator at `executeSinglePathEntries` recorded
the later entry's small `oldText`/`newText` as the whole-file
transition. ACP clients would then render a misleading partial diff
instead of degrading to text-only for the over-budget edit.
Add an explicit `snapshotsPruned` marker on `EditToolDetails` /
`EditToolPerFileResult`, set by `pruneSnapshot` whenever it strips a
payload. `executeSinglePathEntries` tracks the flag across child
results and suppresses aggregate `oldText`/`newText` (re-stamping the
marker on the aggregate) the moment any child was pruned;
`executeApplyPatchPerFile` propagates the flag onto each per-file entry.
Regression test exercises the exact scenario the reviewer raised on
#3787: replace mode where entry 1 collapses a >1 MB file to a single
line (pruned) and entry 2 trivially renames the now-tiny result; the
aggregate result now carries `snapshotsPruned: true` with both
snapshot fields omitted.