diff --git a/docs/tools/edit.md b/docs/tools/edit.md index d9951083b..9128ce5e3 100644 --- a/docs/tools/edit.md +++ b/docs/tools/edit.md @@ -8,18 +8,13 @@ - Key collaborators: - `packages/coding-agent/src/utils/edit-mode.ts` — selects active edit mode - `packages/hashline/src/grammar.lark` — hashline grammar - - `packages/hashline/src/format.ts` — sigils and header constants (`¶`, `#`, `:`, `:-`, `+`, `^`) + - `packages/hashline/src/format.ts` — sigils and header constants (`¶`, `#`, `@@`, `+`, `&`, `,`) - `packages/hashline/src/input.ts` — parses `¶PATH#TAG` sections - `packages/hashline/src/tokenizer.ts` / `packages/hashline/src/parser.ts` — tokenizes and parses ops - `packages/hashline/src/apply.ts` — applies parsed edits to file text - `packages/hashline/src/mismatch.ts` — stale-anchor mismatch formatting - `packages/hashline/src/recovery.ts` — snapshot-based stale-anchor recovery - - `packages/hashline/src/snapshots.ts` — mints and resolves per-path two-hex opaque snapshot tags - - `packages/coding-agent/src/edit/file-snapshot-store.ts` — per-session read/search snapshot store wiring - - `packages/coding-agent/src/tools/read.ts` — emits anchored lines and records read snapshots - - `packages/coding-agent/src/tools/search.ts` — records sparse snapshots from matches/context - - `packages/coding-agent/src/tools/fs-cache-invalidation.ts` — invalidates FS scan caches after writes - - `packages/coding-agent/src/edit/streaming.ts` — computes in-flight diff previews for the TUI + - `packages/hashline/src/snapshots.ts` — mints and resolves per-path opaque snapshot tags ## Inputs @@ -27,46 +22,37 @@ | Field | Type | Required | Description | | --- | --- | --- | --- | -| `input` | `string` | Yes | One or more edit sections. Anchored sections must start with `¶PATH#TAG`; unbound `¶PATH` is allowed only for new-file / `BOF` / `EOF` boundary inserts. Optional `*** Begin Patch` / `*** End Patch` envelope is ignored if present. | +| `input` | `string` | Yes | One or more file sections. Anchored sections start with `¶PATH#TAG`; hashless `¶PATH` is allowed only for new-file creation or BOF/EOF-only inserts. Optional `*** Begin Patch` / `*** End Patch` envelope is ignored if present. | Patch language inside `input`: -- **Section header**: `¶PATH#TAG` for anchored edits, `¶PATH` for BOF/EOF-only inserts. `TAG` is two lowercase hex chars minted by the session snapshot store. -- **Anchor blocks** select a range of original lines: - - `A-B:` — select lines A..B; the body rows below describe their new content. - - `A-B:-` — select lines A..B and delete them. No body permitted. - - `A:` is accepted as `A-A:`. `A:-` is accepted as `A-A:-`. - - `BOF:` — virtual position before line 1; body rows insert there. - - `EOF:` — virtual position after the last line; body rows insert there. - - `BOF-BOF:` / `EOF-EOF:` / `BOF-EOF:` are silently normalized to the virtual anchor (range suffix carries no information for virtual positions). -- **Body rows** (one per line, immediately under the anchor): +- **File header**: `¶PATH#TAG` (or `¶PATH` for new-file / virtual-only hunks). `TAG` is three uppercase-hex chars minted by the session snapshot store. +- **Hunk header**: bare `A B` selects original lines A..B. The range separator is normally whitespace; the parser also silently accepts `A-B`, `A..B`, and `A…B` (unicode ellipsis). Virtual variants `BOF` and `EOF` target positions before line 1 / after the last line. The bare single-line shorthand `A` is accepted as `A A`. +- **Body rows** (one per line, immediately under the hunk header): - `+TEXT` — add the literal line `TEXT` verbatim, including all leading whitespace. - `+` alone — add one blank line. - - `^A-B` — re-emit original file lines A..B. Use this to keep some of the lines you selected. `^A` is accepted as `^A-A`. -- **Semantics of the body**: + - `&A..B` — re-emit original file lines A..B. Use this to keep some of the lines you selected. `&A` is accepted as `&A..A`. +- **Semantics**: - The new content of the selected range is just the body rows top-to-bottom. - - `A-B:` with no body rows REPLACES the range with one blank line. Use `A-B:-` to delete. - - `BOF:` / `EOF:` with no body inserts one blank line at that virtual position. + - **Empty body deletes the range entirely.** + - `BOF` / `EOF` with empty body is a no-op (nothing to insert). -Anchors come from `read`/`search` output. `read` emits a `¶PATH#TAG` header from the session snapshot store and lines as `LINE:TEXT`; copy the header into the edit section and copy only the line number into anchor lines. - -Other edit modes exist (`replace`, `patch`, `apply_patch`) and are selected outside the tool payload by `resolveEditMode()` in `packages/coding-agent/src/utils/edit-mode.ts`. Their schemas are different; this document covers the default hashline mode. +Anchors come from `read`/`search` output. `read` emits a `¶PATH#TAG` header from the session snapshot store and lines as `LINE:TEXT`; copy the header into the edit section and copy only the line number into hunk headers. ### Tolerated input shapes (lenient parsing) Because models reproduce nearby shapes (`read` output, `apply_patch` envelopes, unified-diff hunks), the parser is liberal about a handful of harmless variants: -- `A:` / `A:-` — single-line shorthand for `A-A:` / `A-A:-`. -- `^A` — shorthand for `^A-A`. -- Bare body rows with no `+`/`^` prefix are auto-prepended with `+` and a `BARE_BODY_AUTO_PIPED_WARNING` is appended, BUT only when every row in that block is uniformly bare. Mixed `+`/raw blocks still throw. -- Lone `-` body row immediately after a bare anchor is retroactively converted to a `:-` delete with a `DASH_PAYLOAD_AUTO_DELETE_WARNING`. -- An overlapping bare anchor followed by a concrete delete or replace block is treated as a stale "before then after" pair: the bare block is dropped with a `REPLACE_PAIR_COALESCED_OVERLAP_WARNING`. Identical-range pairs use the same coalesce as a stronger guarantee with `REPLACE_PAIR_COALESCED_WARNING`. -- Two or more consecutive single-line `A-A:` blocks with empty bodies emit a `STACKED_BLANK_REPLACE_WARNING` (the model probably meant `A-B:-`). -- `*`/`>` decoration prefixes from grep-style output are stripped from anchors. -- `*** Update File:` / `*** Add File:` / `*** Delete File:` sentinels and unified-diff `@@` headers throw an `apply_patch sentinel … is not valid in hashline` error so the model knows it shipped the wrong format envelope. -- `-N:` / `-N-M:` apply_patch hunk-anchor prefixes throw an `apply_patch line prefix … is not valid in hashline` error. -- A lone `-` outside any pending block throws a focused `a lone "-" is not a valid hashline op` error pointing at `A-B:-`. -- `*** Begin Patch` / `*** End Patch` envelopes are silently consumed. `*** Abort` terminates parsing silently — ops parsed before the marker still apply, no warning is surfaced. +- `A` — accepted as `A A` (single-line shorthand). +- `A-B`, `A..B`, `A…B` — accepted as `A B` (any of hyphen, double-dot, or unicode ellipsis works as a silent separator). +- `&A` — accepted as `&A..A`. +- Bare body rows with no `+`/`&` prefix are auto-prepended with `+` and a `BARE_BODY_AUTO_PIPED_WARNING` is appended, BUT only when every row in that block is uniformly bare. Mixed `+`/raw blocks still throw. +- `+&A..B` rows (model mistakenly prefixed a repeat with `+`) are silently rerouted as `&A..B` repeats with `PLUS_PREFIXED_REPEAT_WARNING`. +- Identical-range hunks in the same patch are coalesced last-wins with `REPLACE_PAIR_COALESCED_WARNING`. +- An overlapping bare hunk followed by a concrete hunk is treated as a stale "before then after" pair; the bare hunk is dropped with `REPLACE_PAIR_COALESCED_OVERLAP_WARNING`. +- `*** Begin Patch` / `*** End Patch` envelopes are silently consumed. `*** Abort` terminates parsing silently — ops parsed before the marker still apply, no warning surfaced. +- `*** Update File:` / `*** Add File:` / `*** Delete File:` / `*** Move to:` apply_patch sentinels throw an `apply_patch sentinel … is not valid in hashline` error. +- `@@`-bracketed hunk headers (whether the apply_patch `@@ context @@` form or the unified-diff `@@ -N,M +N,M @@` shape) are rejected with an explicit "drop the `@@ ... @@` brackets" message — hashline hunks are bare `A B` lines. ## Outputs - Single-shot tool result; hashline mode does not use a `resolve` preview/apply handshake. @@ -88,46 +74,13 @@ Warnings: - `meta`: output metadata - `perFileResults`: present for multi-section input - Multi-section input returns one aggregated result with combined text and per-file details. -- While the model is still typing arguments, the TUI can compute a diff preview with `packages/coding-agent/src/edit/streaming.ts`; that preview is not a deferred action and does not block execution. - -## Flow -1. `EditTool.execute()` in `packages/coding-agent/src/edit/index.ts` resolves the active mode. Default is `hashline`; `customFormat` exposes `packages/hashline/src/grammar.lark` as a constant string for prompt embedding. -2. `executeHashlineSingle()` in `packages/coding-agent/src/edit/hashline/execute.ts` parses the raw `input` via `Patch.parse()` (`packages/hashline/src/input.ts`), which: - - strips a leading BOM and `*** Begin Patch` markers, - - splits the input into `¶PATH#TAG` sections, - - merges multiple sections targeting the same path so every op refers to the original file snapshot, - - rejects malformed headers. -3. For each section, `Patcher.prepare()` (`packages/hashline/src/patcher.ts`): - - parses the diff body via `parsePatch()` (tokenizer + parser), - - reads the current file, - - resolves the section tag against the session snapshot store, - - runs recovery if the tag is stale (recorded snapshot replay + 3-way merge against current disk), - - validates anchor line bounds against the resolved file content, - - applies the edits in memory via `applyEdits()`. -4. Multi-section calls preflight every section before any write hits the filesystem so a partial batch never lands. -5. `applyEdits()` in `packages/hashline/src/apply.ts`: - - expands `^A-B` repeat edits into concrete inserts, - - runs `absorbReplacementBoundaryDuplicates()` to widen replacement deletes when the payload's leading/trailing rows match adjacent file lines (with an `Auto-absorbed …` warning), - - emits a per-line `Deleted line N contains a structural bracket/brace boundary …` warning ONLY when the block's net brace/paren/bracket balance is not preserved by its replacement payload (so well-formed multi-line replaces no longer false-positive), - - applies anchor-targeted edits bottom-up so later splices do not invalidate earlier line numbers, - - applies BOF and EOF inserts after the per-line bucket. -6. `Patcher.commit()` writes the result. The writethrough callback from `createLspWritethrough()` may format the file and fetch diagnostics. -7. `invalidateFsScanAfterWrite()` calls native `invalidateFsScanCache(path)` so filesystem-backed tools do not serve stale scan results. -8. The session file-read cache is refreshed with the post-edit file text via `recordContiguous()`, making the just-written content the new recovery base for subsequent stale-anchor merges. -9. The final response is built from a unified diff (`generateDiffString()`), a compact preview, and any accumulated warnings. - -## Modes / Variants -- `hashline` — default mode; line-anchored patch language described here (`packages/coding-agent/src/utils/edit-mode.ts`). -- `replace` — exact/fuzzy old/new text replacement (`packages/coding-agent/src/edit/modes/replace.ts`). -- `patch` — structured JSON diff-hunk mode (`packages/coding-agent/src/edit/modes/patch.ts`). -- `apply_patch` — freeform Codex-style `*** Begin Patch` envelope, internally expanded into patch-mode entries (`packages/coding-agent/src/edit/modes/apply-patch.ts`). ## Worked examples Reference file (the exact shape `read` returns): ```text -¶a.ts#0a +¶a.ts#0A3 1:const X = "a"; 2:const Y = X; 3: @@ -139,8 +92,8 @@ Reference file (the exact shape `read` returns): Replace line 1 with two lines: ```text -¶a.ts#0a -1-1: +¶a.ts#0A3 +1 +const X = "b"; +export const Y = X; ``` @@ -148,106 +101,71 @@ Replace line 1 with two lines: Insert BELOW line 5 (keep line 5, add after): ```text -¶a.ts#0a -5-5: -^5-5 +¶a.ts#0A3 +5 +&5 +console.log(X + Y); ``` Insert ABOVE line 5 (add before, keep line 5): ```text -¶a.ts#0a -5-5: +¶a.ts#0A3 +5 +console.log(X + Y); -^5-5 +&5 ``` Delete lines 4..5 entirely: ```text -¶a.ts#0a -4-5:- -``` - -Replace lines 4..5 with one blank line (NOT a delete): - -```text -¶a.ts#0a -4-5: +¶a.ts#0A3 +4 5 ``` Insert at start and end of file: ```text -¶a.ts#0a -BOF: +¶a.ts#0A3 +BOF +// header -EOF: +EOF +// trailer ``` Multi-file: ```text -¶src/a.ts#0a -4-4: +¶src/a.ts#0A3 +4 +const enabled = true; -¶src/b.ts#1f -20-20:- +¶src/b.ts#1F7 +20 ``` -## Side Effects -- Filesystem - - Reads target files with `readEditFileText()`. - - Writes full updated file contents with `serializeEditFileText()`. - - Preserves BOM and original line-ending style. -- Subprocesses / native bindings - - `createLspWritethrough()` may trigger formatter / diagnostics work through the LSP subsystem. - - `invalidateFsScanAfterWrite()` calls native `invalidateFsScanCache()` from `@oh-my-pi/pi-natives`. -- Session state - - Reads and updates the per-session `FileReadCache` used for stale-anchor recovery. - - Stores pending deferred-diagnostics abort controllers per path inside `EditTool`. - - Queues late diagnostics back into the session transcript as a hidden custom message. -- Background work / cancellation - - A new edit to the same path aborts the prior deferred diagnostics fetch for that path (`packages/coding-agent/src/edit/index.ts`). - - The tool itself is marked `nonAbortable = true` and `concurrency = "exclusive"` in `packages/coding-agent/src/edit/index.ts`. - ## Limits & Caps -- Default mode is `hashline` (`DEFAULT_EDIT_MODE`) in `packages/coding-agent/src/utils/edit-mode.ts`. -- File snapshot tags are exactly two lowercase hex chars minted by the per-session snapshot store. -- Each path gets a 256-slot ring. The initial slot is random, and each store randomizes slot→tag encoding, so tags are opaque rather than predictable counters. +- File snapshot tags are exactly three uppercase-hex chars minted by the per-session snapshot store. - The visible mismatch report shows 2 lines of context on each side (`MISMATCH_CONTEXT`) in `packages/hashline/src/messages.ts`. - Stale-anchor recovery uses `fuzzFactor: 0` in `packages/hashline/src/recovery.ts`. -- `HL_OP_REPLACE` is `:`, `HL_OP_DELETE_SUFFIX` is `:-`, `HL_PAYLOAD_REPLACE` is `+`, `HL_PAYLOAD_REPEAT` is `^`, `HL_FILE_PREFIX` is `¶`, and `HL_FILE_HASH_SEP` is `#` (`packages/hashline/src/format.ts`). +- `HL_FILE_PREFIX` is `¶`, `HL_PAYLOAD_REPLACE` is `+`, `HL_PAYLOAD_REPEAT` is `&`, `HL_RANGE_SEP` is `..` (repeat-row bodies only), and `HL_FILE_HASH_SEP` is `#` (`packages/hashline/src/format.ts`). Hunk headers carry no sigil; the range is just two whitespace-separated line numbers. ## Errors - Missing section header: - `input must begin with "¶PATH#HASH" on the first non-blank line for anchored edits; got: ...` -- Empty header: - - `Input header "¶" is empty; provide a file path.` - Missing tag for anchored edit: - `Missing hashline snapshot tag for anchored edit to ; use ¶#tag from your latest read/search output.` -- Inline payload on the anchor line: - - `line N: Inline payload on the anchor line is rejected. Write the anchor on its own line (e.g. A-B:), then put the body content on the next line prefixed with + (literal) or ^A-B (repeat). …` - Stray payload line: - - `line N: payload line has no preceding A-B:, BOF:, or EOF: anchor. Got "...".` -- Raw body row with no `+` / `^` prefix in a mixed-prefix block: - - `line N: payload row in a hashline block must start with + or ^A-B. Got "...".` + - `line N: payload line has no preceding hunk header. Use an \`A B\` (or \`BOF\` / \`EOF\`) line above the body. Got "...".` +- Raw body row with no `+` / `&` prefix in a mixed-prefix block: + - `line N: payload row in a hashline hunk must start with + or &A..B. Got "...".` - Range out of order: - - `line N: range A-B ends before it starts.` -- Overlapping ops on the same anchor: - - `line N: anchor line X is already targeted by another op on line Y. Issue ONE block per range; payload is only the final desired content, never a before/after pair.` -- BOF/EOF used with `:-`: - - `line N: BOF:/EOF: anchors are virtual positions and cannot use :-. Use +TEXT or ^A-B body rows to insert at a virtual position.` -- Lone `-` op at top level: - - `line N: a lone "-" is not a valid hashline op. To delete a range, write A-B:- on the anchor line itself (e.g. 5-7:-).` + - `line N: range A..B ends before it starts.` +- Overlapping hunks on the same anchor: + - `line N: anchor line X is already targeted by another hunk on line Y. Issue ONE hunk per range; payload is only the final desired content, never a before/after pair.` - apply_patch / unified-diff contamination: - - `line N: apply_patch sentinel "*** …" is not valid in hashline. Use ¶PATH#HASH then A-B: / A-B:- / BOF: / EOF: blocks …` - - `line N: unified-diff hunk header (@@) is not valid in hashline. Use a ¶PATH#HASH header and bare A-B: anchor blocks.` - - `line N: apply_patch line prefix (-N: / -N-M:) is not valid in hashline. Drop the - prefix; use A-B: (replace) or A-B:- (delete) on the anchor line itself.` -- Missing file for anchor-scoped edits: - - `File not found: ` + - `line N: apply_patch sentinel "*** …" is not valid in hashline. File sections start with \`¶path#HASH\` (no \`Update File:\` / \`Add File:\` keyword). Hunks are bare \`A B\` lines with \`+TEXT\` / \`&A..B\` body rows.` + - `line N: unified-diff hunk header (\`@@ -N,M +N,M @@\`) is not valid in hashline. Hashline hunks are bare \`A B\` lines (or \`BOF\` / \`EOF\` keywords).` + - `line N: \`@@\`-bracketed hunk header "@@ …" is not valid in hashline. Drop the \`@@ ... @@\` brackets and write the range directly: \`5 7\` (or \`5\` for a single line, \`BOF\` / \`EOF\` for virtual positions).` - Out-of-range anchor: - `Line N does not exist (file has M lines)` - Stale snapshot tag throws `MismatchError`. The error contains re-read guidance and nearby current file lines as `*LINE:TEXT` / ` LINE:TEXT`. @@ -256,25 +174,8 @@ Multi-file: - Recovery failure is silent internally: if cache-based merge cannot prove a valid result, the mismatch error is surfaced unchanged. ## Warnings -- `Detected two identical-range hashline blocks; kept only the second block. …` (`REPLACE_PAIR_COALESCED_WARNING`) -- `Detected an overlapping bare hashline block immediately followed by a concrete block; dropped the earlier bare block. …` (`REPLACE_PAIR_COALESCED_OVERLAP_WARNING`) -- `Auto-prefixed bare body row(s) with +. Always start payload rows with +TEXT (literal) or ^A-B (repeat) …` (`BARE_BODY_AUTO_PIPED_WARNING`) -- `Converted a lone - body row to a :- delete on the preceding anchor. Write A-B:- on the anchor line itself to delete the range.` (`DASH_PAYLOAD_AUTO_DELETE_WARNING`) -- `Detected a run of single-line empty-body blocks (A-A: with no payload). Each one REPLACES its line with a blank; to delete lines use A-B:-.` (`STACKED_BLANK_REPLACE_WARNING`) -- `Auto-absorbed N duplicate line(s) above replacement (file lines A..B matched the payload's leading lines; widened the deletion to start at file line A instead of C).` -- `Auto-absorbed N duplicate line(s) below replacement …` (symmetric variant) -- `Deleted line N contains a structural bracket/brace boundary ("…"); verify the file is still balanced or use '+replacement' payload to keep the boundary intact.` — only fires when the block's net delimiter balance is not preserved by its replacement. +- `Detected two identical-range hashline hunks; kept only the second hunk. …` (`REPLACE_PAIR_COALESCED_WARNING`) +- `Detected an overlapping bare hashline hunk immediately followed by a concrete hunk; dropped the earlier bare hunk. …` (`REPLACE_PAIR_COALESCED_OVERLAP_WARNING`) +- `Auto-prefixed bare body row(s) with +. Always start payload rows with +TEXT (literal) or &A..B (repeat) …` (`BARE_BODY_AUTO_PIPED_WARNING`) +- `A body row started with `+&A..B`. `+` (literal text) and `&A..B` (repeat) are sibling row kinds …` (`PLUS_PREFIXED_REPEAT_WARNING`) - Recovery banners: `RECOVERY_EXTERNAL_WARNING`, `RECOVERY_SESSION_CHAIN_WARNING`, `RECOVERY_SESSION_REPLAY_WARNING` (`packages/hashline/src/messages.ts`). - -## Notes -- `read` and `search` are the authoritative source of section tags. Copy `¶PATH#TAG`; anchor lines use bare line numbers and do not carry the trailing `:TEXT`. -- Multi-op patches are parsed against the original file snapshot. Do not renumber later anchors after earlier ops; `applyEdits()` buckets and applies them bottom-up. -- Failed hand-edits often come from sequentially shifting later anchors inside the same patch. Treat every op as using the line numbers from the original section header. -- Inline payload on the anchor line is rejected. Put the body content on the next line prefixed with `+` (literal) or `^A-B` (repeat). -- Trailing whitespace on body rows is preserved exactly. To preserve trailing spaces, put them in the `+TEXT` row. -- Section tags are opaque snapshot-store slots, not content hashes. A tag is valid only in the session store that minted it; if the live file no longer matches the recorded snapshot, stale-anchor recovery must prove a safe merge before writing. -- `splitRawSections()` (in `packages/hashline/src/input.ts`) normalizes absolute `¶PATH#TAG` headers back to a cwd-relative path when the file is inside the current working tree. Headers with any run of leading `¶` chars (e.g. `¶foo.ts`, `¶¶foo.ts`) are accepted; the canonical form is `¶PATH#TAG` for anchored edits. -- Optional `*** Begin Patch` / `*** End Patch` markers are accepted, but the file sections are still `¶PATH#TAG`-based, not Codex `*** Update File:` hunks. -- `*** Abort` terminates parsing silently; ops parsed before the marker still apply, but no warning is surfaced. -- Snapshot tags are not invalidated on write-through; a tag remains in its path ring until that slot wraps. If a later read records different content, it mints a new tag while old snapshots remain available for recovery until overwritten. -- There is no resolve-style apply/discard phase for hashline edits. The only preview path is the transient TUI diff preview in `packages/coding-agent/src/edit/streaming.ts`. diff --git a/packages/coding-agent/CHANGELOG.md b/packages/coding-agent/CHANGELOG.md index 25bcb3589..79cf2dc14 100644 --- a/packages/coding-agent/CHANGELOG.md +++ b/packages/coding-agent/CHANGELOG.md @@ -1,16 +1,18 @@ # Changelog ## [Unreleased] -### Fixed +### Breaking Changes -- Fixed `omp auth-broker serve` crashing at startup with `logger.setTransports is not a function` — switched the call site to `import { setTransports } from "@oh-my-pi/pi-utils/logger"`, bypassing the `logger` namespace re-export that some Bun versions failed to expose at runtime +- Changed hashline edit parsing to require wrapped hunk headers such as `@@ A..B @@` (including `@@ BOF @@` and `@@ EOF @@`), with empty `@@ A..B @@` blocks deleting the anchored range and legacy inline payload forms treated as malformed ### Added - Added strict-mode indicators to `omp auth-gateway check` output by appending `[strict]` to strict-mode text headers and adding a top-level `strict` field in `--json` output - `omp auth-gateway check --strict` exercises each broker-supplied credential against its provider's chat-completion endpoint (cheapest bundled chat model per provider, with 15s/attempt timeout and up to 4 catalog fall-throughs on "model not found / invalid model" errors). Surfaces failures where the usage endpoint reports 200 but the chat endpoint 401s the same bearer (revoked OAuth scope, mislabeled provider row, …). Output gains a `[chat: ok|FAIL|skip]` column in text mode and a `completion` field on each credential in `--json` mode; the chat-failed count contributes to the non-zero exit code. + ### Changed +- Changed hashline apply behavior to preserve duplicated boundary and context lines in replacement and insert payloads instead of auto-absorbing or dropping them - Updated hashline syntax: replaced `↑`/`↓` payload sigils with `^` repeat syntax and `|` literal rows for clearer edit semantics - Changed hashline delete syntax from bare `A:` or `A-B:` to explicit `A-B:-` inline delete marker - Modified hashline anchor syntax to require explicit range notation `A-B:` instead of shorthand `A:` for single-line operations @@ -32,6 +34,19 @@ - Fixed `extractRetryHint` not recognising Codex's `Try again in ~N min.` / `… hour` / `… hours` phrasing, which left the gateway and TUI without a server-suggested retry window when an upstream account hit its usage cap. The shared `try again in` pattern now accepts `min`, `minutes`, `mins`, `h`, `hr`, `hour`, `hours` units in addition to `ms` / `s` / `sec`, and tolerates a leading `~` and embedded whitespace. - Fixed the auth-gateway threading `sessionId: undefined` into `AuthStorage.getApiKey`, which left `#sessionLastCredential` empty and made `markUsageLimitReached` a no-op for gateway-mediated requests. Both `/v1/chat/completions`-style endpoints and the `/v1/pi/stream` fast path now derive a stable `sessionId` from the client's `prompt_cache_key` (or the existing model+system+tools+first-message hash when absent) and reuse the same identity for credential-stickiness and prefix-cache routing. +### Removed + +- Removed the `edit.hashlineAutoDropPureInsertDuplicates` setting +- Removed the `edit.hashlineAutoDropPureInsertDuplicates` setting from configuration and execution paths + +### Fixed + +- Fixed `eval` tool to resize large displayed images and append dimension notes to text output +- Fixed `write` tool to strip malformed or loose hashline section headers before writing file content +- Fixed `eval` tool image rendering to resize displayed images before returning them and append image-dimension notes to text output +- Fixed `write` tool output sanitation to strip malformed or loose hashline section headers before writing file content +- Fixed `omp auth-broker serve` crashing at startup with `logger.setTransports is not a function` — switched the call site to `import { setTransports } from "@oh-my-pi/pi-utils/logger"`, bypassing the `logger` namespace re-export that some Bun versions failed to expose at runtime + ## [15.5.7] - 2026-05-27 ### Added - `providers.openrouterVariant` setting (Settings → Providers → "OpenRouter Routing") to default OpenRouter requests to a routing-variant suffix (`:nitro`, `:floor`, `:online`, `:exacto`). Selectors that already name a variant (e.g. `openrouter/anthropic/claude-haiku:nitro`) keep precedence. diff --git a/packages/coding-agent/src/config/settings-schema.ts b/packages/coding-agent/src/config/settings-schema.ts index 4c971dd16..7f663a280 100644 --- a/packages/coding-agent/src/config/settings-schema.ts +++ b/packages/coding-agent/src/config/settings-schema.ts @@ -1568,16 +1568,6 @@ export const SETTINGS_SCHEMA = { }, }, - "edit.hashlineAutoDropPureInsertDuplicates": { - type: "boolean", - default: false, - ui: { - tab: "editing", - label: "Hashline Duplicate Insert Drop", - description: - "Drop payload lines that duplicate adjacent file context — 2+-line context echoes on pure inserts, and a single boundary line at either edge of an `A-B:` replacement", - }, - }, "edit.blockAutoGenerated": { type: "boolean", default: true, diff --git a/packages/coding-agent/src/edit/hashline/diff.ts b/packages/coding-agent/src/edit/hashline/diff.ts index 233c31559..5c7359035 100644 --- a/packages/coding-agent/src/edit/hashline/diff.ts +++ b/packages/coding-agent/src/edit/hashline/diff.ts @@ -23,7 +23,6 @@ import { generateDiffString } from "../diff"; import { readEditFileText } from "../read-file"; export interface HashlineDiffOptions { - autoDropPureInsertDuplicates?: boolean; /** * Use the streaming-tolerant applier ({@link PatchSection.applyPartialTo}) * so trailing in-flight ops do not throw or emit phantom edits. Streaming @@ -82,9 +81,7 @@ export async function computeHashlineSectionDiff( const normalized = normalizeToLF(content); const hashError = validateSectionHash(section, absolutePath, normalized, snapshots); if (hashError) return { error: hashError }; - const result = options.streaming - ? section.applyPartialTo(normalized, options) - : section.applyTo(normalized, options); + const result = options.streaming ? section.applyPartialTo(normalized) : section.applyTo(normalized); if (normalized === result.text) return { error: `No changes would be made to ${section.path}.` }; return generateDiffString(normalized, result.text); } catch (err) { diff --git a/packages/coding-agent/src/edit/hashline/execute.ts b/packages/coding-agent/src/edit/hashline/execute.ts index ca50038b8..d7347bbd2 100644 --- a/packages/coding-agent/src/edit/hashline/execute.ts +++ b/packages/coding-agent/src/edit/hashline/execute.ts @@ -37,12 +37,6 @@ export interface ExecuteHashlineSingleOptions { beginDeferredDiagnosticsForPath: (path: string) => WritethroughDeferredHandle; } -function getHashlineApplyOptions(session: ToolSession): { autoDropPureInsertDuplicates: boolean } { - return { - autoDropPureInsertDuplicates: session.settings.get("edit.hashlineAutoDropPureInsertDuplicates"), - }; -} - function noChangeDiagnostic(path: string): string { // The patch parsed and applied cleanly but produced no change — the // `|literal` body rows matched the file content at the targeted lines @@ -139,8 +133,7 @@ export async function executeHashlineSingle( batchRequest: options.batchRequest, }); const snapshots = getFileSnapshotStore(options.session); - const applyOptions = getHashlineApplyOptions(options.session); - const patcher = new Patcher({ fs, snapshots, applyOptions }); + const patcher = new Patcher({ fs, snapshots }); // Single-section fast path: prepare, commit, render. if (patch.sections.length === 1) { diff --git a/packages/coding-agent/src/edit/streaming.ts b/packages/coding-agent/src/edit/streaming.ts index 5aac5d94e..50a75fa6d 100644 --- a/packages/coding-agent/src/edit/streaming.ts +++ b/packages/coding-agent/src/edit/streaming.ts @@ -43,7 +43,6 @@ export interface StreamingDiffContext { snapshots: SnapshotStore; fuzzyThreshold?: number; allowFuzzy?: boolean; - hashlineAutoDropPureInsertDuplicates?: boolean; /** * True while the tool's arguments are still streaming in. Strategies that * accept free-form text input (apply_patch, hashline) trim the trailing @@ -327,9 +326,7 @@ const hashlineStrategy: EditStreamingStrategy = { // to parse; suppress until the next chunk arrives. Once args are // complete, surface the error so the model sees what went wrong. if (ctx.isStreaming) return null; - const result = await computeHashlineDiff({ input }, ctx.cwd, ctx.snapshots, { - autoDropPureInsertDuplicates: ctx.hashlineAutoDropPureInsertDuplicates, - }); + const result = await computeHashlineDiff({ input }, ctx.cwd, ctx.snapshots); ctx.signal.throwIfAborted(); return [toPerFilePreview("", result)]; } @@ -349,7 +346,6 @@ const hashlineStrategy: EditStreamingStrategy = { ctx.signal.throwIfAborted(); const section = sectionsToProcess[i]; const result = await computeHashlineSectionDiff(section, ctx.cwd, ctx.snapshots, { - autoDropPureInsertDuplicates: ctx.hashlineAutoDropPureInsertDuplicates, streaming: ctx.isStreaming, }); ctx.signal.throwIfAborted(); diff --git a/packages/coding-agent/src/modes/components/tool-execution.ts b/packages/coding-agent/src/modes/components/tool-execution.ts index da8b73547..f7a3d1cd1 100644 --- a/packages/coding-agent/src/modes/components/tool-execution.ts +++ b/packages/coding-agent/src/modes/components/tool-execution.ts @@ -110,7 +110,6 @@ export interface ToolExecutionOptions { showImages?: boolean; // default: true (only used if terminal supports images) editFuzzyThreshold?: number; editAllowFuzzy?: boolean; - hashlineAutoDropPureInsertDuplicates?: boolean; } export interface ToolExecutionHandle { @@ -144,7 +143,6 @@ export class ToolExecutionComponent extends Container { #showImages: boolean; #editFuzzyThreshold: number | undefined; #editAllowFuzzy: boolean | undefined; - #hashlineAutoDropPureInsertDuplicates: boolean | undefined; #snapshots?: SnapshotStore; #isPartial = true; #tool?: AgentTool; @@ -192,7 +190,6 @@ export class ToolExecutionComponent extends Container { this.#showImages = options.showImages ?? true; this.#editFuzzyThreshold = options.editFuzzyThreshold; this.#editAllowFuzzy = options.editAllowFuzzy; - this.#hashlineAutoDropPureInsertDuplicates = options.hashlineAutoDropPureInsertDuplicates; this.#snapshots = options.snapshots; this.#tool = tool; this.#ui = ui; @@ -277,7 +274,6 @@ export class ToolExecutionComponent extends Container { snapshots: this.#snapshots!, fuzzyThreshold: this.#editFuzzyThreshold, allowFuzzy: this.#editAllowFuzzy, - hashlineAutoDropPureInsertDuplicates: this.#hashlineAutoDropPureInsertDuplicates, isStreaming, }); if (controller.signal.aborted) return; diff --git a/packages/coding-agent/src/modes/controllers/event-controller.ts b/packages/coding-agent/src/modes/controllers/event-controller.ts index 952fe41c5..c0cdfd7c5 100644 --- a/packages/coding-agent/src/modes/controllers/event-controller.ts +++ b/packages/coding-agent/src/modes/controllers/event-controller.ts @@ -334,7 +334,6 @@ export class EventController { showImages: settings.get("terminal.showImages"), editFuzzyThreshold: settings.get("edit.fuzzyThreshold"), editAllowFuzzy: settings.get("edit.fuzzyMatch"), - hashlineAutoDropPureInsertDuplicates: settings.get("edit.hashlineAutoDropPureInsertDuplicates"), }, tool, this.ctx.ui, @@ -450,7 +449,6 @@ export class EventController { showImages: settings.get("terminal.showImages"), editFuzzyThreshold: settings.get("edit.fuzzyThreshold"), editAllowFuzzy: settings.get("edit.fuzzyMatch"), - hashlineAutoDropPureInsertDuplicates: settings.get("edit.hashlineAutoDropPureInsertDuplicates"), }, tool, this.ctx.ui, diff --git a/packages/coding-agent/src/modes/utils/ui-helpers.ts b/packages/coding-agent/src/modes/utils/ui-helpers.ts index 3508d2f67..7dd7b47d7 100644 --- a/packages/coding-agent/src/modes/utils/ui-helpers.ts +++ b/packages/coding-agent/src/modes/utils/ui-helpers.ts @@ -382,7 +382,6 @@ export class UiHelpers { showImages: settings.get("terminal.showImages"), editFuzzyThreshold: settings.get("edit.fuzzyThreshold"), editAllowFuzzy: settings.get("edit.fuzzyMatch"), - hashlineAutoDropPureInsertDuplicates: settings.get("edit.hashlineAutoDropPureInsertDuplicates"), }, tool, this.ctx.ui, diff --git a/packages/coding-agent/src/tools/eval.ts b/packages/coding-agent/src/tools/eval.ts index 28cf165ab..1422e9d4a 100644 --- a/packages/coding-agent/src/tools/eval.ts +++ b/packages/coding-agent/src/tools/eval.ts @@ -14,6 +14,7 @@ import { getMarkdownTheme, type Theme } from "../modes/theme/theme"; import evalDescription from "../prompts/tools/eval.md" with { type: "text" }; import { DEFAULT_MAX_BYTES, OutputSink, type OutputSummary, TailBuffer } from "../session/streaming-output"; import { getTreeBranch, getTreeContinuePrefix, renderCodeCell } from "../tui"; +import { formatDimensionNote, resizeImage } from "../utils/image-resize"; import { resolveEvalBackends, type ToolSession } from "."; import { truncateForPrompt } from "./approval"; import { @@ -403,6 +404,7 @@ export class EvalTool implements AgentTool { const cellStatusEvents: EvalStatusEvent[] = []; const cellDisplayOutputs: EvalDisplayOutput[] = []; + const cellImageNotes: string[] = []; let cellHasMarkdown = false; for (const output of result.displayOutputs) { if (output.type === "json") { @@ -410,8 +412,26 @@ export class EvalTool implements AgentTool { cellDisplayOutputs.push(output); } if (output.type === "image") { - images.push({ type: "image", data: output.data, mimeType: output.mimeType }); - cellDisplayOutputs.push(output); + const resized = await resizeImage({ + type: "image", + data: output.data, + mimeType: output.mimeType, + }); + const image: ImageContent = { + type: "image", + data: resized.data, + mimeType: resized.mimeType, + }; + images.push(image); + cellDisplayOutputs.push({ + type: "image", + data: image.data, + mimeType: image.mimeType, + }); + const dimensionNote = formatDimensionNote(resized); + if (dimensionNote) { + cellImageNotes.push(`display image ${cellImageNotes.length + 1}: ${dimensionNote}`); + } } if (output.type === "status") { statusEvents.push(output.event); @@ -423,9 +443,14 @@ export class EvalTool implements AgentTool { } const stdoutTrimmed = result.output.trim(); + const imageText = cellImageNotes.join("\n"); const displayText = formatDisplayOutputsForText(cellDisplayOutputs); + const visibleDisplayText = + displayText && imageText ? `${displayText}\n\n${imageText}` : displayText || imageText; const cellOutput = - stdoutTrimmed && displayText ? `${stdoutTrimmed}\n\n${displayText}` : stdoutTrimmed || displayText; + stdoutTrimmed && visibleDisplayText + ? `${stdoutTrimmed}\n\n${visibleDisplayText}` + : stdoutTrimmed || visibleDisplayText; cellResult.output = cellOutput; cellResult.exitCode = result.exitCode; cellResult.durationMs = durationMs; diff --git a/packages/coding-agent/src/tools/write.ts b/packages/coding-agent/src/tools/write.ts index 3df67ea04..3f18bfec3 100644 --- a/packages/coding-agent/src/tools/write.ts +++ b/packages/coding-agent/src/tools/write.ts @@ -1,12 +1,14 @@ import { Database } from "bun:sqlite"; import * as fs from "node:fs/promises"; import * as path from "node:path"; + import { stripHashlinePrefixes } from "@oh-my-pi/hashline"; import type { AgentTool, AgentToolContext, AgentToolResult, AgentToolUpdateCallback } from "@oh-my-pi/pi-agent-core"; import type { Component } from "@oh-my-pi/pi-tui"; import { Text } from "@oh-my-pi/pi-tui"; import { isEnoent, isRecord, prompt, untilAborted } from "@oh-my-pi/pi-utils"; import * as z from "zod/v4"; + import type { RenderResultOptions } from "../extensibility/custom-tools/types"; import { InternalUrlRouter } from "../internal-urls"; import { parseInternalUrl } from "../internal-urls/parse"; @@ -53,6 +55,8 @@ import { import { ToolError } from "./tool-errors"; import { toolResult } from "./tool-result"; +const LOOSE_HASHLINE_HEADER_RE = /^\s*¶\S+#[^ \t\r\n]*\s*$/; + let fflateModulePromise: Promise | undefined; async function loadFflate(): Promise { if (!fflateModulePromise) fflateModulePromise = import("fflate"); @@ -74,6 +78,31 @@ export interface WriteToolDetails { madeExecutable?: boolean; } +/** + * Strip hashline display prefixes from write content. + * + * Includes a fallback for loosely-formed section headers that still carry + * line-number prefixes (for example legacy or malformed hashline echoes). + */ +function stripWriteContentWithPotentialLooseHeader(lines: string[]): { text: string; stripped: boolean } { + const cleaned = stripHashlinePrefixes(lines); + if (cleaned !== lines) { + return { text: cleaned.join("\n"), stripped: true }; + } + + const headerIndex = lines.findIndex(line => line.trim().length > 0); + if (headerIndex === -1 || !LOOSE_HASHLINE_HEADER_RE.test(lines[headerIndex])) { + return { text: lines.join("\n"), stripped: false }; + } + + const linesWithoutHeader = lines.slice(0, headerIndex).concat(lines.slice(headerIndex + 1)); + const cleanedWithoutHeader = stripHashlinePrefixes(linesWithoutHeader); + if (cleanedWithoutHeader === linesWithoutHeader) { + return { text: lines.join("\n"), stripped: false }; + } + return { text: cleanedWithoutHeader.join("\n"), stripped: true }; +} + /** * Strip hashline display prefixes from write content. * @@ -84,10 +113,7 @@ function stripWriteContent(session: ToolSession, content: string): { text: strin if (!resolveFileDisplayMode(session).hashLines) { return { text: content, stripped: false }; } - const lines = content.split("\n"); - const cleaned = stripHashlinePrefixes(lines); - if (cleaned === lines) return { text: content, stripped: false }; - return { text: cleaned.join("\n"), stripped: true }; + return stripWriteContentWithPotentialLooseHeader(content.split("\n")); } /** diff --git a/packages/coding-agent/test/core/hashline.test.ts b/packages/coding-agent/test/core/hashline.test.ts index a2b88c3ca..a70bbcdf3 100644 --- a/packages/coding-agent/test/core/hashline.test.ts +++ b/packages/coding-agent/test/core/hashline.test.ts @@ -3,7 +3,6 @@ import * as fs from "node:fs/promises"; import * as os from "node:os"; import * as path from "node:path"; import { - type ApplyOptions, applyEdits, buildCompactDiffPreview as buildCompactHashlineDiffPreview, detectLineEnding, @@ -38,9 +37,13 @@ import * as z from "zod/v4"; function applyHashlineEdits( text: string, edits: readonly Edit[], - options: ApplyOptions = {}, -): { text: string; lines: string; firstChangedLine?: number; warnings?: string[] } { - const r = applyEdits(text, [...edits], options); +): { + text: string; + lines: string; + firstChangedLine?: number; + warnings?: string[]; +} { + const r = applyEdits(text, [...edits]); return { ...r, lines: r.text }; } @@ -67,14 +70,12 @@ function tryRecoverHashlineWithCache(args: { currentText: string; tag: string; edits: readonly Edit[]; - options?: ApplyOptions; }): { text: string; lines: string; firstChangedLine: number | undefined; warnings: string[] } | null { const recovered = new Recovery(args.cache).tryRecover({ path: args.absolutePath, currentText: args.currentText, fileHash: args.tag, edits: args.edits, - options: args.options, }); return recovered ? { ...recovered, lines: recovered.text } : null; } @@ -87,7 +88,7 @@ beforeAll(async () => { }); const repl = (text: string): string => `+${text}`; -const repeat = (start: string, end = start): string => `^${start}-${end}`; +const repeat = (start: string, end = start): string => `&${start}..${end}`; const outputSep = ":"; const outputSepRe = ":"; @@ -104,17 +105,13 @@ function header(filePath: string, tag: string): string { } function sameLineRange(anchor: string): string { - return `${anchor}-${anchor}`; + return `${anchor} ${anchor}`; } function applyDiff(content: string, diff: string): string { return applyHashlineEdits(content, parseHashline(diff).edits).lines; } -function applyDiffWithPureInsertAutoDrop(content: string, diff: string): string { - return applyHashlineEdits(content, parseHashline(diff).edits, { autoDropPureInsertDuplicates: true }).lines; -} - async function withTempDir(fn: (tempDir: string) => Promise): Promise { const tempDir = await fs.mkdtemp(path.join(os.tmpdir(), "hashline-edit-")); try { @@ -189,7 +186,7 @@ describe("hashline normalization", () => { describe("hashline parser — range-anchor syntax", () => { it("keeps parsed edits reusable across different target snapshots", () => { const section = Patch.parseSingle( - ["¶a.ts", `${sameLineRange(tag(2, "bbb"))}:`, repeat(tag(2, "bbb")), repl("tail")].join("\n"), + ["¶a.ts", `${sameLineRange(tag(2, "bbb"))}`, repeat(tag(2, "bbb")), repl("tail")].join("\n"), ); expect(section.applyTo("aaa\nbbb").text).toBe("aaa\nbbb\ntail"); @@ -200,66 +197,66 @@ describe("hashline parser — range-anchor syntax", () => { it("inserts payload before/after a Lid, and at BOF/EOF", () => { const diff = [ - `${sameLineRange(tag(2, "bbb"))}:`, + `${sameLineRange(tag(2, "bbb"))}`, repl("before b"), repeat(tag(2, "bbb")), repl("after b"), - "BOF:", + "BOF", repl("top"), - "EOF:", + "EOF", repl("tail"), ].join("\n"); expect(applyDiff(content, diff)).toBe("top\naaa\nbefore b\nbbb\nafter b\nccc\ntail"); }); it("inserts after the final line without falling off the file", () => { - const diff = [`${sameLineRange(tag(3, "ccc"))}:`, repeat(tag(3, "ccc")), repl("tail")].join("\n"); + const diff = [`${sameLineRange(tag(3, "ccc"))}`, repeat(tag(3, "ccc")), repl("tail")].join("\n"); expect(applyDiff(content, diff)).toBe("aaa\nbbb\nccc\ntail"); }); - it("deletes a line or range via inline delete", () => { - expect(applyDiff(content, `${sameLineRange(tag(2, "bbb"))}:-`)).toBe("aaa\nccc"); - expect(applyDiff(content, `${tag(2, "bbb")}-${tag(3, "ccc")}:-`)).toBe("aaa"); + it("deletes a line or range via the standalone -A..B op", () => { + expect(applyDiff(content, `${tag(2, "bbb")} ${tag(2, "bbb")}`)).toBe("aaa\nccc"); + expect(applyDiff(content, `${tag(2, "bbb")} ${tag(3, "ccc")}`)).toBe("aaa"); }); it("replaces a line with one blank when given an explicit empty replace payload", () => { - const explicit = [`${sameLineRange(tag(2, "bbb"))}:`, repl("")].join("\n"); + const explicit = [`${sameLineRange(tag(2, "bbb"))}`, repl("")].join("\n"); expect(applyDiff(content, explicit)).toBe("aaa\n\nccc"); }); it("replaces one line or an inclusive range with payload lines", () => { - const single = [`${sameLineRange(tag(2, "bbb"))}:`, repl("BBB")].join("\n"); + const single = [`${sameLineRange(tag(2, "bbb"))}`, repl("BBB")].join("\n"); expect(applyDiff(content, single)).toBe("aaa\nBBB\nccc"); - const range = [`${tag(2, "bbb")}-${tag(3, "ccc")}:`, repl("BBB"), repl("CCC")].join("\n"); + const range = [`${tag(2, "bbb")} ${tag(3, "ccc")}`, repl("BBB"), repl("CCC")].join("\n"); expect(applyDiff(content, range)).toBe("aaa\nBBB\nCCC"); }); - it("accepts bare `A:` as a shorthand for `A-A:`", () => { + it("accepts bare `A=` as a shorthand for `A..A=`", () => { const anchor = tag(2, "bbb"); // `LINE:` is the exact shape `read` renders each file row as, so the // parser leniently treats it as `LINE-LINE:` for models that // reproduce the read-output shape as an anchor. - expect(applyDiff(content, `${anchor}:\n${repl("BBB")}`)).toBe("aaa\nBBB\nccc"); + expect(applyDiff(content, `${anchor}\n${repl("BBB")}`)).toBe("aaa\nBBB\nccc"); }); - it("replaces empty anchor blocks with one blank line", () => { + it("empty anchor body deletes the range entirely", () => { const anchor = tag(2, "bbb"); - expect(applyDiff(content, `${sameLineRange(anchor)}:`)).toBe("aaa\n\nccc"); - expect(applyDiff(content, `${anchor}-${tag(3, "ccc")}:`)).toBe("aaa\n"); + expect(applyDiff(content, `${sameLineRange(anchor)}`)).toBe("aaa\nccc"); + expect(applyDiff(content, `${anchor} ${tag(3, "ccc")}`)).toBe("aaa"); }); - it("rejects inline payload on anchor rows", () => { + it("rejects orphan inline-anchor shapes from old format", () => { const anchor = tag(2, "bbb"); - for (const diff of [`${sameLineRange(anchor)}:NEW`, `${anchor}-${tag(3, "ccc")}:NEW`, "BOF:NEW", "EOF:NEW"]) { - expect(() => parseHashline(diff)).toThrow(/Inline payload on the anchor line is rejected/); + for (const diff of [`${anchor}..${tag(3, "ccc")}=NEW`, "BOF=NEW", "EOF=NEW"]) { + expect(() => parseHashline(diff)).toThrow(/payload line has no preceding hunk header/); } }); it("emits body rows in textual order", () => { const diff = [ - `${sameLineRange(tag(2, "bbb"))}:`, + `${sameLineRange(tag(2, "bbb"))}`, repl("above 1"), repl("above 2"), repl("BBB"), @@ -270,9 +267,7 @@ describe("hashline parser — range-anchor syntax", () => { }); it("preserves the anchor when repeat rows re-emit it", () => { - const diff = [`${sameLineRange(tag(2, "bbb"))}:`, repl("before"), repeat(tag(2, "bbb")), repl("after")].join( - "\n", - ); + const diff = [`${sameLineRange(tag(2, "bbb"))}`, repl("before"), repeat(tag(2, "bbb")), repl("after")].join("\n"); expect(applyDiff(content, diff)).toBe("aaa\nbefore\nbbb\nafter\nccc"); }); @@ -280,230 +275,124 @@ describe("hashline parser — range-anchor syntax", () => { // `+` is the canonical sigil; payload rows like `+|literal` emit // `|literal` verbatim. Same for `^literal` and `↓literal` — none of // these are recognized sigils once they sit inside a `+TEXT` row. - const diff = [`${sameLineRange(tag(2, "bbb"))}:`, repl("|literal"), repl("^literal"), repl("↓literal")].join( - "\n", - ); + const diff = [`${sameLineRange(tag(2, "bbb"))}`, repl("|literal"), repl("^literal"), repl("↓literal")].join("\n"); expect(applyDiff(content, diff)).toBe("aaa\n|literal\n^literal\n↓literal\nccc"); }); it("accepts literal payload at virtual BOF/EOF anchors", () => { - expect(applyDiff(content, ["BOF:", repl("HEAD")].join("\n"))).toBe("HEAD\naaa\nbbb\nccc"); - expect(applyDiff(content, ["EOF:", repl("TAIL")].join("\n"))).toBe("aaa\nbbb\nccc\nTAIL"); + expect(applyDiff(content, ["BOF", repl("HEAD")].join("\n"))).toBe("HEAD\naaa\nbbb\nccc"); + expect(applyDiff(content, ["EOF", repl("TAIL")].join("\n"))).toBe("aaa\nbbb\nccc\nTAIL"); }); - it("rejects unprefixed payload continuation lines", () => { + it("auto-pipes unprefixed payload continuation lines as literal text", () => { const anchor = tag(2, "bbb"); - expect(() => parseHashline(`${sameLineRange(anchor)}:\n${repl("FIRST")}\nSECOND`)).toThrow(/must start with/); + const { edits, warnings } = parseHashline(`${sameLineRange(anchor)}\n${repl("FIRST")}\nSECOND`); + expect(applyHashlineEdits("aaa\nbbb\nccc", edits).lines).toBe("aaa\nFIRST\nSECOND\nccc"); + expect(warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); it("preserves whitespace-bearing payload exactly", () => { const anchor = tag(2, "bbb"); const payload = "\tconst streamKeepaliveMs = opts.streamKeepaliveMs;"; - expect(applyDiff(content, [`${sameLineRange(anchor)}:`, repeat(anchor), repl(payload)].join("\n"))).toBe( + expect(applyDiff(content, [`${sameLineRange(anchor)}`, repeat(anchor), repl(payload)].join("\n"))).toBe( `aaa\nbbb\n${payload}\nccc`, ); - expect(applyDiff(content, [`${sameLineRange(anchor)}:`, repl(payload), repeat(anchor)].join("\n"))).toBe( + expect(applyDiff(content, [`${sameLineRange(anchor)}`, repl(payload), repeat(anchor)].join("\n"))).toBe( `aaa\n${payload}\nbbb\nccc`, ); }); - it("auto-absorbs duplicated multiline prefix boundaries during replacement", () => { + it("keeps duplicated multiline replacement boundaries literal", () => { + const prefixSource = ["// one", "// two", "old();"].join("\n"); + const prefixDiff = [`${sameLineRange(tag(3, "old();"))}`, repl("// one"), repl("// two"), repl("new();")].join( + "\n", + ); + expect(applyDiff(prefixSource, prefixDiff)).toBe(["// one", "// two", "// one", "// two", "new();"].join("\n")); + + const suffixSource = ["old();", "// one", "// two"].join("\n"); + const suffixDiff = [`${sameLineRange(tag(1, "old();"))}`, repl("new();"), repl("// one"), repl("// two")].join( + "\n", + ); + expect(applyDiff(suffixSource, suffixDiff)).toBe(["new();", "// one", "// two", "// one", "// two"].join("\n")); + }); + + it("keeps duplicated structural replacement boundaries literal", () => { + const suffixSource = ["old();", "};"].join("\n"); + const suffixDiff = [`${sameLineRange(tag(1, "old();"))}`, repl("new();"), repl("};")].join("\n"); + expect(applyDiff(suffixSource, suffixDiff)).toBe(["new();", "};", "};"].join("\n")); + + const prefixSource = ["};", "old();"].join("\n"); + const prefixDiff = [`${sameLineRange(tag(2, "old();"))}`, repl("};"), repl("new();")].join("\n"); + expect(applyDiff(prefixSource, prefixDiff)).toBe(["};", "};", "new();"].join("\n")); + }); + + it("keeps duplicated single non-structural replacement boundaries literal", () => { + const prefixSource = ["const X = …", "", "const LEGACY = {", " a: 1,", "}"].join("\n"); + const prefixDiff = [`${tag(2, "")} ${tag(5, "}")}`, repl("const X = …")].join("\n"); + expect(applyDiff(prefixSource, prefixDiff)).toBe(["const X = …", "const X = …"].join("\n")); + + const suffixSource = ["## Legacy", "", "stale content", "", "## Subagents"].join("\n"); + const suffixDiff = [`${tag(1, "## Legacy")} ${tag(4, "")}`, repl("## Subagents")].join("\n"); + expect(applyDiff(suffixSource, suffixDiff)).toBe(["## Subagents", "## Subagents"].join("\n")); + }); + + it("does not emit warnings for duplicated replacement boundaries", () => { const source = ["// one", "// two", "old();"].join("\n"); - const diff = [`${sameLineRange(tag(3, "old();"))}:`, repl("// one"), repl("// two"), repl("new();")].join("\n"); - - expect(applyDiff(source, diff)).toBe(["// one", "// two", "new();"].join("\n")); - }); - - it("auto-absorbs duplicated multiline suffix boundaries during replacement", () => { - const source = ["old();", "// one", "// two"].join("\n"); - const diff = [`${sameLineRange(tag(1, "old();"))}:`, repl("new();"), repl("// one"), repl("// two")].join("\n"); - - expect(applyDiff(source, diff)).toBe(["new();", "// one", "// two"].join("\n")); - }); - - it("auto-absorbs a duplicated single structural suffix during replacement", () => { - const source = ["old();", "};"].join("\n"); - const diff = [`${sameLineRange(tag(1, "old();"))}:`, repl("new();"), repl("};")].join("\n"); - - expect(applyDiff(source, diff)).toBe(["new();", "};"].join("\n")); - }); - - it("auto-absorbs a duplicated single structural prefix during replacement", () => { - const source = ["};", "old();"].join("\n"); - const diff = [`${sameLineRange(tag(2, "old();"))}:`, repl("};"), repl("new();")].join("\n"); - - expect(applyDiff(source, diff)).toBe(["};", "new();"].join("\n")); - }); - - it("does not absorb a single structural replacement suffix when it preserves balance", () => { - // The replacement payload `if ok {` + `}` is itself net-zero, so the trailing - // `}` is a legitimate part of the new block, not a duplicate of the file's - // existing `}`. The single-line structural absorb must NOT fire here. - const source = ["old();", "}"].join("\n"); - const diff = [`${sameLineRange(tag(1, "old();"))}:`, repl("if ok {"), repl("}")].join("\n"); - - expect(applyDiff(source, diff)).toBe(["if ok {", "}", "}"].join("\n")); - }); - - it("does not auto-absorb a single duplicated boundary line", () => { - const source = ["keep", "old();"].join("\n"); - const diff = [`${sameLineRange(tag(2, "old();"))}:`, repl("keep"), repl("new();")].join("\n"); - - expect(applyDiff(source, diff)).toBe(["keep", "keep", "new();"].join("\n")); - }); - - it("does not auto-absorb a duplicate boundary that another op already targets", () => { - // Lines 3-4 ("X","Y") match the payload's trailing block, but line 4 - // is also the anchor of a separate insert. Absorbing it would silently - // steal that anchor and turn the insert into a replacement. - const source = ["A", "B", "X", "Y", "Z"].join("\n"); - const diff = [ - `${tag(1, "A")}-${tag(2, "B")}:`, - repl("alpha"), - repl("X"), - repl("Y"), - `${sameLineRange(tag(4, "Y"))}:`, - repl("extra"), - repeat(tag(4, "Y")), - ].join("\n"); - - expect(applyDiff(source, diff)).toBe(["alpha", "X", "Y", "X", "extra", "Y", "Z"].join("\n")); - }); - - it("surfaces a warning when boundary duplicates are auto-absorbed", () => { - const source = ["// one", "// two", "old();"].join("\n"); - const diff = [`${sameLineRange(tag(3, "old();"))}:`, repl("// one"), repl("// two"), repl("new();")].join("\n"); + const diff = [`${sameLineRange(tag(3, "old();"))}`, repl("// one"), repl("// two"), repl("new();")].join("\n"); const result = applyHashlineEdits(source, parseHashline(diff).edits); - expect(result.lines).toBe(["// one", "// two", "new();"].join("\n")); - expect(result.warnings).toBeDefined(); - expect(result.warnings).toEqual( - expect.arrayContaining([expect.stringMatching(/Auto-absorbed 2 duplicate line\(s\) above replacement/)]), - ); + expect(result.lines).toBe(["// one", "// two", "// one", "// two", "new();"].join("\n")); + expect(result.warnings).toBeUndefined(); }); - it("auto-absorbs a single duplicated non-structural prefix during replacement when opt-in is set", () => { - // Regression: `103-138:const X = …` over a file whose line 102 already - // reads `const X = …` produced two consecutive declarations. With the - // opt-in on, the leading boundary line gets dropped. - const source = ["const X = …", "", "const LEGACY = {", " a: 1,", "}"].join("\n"); - const diff = [`${tag(2, "")}-${tag(5, "}")}:`, repl("const X = …")].join("\n"); - - const result = applyHashlineEdits(source, parseHashline(diff).edits, { autoDropPureInsertDuplicates: true }); - expect(result.lines).toBe(["const X = …"].join("\n")); - expect(result.warnings).toEqual( - expect.arrayContaining([expect.stringMatching(/Auto-absorbed 1 duplicate line\(s\) above replacement/)]), - ); - }); - - it("auto-absorbs a single duplicated non-structural suffix during replacement when opt-in is set", () => { - // Regression: `93-104:## Subagents` over a file whose line 105 already - // reads `## Subagents` produced two consecutive headings. With the - // opt-in on, the trailing boundary line gets dropped. - const source = ["## Legacy", "", "stale content", "", "## Subagents"].join("\n"); - const diff = [`${tag(1, "## Legacy")}-${tag(4, "")}:`, repl("## Subagents")].join("\n"); - - const result = applyHashlineEdits(source, parseHashline(diff).edits, { autoDropPureInsertDuplicates: true }); - expect(result.lines).toBe(["## Subagents"].join("\n")); - expect(result.warnings).toEqual( - expect.arrayContaining([expect.stringMatching(/Auto-absorbed 1 duplicate line\(s\) below replacement/)]), - ); - }); - - it("preserves a legitimate single-line replacement that happens to match an adjacent line by default", () => { - // Without the opt-in, `2:foo` over `[1]foo,[2]bar,[3]baz` must still - // produce two consecutive `foo` lines. The non-structural single-line - // absorber stays gated on `autoDropPureInsertDuplicates`. + it("preserves a legitimate single-line replacement that happens to match an adjacent line", () => { const source = ["foo", "bar", "baz"].join("\n"); - const diff = [`${sameLineRange(tag(2, "bar"))}:`, repl("foo")].join("\n"); + const diff = [`${sameLineRange(tag(2, "bar"))}`, repl("foo")].join("\n"); expect(applyDiff(source, diff)).toBe(["foo", "foo", "baz"].join("\n")); }); - it("does not auto-drop generic (multi-line) pure-insert duplicate boundaries by default", () => { - // Multi-line context echo (`bbb`, `ccc`) is gated on the - // `autoDropPureInsertDuplicates` opt-in. Single-line pure-insert - // duplicates stay literal because they are ambiguous. - const source = ["aaa", "bbb", "ccc"].join("\n"); - const diff = ["EOF:", repl("bbb"), repl("ccc"), repl("NEW")].join("\n"); - expect(applyDiff(source, diff)).toBe("aaa\nbbb\nccc\nbbb\nccc\nNEW"); + it("keeps pure-insert payload that duplicates adjacent file context", () => { + const eofSource = ["aaa", "bbb", "ccc"].join("\n"); + const eofDiff = ["EOF", repl("bbb"), repl("ccc"), repl("NEW")].join("\n"); + expect(applyDiff(eofSource, eofDiff)).toBe("aaa\nbbb\nccc\nbbb\nccc\nNEW"); + + const bofSource = ["aaa", "bbb", "ccc", "ddd"].join("\n"); + const bofDiff = ["BOF", repl("NEW"), repl("aaa"), repl("bbb")].join("\n"); + expect(applyDiff(bofSource, bofDiff)).toBe("NEW\naaa\nbbb\naaa\nbbb\nccc\nddd"); }); - it("preserves a duplicated single structural suffix for pure insert by default", () => { + it("preserves duplicated structural pure-insert payload", () => { const source = ["if ok {", " keep();", " }"].join("\n"); - const diff = ["EOF:", repl(" added();"), repl(" }")].join("\n"); + const diff = ["EOF", repl(" added();"), repl(" }")].join("\n"); expect(applyDiff(source, diff)).toBe(["if ok {", " keep();", " }", " added();", " }"].join("\n")); }); - it("preserves a duplicated single structural prefix even when duplicate absorption is enabled", () => { - const source = [" });", "next();"].join("\n"); - const diff = [ - `${sameLineRange(tag(1, " });"))}:`, - repeat(tag(1, " });")), - repl(" });"), - repl("added();"), - ].join("\n"); - const result = applyHashlineEdits(source, parseHashline(diff).edits, { autoDropPureInsertDuplicates: true }); - - expect(result.lines).toBe([" });", " });", "added();", "next();"].join("\n")); - expect(result.warnings).toBeUndefined(); - }); - - it("preserves an intentional non-structural anchor duplicate for below insert by default", () => { + it("preserves an intentional non-structural anchor duplicate for below insert", () => { const source = ["aaa", "bbb", "ccc"].join("\n"); - const diff = [`${sameLineRange(tag(2, "bbb"))}:`, repeat(tag(2, "bbb")), repl("bbb"), repl("NEW")].join("\n"); + const diff = [`${sameLineRange(tag(2, "bbb"))}`, repeat(tag(2, "bbb")), repl("bbb"), repl("NEW")].join("\n"); expect(applyDiff(source, diff)).toBe("aaa\nbbb\nbbb\nNEW\nccc"); }); - it("preserves an intentional non-structural anchor duplicate for above insert by default", () => { + it("preserves an intentional non-structural anchor duplicate for above insert", () => { const source = ["aaa", "bbb", "ccc"].join("\n"); - const diff = [`${sameLineRange(tag(2, "bbb"))}:`, repl("NEW"), repl("bbb"), repeat(tag(2, "bbb"))].join("\n"); + const diff = [`${sameLineRange(tag(2, "bbb"))}`, repl("NEW"), repl("bbb"), repeat(tag(2, "bbb"))].join("\n"); expect(applyDiff(source, diff)).toBe("aaa\nNEW\nbbb\nbbb\nccc"); }); - it("does not drop a single structural pure-insert suffix when it preserves balance", () => { + it("keeps a single structural pure-insert suffix when it preserves balance", () => { const source = ["if outer {", "}"].join("\n"); - const diff = [`${sameLineRange(tag(2, "}"))}:`, repl("if inner {"), repl("}"), repeat(tag(2, "}"))].join("\n"); + const diff = [`${sameLineRange(tag(2, "}"))}`, repl("if inner {"), repl("}"), repeat(tag(2, "}"))].join("\n"); expect(applyDiff(source, diff)).toBe(["if outer {", "if inner {", "}", "}"].join("\n")); }); - it("auto-absorbs duplicated leading payload at EOF insert", () => { - const source = ["aaa", "bbb", "ccc"].join("\n"); - const diff = ["EOF:", repl("bbb"), repl("ccc"), repl("NEW")].join("\n"); - expect(applyDiffWithPureInsertAutoDrop(source, diff)).toBe("aaa\nbbb\nccc\nNEW"); - }); - - it("auto-absorbs duplicated trailing payload at BOF insert", () => { - const source = ["aaa", "bbb", "ccc", "ddd"].join("\n"); - const diff = ["BOF:", repl("NEW"), repl("aaa"), repl("bbb")].join("\n"); - expect(applyDiffWithPureInsertAutoDrop(source, diff)).toBe("NEW\naaa\nbbb\nccc\nddd"); - }); - - it("preserves a single duplicated anchor line in a pure insert even when generic duplicate absorption is enabled", () => { - const source = ["aaa", "bbb", "ccc"].join("\n"); - const diff = ["EOF:", repl("ccc"), repl("NEW")].join("\n"); - - expect(applyDiffWithPureInsertAutoDrop(source, diff)).toBe("aaa\nbbb\nccc\nccc\nNEW"); - }); - - it("surfaces a warning when pure-insert duplicates are auto-dropped", () => { - const source = ["aaa", "bbb", "ccc"].join("\n"); - const diff = ["EOF:", repl("bbb"), repl("ccc"), repl("NEW")].join("\n"); - const result = applyHashlineEdits(source, parseHashline(diff).edits, { autoDropPureInsertDuplicates: true }); - expect(result.lines).toBe("aaa\nbbb\nccc\nNEW"); - expect(result.warnings).toBeDefined(); - expect(result.warnings).toEqual( - expect.arrayContaining([expect.stringMatching(/Auto-dropped 2 duplicate line\(s\) at the start of insert/)]), - ); - }); - it("preserves payload text exactly", () => { const diff = [ - `${sameLineRange(tag(2, "bbb"))}:`, + `${sameLineRange(tag(2, "bbb"))}`, repl(""), repl("# not a header"), repl("+ not an op"), @@ -514,7 +403,7 @@ describe("hashline parser — range-anchor syntax", () => { }); it("treats explicit empty replace payload rows as blank lines", () => { - const diff = [`${sameLineRange(tag(2, "bbb"))}:`, repl("first"), repl(""), repl(""), repl("after")].join("\n"); + const diff = [`${sameLineRange(tag(2, "bbb"))}`, repl("first"), repl(""), repl(""), repl("after")].join("\n"); expect(applyDiff(content, diff)).toBe("aaa\nfirst\n\n\nafter\nccc"); }); @@ -522,34 +411,34 @@ describe("hashline parser — range-anchor syntax", () => { const diff = [ "# This is a comment line from a model explanation.", "## Another comment line.", - `${sameLineRange(tag(2, "bbb"))}:`, + `${sameLineRange(tag(2, "bbb"))}`, repl("BBB"), ].join("\n"); expect(applyDiff(content, diff)).toBe("aaa\nBBB\nccc"); }); it("does not skip comment lines when they are not immediately before an operation", () => { - const diff = ["# This is a stray comment.", "", `${sameLineRange(tag(2, "bbb"))}:`, repl("BBB")].join("\n"); + const diff = ["# This is a stray comment.", "", `${sameLineRange(tag(2, "bbb"))}`, repl("BBB")].join("\n"); expect(() => parseHashline(diff)).toThrow(/payload line has no preceding/); }); it("preserves raw blank separators between ops", () => { const diff = [ - `${sameLineRange(tag(1, "aaa"))}:`, + `${sameLineRange(tag(1, "aaa"))}`, repl("AAA"), "", "", - `${sameLineRange(tag(3, "ccc"))}:`, + `${sameLineRange(tag(3, "ccc"))}`, repl("CCC"), ].join("\n"); expect(applyDiff(content, diff)).toBe("AAA\nbbb\nCCC"); }); it("inserts explicit blank lines above and below an anchor", () => { - expect(applyDiff(content, `${sameLineRange(tag(1, "aaa"))}:\n${repl("")}\n${repeat(tag(1, "aaa"))}`)).toBe( + expect(applyDiff(content, `${sameLineRange(tag(1, "aaa"))}\n${repl("")}\n${repeat(tag(1, "aaa"))}`)).toBe( "\naaa\nbbb\nccc", ); - expect(applyDiff(content, `${sameLineRange(tag(1, "aaa"))}:\n${repeat(tag(1, "aaa"))}\n${repl("")}`)).toBe( + expect(applyDiff(content, `${sameLineRange(tag(1, "aaa"))}\n${repeat(tag(1, "aaa"))}\n${repl("")}`)).toBe( "aaa\n\nbbb\nccc", ); }); @@ -558,31 +447,15 @@ describe("hashline parser — range-anchor syntax", () => { expect(() => parseHashline(repl("orphan")).edits).toThrow(/payload line has no preceding/); }); - it("rejects ranges with `..` separator", () => { - expect(() => parseHashline(`${tag(2, "bbb")}..${sameLineRange(tag(3, "ccc"))}:\n${repl("BBB")}`).edits).toThrow( - /payload line has no preceding/, - ); + it("accepts `A..B` with empty body as a delete", () => { + const result = parseHashline(`${sameLineRange(tag(2, "bbb"))}`); + expect(result.edits).toEqual([{ kind: "delete", anchor: { line: 2 }, lineNum: 1, index: 0 }]); }); - - it("describes the new block shape on unknown-op lines", () => { - expect(() => parseHashline(`-${sameLineRange(tag(2, "bbb"))}`).edits).toThrow(/Use A-B:, A-B:-, BOF:, or EOF:/); - }); - it("rejects `LINE:TEXT` copied verbatim from read output", () => { const anchor = tag(2, "bbb"); - expect(() => parseHashline(`${sameLineRange(anchor)}:BBB`)).toThrow( - /Inline payload on the anchor line is rejected/, - ); - expect(() => parseHashline(`${anchor}-${tag(3, "ccc")}:BBB`)).toThrow( - /Inline payload on the anchor line is rejected/, - ); - }); - - it("leniently strips `*`/`>` line-marker decoration from anchors", () => { - const anchor = tag(2, "bbb"); - expect(applyDiff(content, `*${sameLineRange(anchor)}:\n${repl("BBB")}`)).toBe("aaa\nBBB\nccc"); - expect(applyDiff(content, `>${sameLineRange(anchor)}:\n${repl("X")}\n${repeat(anchor)}`)).toBe( - "aaa\nX\nbbb\nccc", + expect(() => parseHashline(`${sameLineRange(anchor)}:BBB`)).toThrow(/payload line has no preceding hunk header/); + expect(() => parseHashline(`${anchor}..${tag(3, "ccc")}:BBB`)).toThrow( + /payload line has no preceding hunk header/, ); }); @@ -593,31 +466,31 @@ describe("hashline parser — range-anchor syntax", () => { it("preserves payload text containing arrow sigils after the leading payload sigil", () => { const anchor = tag(2, "bbb"); - expect(applyDiff(content, `${sameLineRange(anchor)}:\n${repl("bbb↑")}\n${repl("tail↓")}`)).toBe( + expect(applyDiff(content, `${sameLineRange(anchor)}\n${repl("bbb↑")}\n${repl("tail↓")}`)).toBe( "aaa\nbbb↑\ntail↓\nccc", ); }); it("accepts BOF/EOF inserts with literal payload rows", () => { - expect(applyDiff(content, `BOF:\n${repl("HEAD")}`)).toBe("HEAD\naaa\nbbb\nccc"); - expect(applyDiff(content, `EOF:\n${repl("TAIL")}`)).toBe("aaa\nbbb\nccc\nTAIL"); + expect(applyDiff(content, `BOF\n${repl("HEAD")}`)).toBe("HEAD\naaa\nbbb\nccc"); + expect(applyDiff(content, `EOF\n${repl("TAIL")}`)).toBe("aaa\nbbb\nccc\nTAIL"); }); it("coalesces two replace ops targeting the same single line (last wins)", () => { - const diff = `${sameLineRange(tag(2, "bbb"))}:\n${repl("BBB")}\n${sameLineRange(tag(2, "bbb"))}:\n${repl("BBB2")}`; + const diff = `${sameLineRange(tag(2, "bbb"))}\n${repl("BBB")}\n${sameLineRange(tag(2, "bbb"))}\n${repl("BBB2")}`; const { edits, warnings } = parseHashline(diff); expect(applyHashlineEdits("aaa\nbbb\nccc", edits).lines).toBe("aaa\nBBB2\nccc"); expect(warnings).toEqual([ - "Detected two identical-range hashline blocks; kept only the second block. Issue ONE block per range — payload is the final desired content, never both old and new.", + "Detected two identical-range hashline hunks; kept only the second hunk. Issue ONE hunk per range — payload is the final desired content, never both old and new.", ]); }); it("coalesces two replace ops covering the same range (before/after-block pattern, last wins)", () => { - const diff = `${tag(2, "bbb")}-${tag(3, "ccc")}:\n${repl("OLD")}\n${repl("OLD2")}\n${tag(2, "bbb")}-${tag(3, "ccc")}:\n${repl("NEW")}\n${repl("NEW2")}`; + const diff = `${tag(2, "bbb")} ${tag(3, "ccc")}\n${repl("OLD")}\n${repl("OLD2")}\n${tag(2, "bbb")} ${tag(3, "ccc")}\n${repl("NEW")}\n${repl("NEW2")}`; const { edits, warnings } = parseHashline(diff); expect(applyHashlineEdits("aaa\nbbb\nccc\nddd", edits).lines).toBe("aaa\nNEW\nNEW2\nddd"); expect(warnings).toEqual([ - "Detected two identical-range hashline blocks; kept only the second block. Issue ONE block per range — payload is the final desired content, never both old and new.", + "Detected two identical-range hashline hunks; kept only the second hunk. Issue ONE hunk per range — payload is the final desired content, never both old and new.", ]); }); @@ -625,12 +498,12 @@ describe("hashline parser — range-anchor syntax", () => { // 3-5 extends past the outer 2-4, so it is neither identical nor contained. // The inner anchors still clash with the outer range's deletes and the // post-hoc validator catches the overlap. - const diff = `${tag(2, "bbb")}-${tag(4, "ddd")}:\n${repl("NEW1")}\n${tag(3, "ccc")}-${tag(5, "eee")}:\n${repl("NEW2")}`; - expect(() => parseHashline(diff).edits).toThrow(/anchor line 3 is already targeted by another op on line 1/); + const diff = `${tag(2, "bbb")} ${tag(4, "ddd")}\n${repl("NEW1")}\n${tag(3, "ccc")} ${tag(5, "eee")}\n${repl("NEW2")}`; + expect(() => parseHashline(diff).edits).toThrow(/anchor line 3 is already targeted by another hunk on line 1/); }); it("uses `|` payload lines inside a multi-line replacement", () => { - const diff = `${tag(2, "bbb")}-${tag(4, "ddd")}:\n${repl("line one")}\n${repl("line two")}\n${repl("line three")}`; + const diff = `${tag(2, "bbb")} ${tag(4, "ddd")}\n${repl("line one")}\n${repl("line two")}\n${repl("line three")}`; const { edits, warnings } = parseHashline(diff); expect(applyHashlineEdits("aaa\nbbb\nccc\nddd\neee", edits).lines).toBe( "aaa\nline one\nline two\nline three\neee", @@ -638,13 +511,15 @@ describe("hashline parser — range-anchor syntax", () => { expect(warnings).toEqual([]); }); - it("rejects read-output `N:TEXT` lines inside a pending `A-B:` block", () => { - const diff = `${tag(2, "bbb")}-${tag(4, "ddd")}:\n${repl("line one")}\n${sameLineRange(tag(3, "ccc"))}:line two`; - expect(() => parseHashline(diff)).toThrow(/Inline payload on the anchor line is rejected/); + it("auto-pipes read-output `N:TEXT` lines inside a pending hunk as literal text", () => { + const diff = `${tag(2, "bbb")} ${tag(4, "ddd")}\n${repl("line one")}\n${sameLineRange(tag(3, "ccc"))}:line two`; + const { edits, warnings } = parseHashline(diff); + expect(applyHashlineEdits("aaa\nbbb\nccc\nddd\neee", edits).lines).toBe("aaa\nline one\n3 3:line two\neee"); + expect(warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); it("treats `N:` outside the pending range as a separate op", () => { - const diff = `${tag(2, "bbb")}-${tag(3, "ccc")}:\n${repl("line one")}\n${sameLineRange(tag(5, "eee"))}:\n${repl("line five")}`; + const diff = `${tag(2, "bbb")} ${tag(3, "ccc")}\n${repl("line one")}\n${sameLineRange(tag(5, "eee"))}\n${repl("line five")}`; const { edits, warnings } = parseHashline(diff); expect(applyHashlineEdits("aaa\nbbb\nccc\nddd\neee\nfff", edits).lines).toBe( "aaa\nline one\nddd\nline five\nfff", @@ -653,12 +528,12 @@ describe("hashline parser — range-anchor syntax", () => { }); it("accepts multiple literal rows before a repeated anchor", () => { - const diff = `${sameLineRange(tag(2, "bbb"))}:\n${repl("X")}\n${repl("Y")}\n${repeat(tag(2, "bbb"))}`; + const diff = `${sameLineRange(tag(2, "bbb"))}\n${repl("X")}\n${repl("Y")}\n${repeat(tag(2, "bbb"))}`; expect(applyDiff(content, diff)).toBe("aaa\nX\nY\nbbb\nccc"); }); it("accepts a replace alongside surrounding literal rows", () => { - const diff = `${sameLineRange(tag(2, "bbb"))}:\n${repl("ABOVE")}\n${repl("NEW")}`; + const diff = `${sameLineRange(tag(2, "bbb"))}\n${repl("ABOVE")}\n${repl("NEW")}`; expect(applyDiff(content, diff)).toBe("aaa\nABOVE\nNEW\nccc"); }); }); @@ -669,66 +544,62 @@ describe("hashline — snapshot tag binding", () => { }); it("applies line-number edits without per-anchor hash validation", () => { - const diff = `${sameLineRange(tag(2, "bbb"))}:\n${repl("BBB")}`; + const diff = `${sameLineRange(tag(2, "bbb"))}\n${repl("BBB")}`; expect(applyDiff("aaa\nbbb\nccc", diff)).toBe("aaa\nBBB\nccc"); }); }); -describe("splitHashlineInput — ¶ headers", () => { - it("extracts path, snapshot tag, and diff body from ¶path#tag header", () => { - const input = [`¶src/foo.ts#0A3`, `${sameLineRange(tag(2, "bbb"))}:`, repl("BBB")].join("\n"); +describe("splitHashlineInput — @ headers", () => { + it("extracts path, snapshot tag, and diff body from @path#tag header", () => { + const input = [`¶src/foo.ts#0A3`, `${sameLineRange(tag(2, "bbb"))}`, repl("BBB")].join("\n"); expect(splitHashlineInput(input)).toEqual({ path: "src/foo.ts", fileHash: "0A3", - diff: `${sameLineRange(tag(2, "bbb"))}:\n${repl("BBB")}`, + diff: `${sameLineRange(tag(2, "bbb"))}\n${repl("BBB")}`, }); }); it("strips leading blank lines", () => { - expect(splitHashlineInput(`\n¶foo.ts\nBOF:\n${repl("x")}`)).toEqual({ + expect(splitHashlineInput(`\n¶foo.ts\nBOF\n${repl("x")}`)).toEqual({ path: "foo.ts", - diff: `BOF:\n${repl("x")}`, + diff: `BOF\n${repl("x")}`, }); }); it("normalizes cwd-prefixed absolute paths to cwd-relative paths", () => { const cwd = process.cwd(); const absolute = path.join(cwd, "src", "foo.ts"); - expect(splitHashlineInput(`¶${absolute}\nBOF:\n${repl("x")}`, { cwd }).path).toBe("src/foo.ts"); + expect(splitHashlineInput(`¶${absolute}\nBOF\n${repl("x")}`, { cwd }).path).toBe("src/foo.ts"); }); it("uses explicit fallback path only when input has recognizable operations", () => { - expect(splitHashlineInput(`BOF:\n${repl("x")}`, { path: "a.ts" })).toEqual({ + expect(splitHashlineInput(`BOF\n${repl("x")}`, { path: "a.ts" })).toEqual({ path: "a.ts", - diff: `BOF:\n${repl("x")}`, + diff: `BOF\n${repl("x")}`, }); expect(() => splitHashlineInput("plain text", { path: "a.ts" })).toThrow(/must begin with/); }); it("splits multiple edit sections", () => { - const input = ["¶a.ts", "BOF:", repl("a"), "¶b.ts", "EOF:", repl("b")].join("\n"); + const input = ["¶a.ts", "BOF", repl("a"), "¶b.ts", "EOF", repl("b")].join("\n"); expect(splitHashlineInputs(input)).toEqual([ - { path: "a.ts", diff: `BOF:\n${repl("a")}` }, - { path: "b.ts", diff: `EOF:\n${repl("b")}` }, + { path: "a.ts", diff: `BOF\n${repl("a")}` }, + { path: "b.ts", diff: `EOF\n${repl("b")}` }, ]); }); - - it("tolerates extra ¶ chars on the section header", () => { - const input = ["¶¶a.ts", "BOF:", repl("a"), "¶¶¶b.ts", "EOF:", repl("b")].join("\n"); - expect(splitHashlineInputs(input)).toEqual([ - { path: "a.ts", diff: `BOF:\n${repl("a")}` }, - { path: "b.ts", diff: `EOF:\n${repl("b")}` }, - ]); + it("rejects a unified-diff hunk header on the first line as contamination", () => { + const input = ["@@ -1,3 +1,3 @@", "BOF", repl("x")].join("\n"); + expect(() => splitHashlineInputs(input)).toThrow(/unified-diff hunk header/); }); - it("silently drops a duplicate header with no operations between them", () => { - const input = ["¶¶src/foo.ts", "¶¶src/foo.ts", "BOF:", repl("x")].join("\n"); - expect(splitHashlineInputs(input)).toEqual([{ path: "src/foo.ts", diff: `BOF:\n${repl("x")}` }]); + it("rejects a unified-diff hunk header (`-N,M +N,M`)", () => { + const input = ["@@ -1,3 +1,3 @@", "BOF", repl("x")].join("\n"); + expect(() => splitHashlineInputs(input)).toThrow(/unified-diff hunk header/); }); it("silently drops a trailing header with no operations", () => { - const input = ["¶¶a.ts", "BOF:", repl("a"), "¶¶b.ts"].join("\n"); - expect(splitHashlineInputs(input)).toEqual([{ path: "a.ts", diff: `BOF:\n${repl("a")}` }]); + const input = ["¶a.ts", "BOF", repl("a"), "¶b.ts"].join("\n"); + expect(splitHashlineInputs(input)).toEqual([{ path: "a.ts", diff: `BOF\n${repl("a")}` }]); }); }); @@ -745,10 +616,10 @@ it("preflights write policy for every section before committing a batch", async const bTag = recordFullSnapshot(snapshots, "b.ts", "bbb\n"); const input = [ header("a.ts", aTag), - `${sameLineRange(tag(1, "aaa"))}:`, + `${sameLineRange(tag(1, "aaa"))}`, repl("AAA"), header("b.ts", bTag), - `${sameLineRange(tag(1, "bbb"))}:`, + `${sameLineRange(tag(1, "bbb"))}`, repl("BBB"), ].join("\n"); @@ -762,26 +633,26 @@ it("preflights write policy for every section before committing a batch", async describe("hashline executor", () => { it("creates a missing file with a file-scoped insert", async () => { await withTempDir(async tempDir => { - const input = `¶new.ts\nBOF:\n${repl("export const x = 1;")}\n`; + const input = `¶new.ts\nBOF\n${repl("export const x = 1;")}\n`; const result = await executeHashlineSingle(hashlineExecuteOptions(tempDir, input)); expect(result.content[0]?.type === "text" ? result.content[0].text : "").toContain("¶new.ts#"); expect(await Bun.file(path.join(tempDir, "new.ts")).text()).toBe("export const x = 1;"); }); }); - it("honors the pure-insert duplicate auto-drop setting", async () => { + it("applies duplicate pure-insert payload literally", async () => { await withTempDir(async tempDir => { const filePath = path.join(tempDir, "a.ts"); const source = ["aaa", "bbb", "ccc"].join("\n"); - const input = `¶a.ts\nEOF:\n${repl("bbb")}\n${repl("ccc")}\n${repl("NEW")}\n`; - await Bun.write(filePath, source); - await executeHashlineSingle(hashlineExecuteOptions(tempDir, input)); - expect(await Bun.file(filePath).text()).toBe("aaa\nbbb\nccc\nbbb\nccc\nNEW"); + const input = `¶a.ts\nEOF\n${repl("bbb")}\n${repl("ccc")}\n${repl("NEW")}\n`; + const session = makeHashlineSession(tempDir); await Bun.write(filePath, source); - const enabled = Settings.isolated({ "edit.hashlineAutoDropPureInsertDuplicates": true }); - const result = await executeHashlineSingle(hashlineExecuteOptions(tempDir, input, enabled)); - expect(await Bun.file(filePath).text()).toBe("aaa\nbbb\nccc\nNEW"); - expect(result.content[0]?.type === "text" ? result.content[0].text : "").toContain("Auto-dropped"); + const result = await executeHashlineSingle(hashlineExecuteOptions(tempDir, input, undefined, session)); + const text = result.content[0]?.type === "text" ? result.content[0].text : ""; + + expect(await Bun.file(filePath).text()).toBe("aaa\nbbb\nccc\nbbb\nccc\nNEW"); + expect(text).not.toContain("Auto-dropped"); + expect(text).not.toContain("Auto-absorbed"); }); }); @@ -794,7 +665,7 @@ describe("hashline executor", () => { const sourceTag = recordFullSnapshot(getFileReadCache(session), filePath, source); // Replace line 2 with `bbb` — identical to the file content. The // patch applies but produces no change. - const input = `${header("a.ts", sourceTag)}\n${sameLineRange(tag(2, "bbb"))}:\n${repl("bbb")}\n`; + const input = `${header("a.ts", sourceTag)}\n${sameLineRange(tag(2, "bbb"))}\n${repl("bbb")}\n`; const result = await executeHashlineSingle(hashlineExecuteOptions(tempDir, input, undefined, session)); const text = result.content[0]?.type === "text" ? result.content[0].text : ""; expect(text).toContain("parsed and applied cleanly, but produced no change"); @@ -816,10 +687,10 @@ describe("hashline executor", () => { const bHeader = "¶b.ts#fff"; const input = [ header("a.ts", aTag), - `${sameLineRange(tag(1, "aaa"))}:`, + `${sameLineRange(tag(1, "aaa"))}`, repl("AAA"), bHeader, - `${sameLineRange(tag(1, "bbb"))}:`, + `${sameLineRange(tag(1, "bbb"))}`, repl("BBB"), ].join("\n"); @@ -840,10 +711,10 @@ describe("hashline executor", () => { const sourceTag = recordFullSnapshot(getFileReadCache(session), filePath, source); const input = [ header("a.ts", sourceTag), - `${sameLineRange(tag(1, "one"))}:`, + `${sameLineRange(tag(1, "one"))}`, repl("ONE"), header("./a.ts", sourceTag), - `${sameLineRange(tag(2, "two"))}:`, + `${sameLineRange(tag(2, "two"))}`, repl("TWO"), ].join("\n"); @@ -869,7 +740,7 @@ describe("hashline executor", () => { // validation outright. const input = [ header("a.ts", originalTag), - `${sameLineRange(tag(2, "L2"))}:`, + `${sameLineRange(tag(2, "L2"))}`, repl("L2a"), repl("L2b"), repl("L2c"), @@ -880,7 +751,7 @@ describe("hashline executor", () => { repl("L2h"), repl("L2i"), header("a.ts", originalTag), - `${sameLineRange(tag(8, "L8"))}:`, + `${sameLineRange(tag(8, "L8"))}`, repeat(tag(8, "L8")), repl("INSERTED"), ].join("\n"); @@ -927,15 +798,15 @@ describe("hashlineEditParamsSchema — payload shape", () => { }); it("tolerates provider extra fields without declaring `path`", () => { - expect(hashlineEditParamsSchema.safeParse({ path: "x.ts", input: `¶x.ts\nBOF:\n${repl("x")}` }).success).toBe( + expect(hashlineEditParamsSchema.safeParse({ path: "x.ts", input: `¶x.ts\nBOF\n${repl("x")}` }).success).toBe( true, ); }); it("accepts `_input` as a provider-emitted alias for `input`", () => { - const parsed = hashlineEditParamsSchema.safeParse({ _input: `¶x.ts\nBOF:\n${repl("x")}` }); + const parsed = hashlineEditParamsSchema.safeParse({ _input: `¶x.ts\nBOF\n${repl("x")}` }); expect(parsed.success).toBe(true); - if (parsed.success) expect(parsed.data.input).toBe(`¶x.ts\nBOF:\n${repl("x")}`); + if (parsed.success) expect(parsed.data.input).toBe(`¶x.ts\nBOF\n${repl("x")}`); }); it("still requires `input`", () => { @@ -1000,7 +871,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { await Bun.write(filePath, `${v1Lines.join("\n")}\n`); // Model authors anchor against V0 — line 2 is "L2" in V0. - const input = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(2, "L2"))}:\n${repl("L2-MODEL")}\n`; + const input = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(2, "L2"))}\n${repl("L2-MODEL")}\n`; const result = await executeHashlineSingle(hashlineExecuteOptions(tempDir, input, undefined, session)); const finalLines = (await Bun.file(filePath).text()).replace(/\n$/, "").split("\n"); @@ -1033,7 +904,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { v1Lines[5] = "L6-CHANGED"; await Bun.write(filePath, `${v1Lines.join("\n")}\n`); - const input = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(6, "L6"))}:\n${repl("L6-MODEL")}\n`; + const input = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(6, "L6"))}\n${repl("L6-MODEL")}\n`; await expect( executeHashlineSingle(hashlineExecuteOptions(tempDir, input, undefined, session)), ).rejects.toThrow(HashlineMismatchError); @@ -1051,7 +922,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { // Live file is completely different — patch context cannot match even // with fuzz tolerance. const currentText = "totally\nunrelated\ncontent\nhere\nnow\n"; - const edits = parseHashline(`${sameLineRange(tag(2, "beta"))}:\n${repl("BETA-MODEL")}`).edits; + const edits = parseHashline(`${sameLineRange(tag(2, "beta"))}\n${repl("BETA-MODEL")}`).edits; const recovered = tryRecoverHashlineWithCache({ cache, @@ -1059,7 +930,6 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { currentText, edits, tag: snapshotTag, - options: {}, }); expect(recovered).toBeNull(); }); @@ -1086,7 +956,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { // First edit: change line 2 : BETA. After the write, the cache should // reflect V1 (post-edit), not V0. - const firstInput = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(2, "beta"))}:\n${repl("BETA")}\n`; + const firstInput = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(2, "beta"))}\n${repl("BETA")}\n`; await executeHashlineSingle(hashlineExecuteOptions(tempDir, firstInput, undefined, session)); const v1Lines = ["alpha", "BETA", "gamma", "delta", "epsilon"]; const v1Text = `${v1Lines.join("\n")}\n`; @@ -1104,7 +974,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { const v2Lines = ["H1", "H2", "H3", "H4", "H5", "H6", "H7", ...v1Lines]; await Bun.write(filePath, `${v2Lines.join("\n")}\n`); - const secondInput = `${header("a.ts", v1Tag)}\n${sameLineRange(tag(3, "gamma"))}:\n${repl("GAMMA")}\n`; + const secondInput = `${header("a.ts", v1Tag)}\n${sameLineRange(tag(3, "gamma"))}\n${repl("GAMMA")}\n`; const result = await executeHashlineSingle(hashlineExecuteOptions(tempDir, secondInput, undefined, session)); const finalLines = (await Bun.file(filePath).text()).replace(/\n$/, "").split("\n"); @@ -1128,7 +998,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { const v0Tag = recordFullSnapshot(getFileReadCache(session), filePath, v0Text); // First edit lands cleanly against v0: line 5 becomes L5-FIRST. - const firstInput = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(5, "L5"))}:\n${repl("L5-FIRST")}\n`; + const firstInput = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(5, "L5"))}\n${repl("L5-FIRST")}\n`; await executeHashlineSingle(hashlineExecuteOptions(tempDir, firstInput, undefined, session)); const v1Lines = [...v0Lines]; @@ -1139,7 +1009,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { // again targets line 5 — the very line the first edit rewrote. // Recovery must refuse so the model re-reads instead of silently // overwriting L5-FIRST with payload authored against L5. - const secondInput = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(5, "L5"))}:\n${repl("L5-SECOND")}\n`; + const secondInput = `${header("a.ts", v0Tag)}\n${sameLineRange(tag(5, "L5"))}\n${repl("L5-SECOND")}\n`; await expect( executeHashlineSingle(hashlineExecuteOptions(tempDir, secondInput, undefined, session)), ).rejects.toThrow(HashlineMismatchError); @@ -1162,8 +1032,7 @@ describe("hashline — anchor-stale recovery via read snapshot cache", () => { absolutePath: fakePath, currentText, tag: v0Tag, - edits: parseHashline(`10-10:\n${repl("L10-EDITED")}`).edits, - options: {}, + edits: parseHashline(`10 10\n${repl("L10-EDITED")}`).edits, }); expect(recovered).not.toBeNull(); @@ -1219,11 +1088,11 @@ describe("hashline *** Abort recovery sentinel (harmony-leak mitigation)", () => it("parser breaks at *** Abort silently (no warning)", () => { const diff = [ - `${sameLineRange(tag(1, "alpha"))}:`, + `${sameLineRange(tag(1, "alpha"))}`, repeat(tag(1, "alpha")), repl("HELLO"), sentinel, - `${sameLineRange(tag(99, "junk"))}:`, + `${sameLineRange(tag(99, "junk"))}`, repeat(tag(99, "junk")), repl("never"), ].join("\n"); @@ -1238,7 +1107,7 @@ describe("hashline *** Abort recovery sentinel (harmony-leak mitigation)", () => it("appended sentinel from harmony-leak truncation: ops above are preserved", () => { // Mirrors the exact shape harmony-leak emits inside a single section. - const diff = `${sameLineRange(tag(1, "alpha"))}:\n${repeat(tag(1, "alpha"))}\n${repl("KEPT")}\n*** Abort\n`; + const diff = `${sameLineRange(tag(1, "alpha"))}\n${repeat(tag(1, "alpha"))}\n${repl("KEPT")}\n*** Abort\n`; const { edits, warnings } = parseHashline(diff); expect(edits).toHaveLength(3); expect(edits[1]).toMatchObject({ text: "KEPT" }); @@ -1248,12 +1117,12 @@ describe("hashline *** Abort recovery sentinel (harmony-leak mitigation)", () => it("splitter respects *** Abort like *** End Patch", () => { const input = [ `¶a.ts`, - `${sameLineRange(tag(1, "alpha"))}:`, + `${sameLineRange(tag(1, "alpha"))}`, repeat(tag(1, "alpha")), repl("a-payload"), sentinel, `¶b.ts`, - `${sameLineRange(tag(1, "beta"))}:`, + `${sameLineRange(tag(1, "beta"))}`, repeat(tag(1, "beta")), repl("never-emitted"), ].join("\n"); @@ -1264,7 +1133,7 @@ describe("hashline *** Abort recovery sentinel (harmony-leak mitigation)", () => }); it("clean input without sentinel produces no warning", () => { - const diff = `${sameLineRange(tag(1, "alpha"))}:\n${repeat(tag(1, "alpha"))}\n${repl("PAYLOAD")}\n`; + const diff = `${sameLineRange(tag(1, "alpha"))}\n${repeat(tag(1, "alpha"))}\n${repl("PAYLOAD")}\n`; const { warnings } = parseHashline(diff); expect(warnings).toEqual([]); }); @@ -1273,33 +1142,33 @@ describe("hashline *** Abort recovery sentinel (harmony-leak mitigation)", () => describe("hashline parser — delete and empty-block semantics", () => { it("inline delete deletes a single line", () => { const text = "line1\nline2\nline3\n"; - const { diff } = splitHashlineInput(`¶a.ts\n2-2:-\n`); + const { diff } = splitHashlineInput(`¶a.ts\n2 2\n`); expect(applyDiff(text, diff)).toBe("line1\nline3\n"); }); it("inline delete deletes the range", () => { const text = "line1\nline2\nline3\nline4\n"; - const { diff } = splitHashlineInput(`¶a.ts\n2-3:-\n`); + const { diff } = splitHashlineInput(`¶a.ts\n2 3\n`); expect(applyDiff(text, diff)).toBe("line1\nline4\n"); }); - it("an `A-B:` anchor with no payload becomes a blank-line replacement", () => { + it("an empty `A..B` deletes the range (not a blank-line replace)", () => { const text = "line1\nline2\nline3\n"; - const { diff } = splitHashlineInput(`¶a.ts\n2-2:\n`); - expect(applyDiff(text, diff)).toBe("line1\n\nline3\n"); + const { diff } = splitHashlineInput(`¶a.ts\n2 2\n`); + expect(applyDiff(text, diff)).toBe("line1\nline3\n"); }); - it("`A-B:` with inline body is still rejected", () => { - const { diff } = splitHashlineInput(`¶a.ts\n2-2:replacement\n`); - expect(() => parseHashline(diff)).toThrow(/Inline payload on the anchor line is rejected/); + it("`2..2=replacement` (old format) parses as orphan body, not as inline payload", () => { + const { diff } = splitHashlineInput(`¶a.ts\n2..2=replacement\n`); + expect(() => parseHashline(diff)).toThrow(/payload line has no preceding hunk header/); }); it("explicit empty literal rows insert blank lines when the anchor is repeated", () => { const text = "line1\nline2\nline3\n"; - const aboveDiff = splitHashlineInput(`¶a.ts\n2-2:\n${repl("")}\n${repeat("2")}\n`).diff; + const aboveDiff = splitHashlineInput(`¶a.ts\n2 2\n${repl("")}\n${repeat("2")}\n`).diff; expect(applyDiff(text, aboveDiff)).toBe("line1\n\nline2\nline3\n"); - const belowDiff = splitHashlineInput(`¶a.ts\n2-2:\n${repeat("2")}\n${repl("")}\n`).diff; + const belowDiff = splitHashlineInput(`¶a.ts\n2 2\n${repeat("2")}\n${repl("")}\n`).diff; expect(applyDiff(text, belowDiff)).toBe("line1\nline2\n\nline3\n"); }); }); @@ -1307,28 +1176,28 @@ describe("hashline parser — delete and empty-block semantics", () => { describe("hashline parser — explicit blank payload rows", () => { it("raw blank lines between ops are ignored", () => { const text = "a\nb\nc\nd\ne\n"; - const ops = `¶a.ts\n1-1:\n${repl("A")}\n\n3-3:\n${repl("C")}\n`; + const ops = `¶a.ts\n1 1\n${repl("A")}\n\n3 3\n${repl("C")}\n`; const { diff } = splitHashlineInput(ops); expect(applyDiff(text, diff)).toBe("A\nb\nC\nd\ne\n"); }); it("empty replace payload rows are appended as blank payload lines", () => { const text = "a\nb\nc\nd\ne\n"; - const ops = `¶a.ts\n1-1:\n${repl("A")}\n${repl("")}\n${repl("")}\n3-3:\n${repl("C")}\n`; + const ops = `¶a.ts\n1 1\n${repl("A")}\n${repl("")}\n${repl("")}\n3 3\n${repl("C")}\n`; const { diff } = splitHashlineInput(ops); expect(applyDiff(text, diff)).toBe("A\n\n\nb\nC\nd\ne\n"); }); - it("`A-A:` followed by two empty replace rows replaces the line with two blanks", () => { + it("`A..A=` followed by two empty replace rows replaces the line with two blanks", () => { const text = "a\nb\nc\nd\ne\n"; - const ops = `¶a.ts\n2-2:\n${repl("")}\n${repl("")}\n4-4:\n${repl("D")}\n`; + const ops = `¶a.ts\n2 2\n${repl("")}\n${repl("")}\n4 4\n${repl("D")}\n`; const { diff } = splitHashlineInput(ops); expect(applyDiff(text, diff)).toBe("a\n\n\nc\nD\ne\n"); }); it("empty replace row inside payload between two content lines is preserved", () => { const text = "a\nb\nc\n"; - const ops = `¶a.ts\n2-2:\n${repl("first")}\n${repl("")}\n${repl("second")}\n`; + const ops = `¶a.ts\n2 2\n${repl("first")}\n${repl("")}\n${repl("second")}\n`; const { diff } = splitHashlineInput(ops); expect(applyDiff(text, diff)).toBe("a\nfirst\n\nsecond\nc\n"); }); diff --git a/packages/coding-agent/test/edit-diff.test.ts b/packages/coding-agent/test/edit-diff.test.ts index 89f47c207..803b931aa 100644 --- a/packages/coding-agent/test/edit-diff.test.ts +++ b/packages/coding-agent/test/edit-diff.test.ts @@ -236,12 +236,12 @@ describe("computeHashlineDiff", () => { const line = "unchanged content"; await Bun.write(sourcePath, `${line}\n`); - // `1-1:` with the same line in the replace bucket is a true no-op: the edit + // `1 1` with the same line in the body is a true no-op: the edit // fires through computeHashlineDiff but produces identical content. const text = `${line}\n`; const snapshotStore = new InMemorySnapshotStore(); const tag = snapshotStore.recordContiguous(sourcePath, 1, text.split("\n"), { fullText: text }); - const input = `${formatHashlineHeader(sourcePath, tag)}\n1-1:\n|${line}\n`; + const input = `${formatHashlineHeader(sourcePath, tag)}\n1 1\n+${line}\n`; const result = await computeHashlineDiff({ input }, tempDir, snapshotStore); expect("error" in result).toBe(true); if ("error" in result) { @@ -254,7 +254,7 @@ describe("computeHashlineDiff", () => { await Bun.write(sourcePath, "first\n"); const result = await computeHashlineDiff( - { input: `¶${sourcePath}\nEOF:\n|second` }, + { input: `¶${sourcePath}\nEOF\n+second` }, tempDir, new InMemorySnapshotStore(), ); @@ -265,7 +265,7 @@ describe("computeHashlineDiff", () => { }); test("returns a handled error when the source path is a local URL", async () => { const result = await computeHashlineDiff( - { input: "¶local://PLAN.md\nEOF:\n|x" }, + { input: "¶local://PLAN.md\nEOF\n+x" }, tempDir, new InMemorySnapshotStore(), ); diff --git a/packages/coding-agent/test/edit-streaming-preview.test.ts b/packages/coding-agent/test/edit-streaming-preview.test.ts index fb6dfbd8f..19826d8b1 100644 --- a/packages/coding-agent/test/edit-streaming-preview.test.ts +++ b/packages/coding-agent/test/edit-streaming-preview.test.ts @@ -54,7 +54,7 @@ describe("hashline streaming preview (multi-section)", () => { const ctx = (cwd: string) => ({ cwd, signal: new AbortController().signal }); test("keeps section A's preview when section B's header just arrived", async () => { - const input = ["¶a.ts", "BOF:", "|// new", "¶b.ts"].join("\n"); + const input = ["¶a.ts", "BOF", "+// new", "¶b.ts"].join("\n"); const previews = await strategy.computeDiffPreview({ input } as never, ctx(tmpDir) as never); expect(previews).not.toBeNull(); expect(previews).toHaveLength(1); @@ -64,8 +64,8 @@ describe("hashline streaming preview (multi-section)", () => { }); test("ignores parse errors from the trailing in-progress section", async () => { - // `7:bad` has inline payload — the trailing section is still being typed. - const input = ["¶a.ts", "BOF:", "|// new", "¶b.ts", "7:bad"].join("\n"); + // `7:bad` has invalid payload — the trailing section is still being typed. + const input = ["¶a.ts", "BOF", "+// new", "¶b.ts", "7:bad"].join("\n"); const previews = await strategy.computeDiffPreview({ input } as never, ctx(tmpDir) as never); expect(previews).not.toBeNull(); expect(previews).toHaveLength(1); @@ -74,7 +74,7 @@ describe("hashline streaming preview (multi-section)", () => { }); test("renders both sections once each has at least one valid op", async () => { - const input = ["¶a.ts", "BOF:", "|// new a", "¶b.ts", "BOF:", "|// new b"].join("\n"); + const input = ["¶a.ts", "BOF", "+// new a", "¶b.ts", "BOF", "+// new b"].join("\n"); const previews = await strategy.computeDiffPreview({ input } as never, ctx(tmpDir) as never); expect(previews).toHaveLength(2); expect(previews?.map(p => p.path).sort()).toEqual(["a.ts", "b.ts"]); diff --git a/packages/coding-agent/test/tools/ast-edit.test.ts b/packages/coding-agent/test/tools/ast-edit.test.ts index ba2b26639..dcb9aacf9 100644 --- a/packages/coding-agent/test/tools/ast-edit.test.ts +++ b/packages/coding-agent/test/tools/ast-edit.test.ts @@ -212,8 +212,8 @@ describe("ast_edit tool schema", () => { | undefined; // Tree-grouped output: `# packages/pkg-…/src/` then `## root.ts# (1 replacement)`. - expect(text).toMatch(/^## root\.ts#[0-9a-f]{4} \(\d+ replacement[s]?\)$/m); - expect(text).toMatch(/^## child\.ts#[0-9a-f]{4} \(\d+ replacement[s]?\)$/m); + expect(text).toMatch(/^## root\.ts#[0-9A-F]{3} \(\d+ replacement[s]?\)$/m); + expect(text).toMatch(/^## child\.ts#[0-9A-F]{3} \(\d+ replacement[s]?\)$/m); expect(text).not.toContain("ignore.js"); expect(text).not.toContain("outside.ts"); expect(details?.totalReplacements).toBe(2); diff --git a/packages/coding-agent/test/tools/ast-grep.test.ts b/packages/coding-agent/test/tools/ast-grep.test.ts index 47258ea73..c6ccaca3f 100644 --- a/packages/coding-agent/test/tools/ast-grep.test.ts +++ b/packages/coding-agent/test/tools/ast-grep.test.ts @@ -101,8 +101,8 @@ describe("ast_grep parse errors", () => { const details = result.details as { matchCount?: number; fileCount?: number } | undefined; // Directory mode uses tree-grouped `# dir/` + `## name#hash` headers. - expect(text).toMatch(/## root\.ts#[0-9a-f]+/); - expect(text).toMatch(/## child\.ts#[0-9a-f]+/); + expect(text).toMatch(/## root\.ts#[0-9A-F]{3}/); + expect(text).toMatch(/## child\.ts#[0-9A-F]{3}/); expect(text).not.toContain("ignore.js"); expect(text).not.toContain("outside.ts"); expect(details?.matchCount).toBe(2); diff --git a/packages/coding-agent/test/tools/eval-display-text.test.ts b/packages/coding-agent/test/tools/eval-display-text.test.ts index b442a25a9..f88f995ec 100644 --- a/packages/coding-agent/test/tools/eval-display-text.test.ts +++ b/packages/coding-agent/test/tools/eval-display-text.test.ts @@ -31,6 +31,15 @@ function baseResult(overrides: Record = {}) { }; } +const RED_1X1_PNG_BASE64 = + "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAIAAACQd1PeAAAADElEQVR4nGP4z8AAAAMBAQDJ/pLvAAAAAElFTkSuQmCC"; + +async function makeRedPng(width: number, height: number): Promise { + const seed = Buffer.from(RED_1X1_PNG_BASE64, "base64"); + const upscaled = await new Bun.Image(seed).resize(width, height, { filter: "nearest" }).png().bytes(); + return Buffer.from(upscaled).toString("base64"); +} + describe("EvalTool display() text surfacing", () => { afterEach(() => { vi.restoreAllMocks(); @@ -105,6 +114,40 @@ describe("EvalTool display() text surfacing", () => { expect(result.details?.images).toBeUndefined(); }); + it("downscales displayed images before returning ImageContent", async () => { + vi.spyOn(pyKernel, "checkPythonKernelAvailability").mockResolvedValue({ ok: true }); + const base64 = await makeRedPng(2400, 1200); + vi.spyOn(evalIndex.jsBackend, "execute").mockResolvedValue( + baseResult({ + displayOutputs: [{ type: "image", data: base64, mimeType: "image/png" }], + }) as never, + ); + + const tool = new EvalTool(makeSession()); + const result = await tool.execute("call-large-image", { + cells: [ + { + language: "js", + code: "```js\ndisplay({ type: 'image', data: largePng, mimeType: 'image/png' });\n```\n", + }, + ], + }); + + const image = result.content.find(c => c.type === "image"); + expect(image).toBeDefined(); + if (!image || image.type !== "image") throw new Error("Expected image content"); + expect(image.data).not.toBe(base64); + + const { width, height } = await new Bun.Image(Buffer.from(image.data, "base64")).metadata(); + expect(width).toBeLessThanOrEqual(1568); + expect(height).toBeLessThanOrEqual(1568); + + const text = result.content.map(c => (c.type === "text" ? c.text : "")).join("\n"); + expect(text).toContain("display image 1:"); + expect(text).toContain("original 2400x1200"); + expect(text).not.toContain(base64); + }); + it("still reports (no text output) when nothing was printed or displayed", async () => { vi.spyOn(pyKernel, "checkPythonKernelAvailability").mockResolvedValue({ ok: true }); vi.spyOn(evalIndex.jsBackend, "execute").mockResolvedValue(baseResult() as never); diff --git a/packages/coding-agent/test/tools/search-internal-urls.test.ts b/packages/coding-agent/test/tools/search-internal-urls.test.ts index 672f292f0..bf288edc5 100644 --- a/packages/coding-agent/test/tools/search-internal-urls.test.ts +++ b/packages/coding-agent/test/tools/search-internal-urls.test.ts @@ -147,7 +147,7 @@ describe("SearchTool internal URL resolution", () => { const text = getResultText(result); expect(text).toContain("needle"); // No hashline section headers or numbered editable lines for immutable sources. - expect(text).not.toMatch(/^¶.*#[0-9a-f]{4}$/m); + expect(text).not.toMatch(/^¶.*#[0-9A-F]{3}$/m); expect(text).not.toMatch(/^\*?\s*\d+:/m); }); @@ -187,7 +187,7 @@ describe("SearchTool internal URL resolution", () => { const text = getResultText(result); expect(text).toContain("needle"); // Mutable local:// sources keep a hashline section header plus numbered match lines. - expect(text).toMatch(/^¶.*#[0-9a-f]{4}$/m); + expect(text).toMatch(/^¶.*#[0-9A-F]{3}$/m); expect(text).toMatch(/^\*\d+:.*needle/m); }); @@ -207,7 +207,7 @@ describe("SearchTool internal URL resolution", () => { const text = getResultText(result); expect(text).toContain("needle"); // Mutable mixed.txt keeps hashlines somewhere in the output. - expect(text).toMatch(/^# mixed\.txt#[0-9a-f]{4}/m); + expect(text).toMatch(/^# mixed\.txt#[0-9A-F]{3}/m); expect(text).toMatch(/^\*\d+:.*mixed needle/m); }); diff --git a/packages/coding-agent/test/tools/search-path-lists.test.ts b/packages/coding-agent/test/tools/search-path-lists.test.ts index f1adebfd1..0408ea072 100644 --- a/packages/coding-agent/test/tools/search-path-lists.test.ts +++ b/packages/coding-agent/test/tools/search-path-lists.test.ts @@ -145,9 +145,9 @@ describe("tool path arrays", () => { const text = getText(result); const details = result.details as { fileCount?: number; scopePath?: string } | undefined; - expect(text).toMatch(/^# apps\/\n## grep\.txt#[0-9a-f]{4}/m); - expect(text).toMatch(/^# packages\/\n## grep\.txt#[0-9a-f]{4}/m); - expect(text).toMatch(/^# phases\/\n## grep\.txt#[0-9a-f]{4}/m); + expect(text).toMatch(/^# apps\/\n## grep\.txt#[0-9A-F]{3}/m); + expect(text).toMatch(/^# packages\/\n## grep\.txt#[0-9A-F]{3}/m); + expect(text).toMatch(/^# phases\/\n## grep\.txt#[0-9A-F]{3}/m); expect(text).toContain("shared-needle"); expect(text).not.toContain("# other"); expect(details?.fileCount).toBe(3); @@ -359,7 +359,7 @@ describe("tool path arrays", () => { const text = getText(result); const details = result.details as { fileCount?: number; scopePath?: string } | undefined; - expect(text).toMatch(/^# apps\/\n## grep\.txt#[0-9a-f]{4}/m); + expect(text).toMatch(/^# apps\/\n## grep\.txt#[0-9A-F]{3}/m); expect(text).toContain("shared-needle"); expect(text).not.toContain(tempDir); expect(details?.fileCount).toBe(1); @@ -416,9 +416,9 @@ describe("tool path arrays", () => { const text = getText(result); const details = result.details as { fileCount?: number; scopePath?: string } | undefined; - expect(text).toMatch(/^# apps\/\n## ast\.ts#[0-9a-f]{4}/m); - expect(text).toMatch(/^# packages\/\n## ast\.ts#[0-9a-f]{4}/m); - expect(text).toMatch(/^# phases\/\n## ast\.ts#[0-9a-f]{4}/m); + expect(text).toMatch(/^# apps\/\n## ast\.ts#[0-9A-F]{3}/m); + expect(text).toMatch(/^# packages\/\n## ast\.ts#[0-9A-F]{3}/m); + expect(text).toMatch(/^# phases\/\n## ast\.ts#[0-9A-F]{3}/m); expect(text).not.toContain("# other"); expect(details?.fileCount).toBe(3); expect(details?.scopePath).toBe("apps/**/*.ts, packages/**/*.ts, phases/**/*.ts"); @@ -444,9 +444,9 @@ describe("tool path arrays", () => { const text = getText(preview); const details = preview.details as { totalReplacements?: number; scopePath?: string } | undefined; - expect(text).toMatch(/^# apps\/\n## ast\.ts#[0-9a-f]{4} \(\d+ replacement/m); - expect(text).toMatch(/^# packages\/\n## ast\.ts#[0-9a-f]{4} \(\d+ replacement/m); - expect(text).toMatch(/^# phases\/\n## ast\.ts#[0-9a-f]{4} \(\d+ replacement/m); + expect(text).toMatch(/^# apps\/\n## ast\.ts#[0-9A-F]{3} \(\d+ replacement/m); + expect(text).toMatch(/^# packages\/\n## ast\.ts#[0-9A-F]{3} \(\d+ replacement/m); + expect(text).toMatch(/^# phases\/\n## ast\.ts#[0-9A-F]{3} \(\d+ replacement/m); expect(text).not.toContain("# other"); expect(details?.totalReplacements).toBe(3); expect(details?.scopePath).toBe("apps/**/*.ts, packages/**/*.ts, phases/**/*.ts"); @@ -556,9 +556,9 @@ describe("tool path arrays", () => { const text = getText(result); const details = result.details as { fileCount?: number; scopePath?: string } | undefined; - expect(text).toMatch(/^# apps\/\n## grep\.txt#[0-9a-f]{4}/m); - expect(text).toMatch(/^# packages\/\n## grep\.txt#[0-9a-f]{4}/m); - expect(text).toMatch(/^# phases\/\n## grep\.txt#[0-9a-f]{4}/m); + expect(text).toMatch(/^# apps\/\n## grep\.txt#[0-9A-F]{3}/m); + expect(text).toMatch(/^# packages\/\n## grep\.txt#[0-9A-F]{3}/m); + expect(text).toMatch(/^# phases\/\n## grep\.txt#[0-9A-F]{3}/m); expect(text).not.toContain("# other"); expect(details?.fileCount).toBe(3); expect(details?.scopePath).toBe("apps, packages, phases"); @@ -583,8 +583,8 @@ describe("tool path arrays", () => { const text = getText(result); const details = result.details as { fileCount?: number; scopePath?: string } | undefined; - expect(text).toMatch(/^# alpha\.txt#[0-9a-f]{4}/m); - expect(text).toMatch(/^# beta\.txt#[0-9a-f]{4}/m); + expect(text).toMatch(/^# alpha\.txt#[0-9A-F]{3}/m); + expect(text).toMatch(/^# beta\.txt#[0-9A-F]{3}/m); expect(text).toContain("exact-needle alpha"); expect(text).toContain("exact-needle beta"); expect(text).not.toContain("nested"); diff --git a/packages/hashline/CHANGELOG.md b/packages/hashline/CHANGELOG.md index b2e4fe8b5..0b66eb5cc 100644 --- a/packages/hashline/CHANGELOG.md +++ b/packages/hashline/CHANGELOG.md @@ -3,6 +3,13 @@ ## [Unreleased] ### Breaking Changes +- Changed hunk header syntax from `A-B:` to `@@ A..B @@` with `@@ A @@` shorthand for single lines +- Changed repeat payload sigil from `^A-B` to `&A..B` with `&A` shorthand for single lines +- Changed range separator from `-` to `..` in all contexts (anchors and repeats) +- Changed empty hunk behavior: concrete ranges now delete (no blank-line insertion); BOF/EOF empty hunks are now no-ops +- Removed `ApplyOptions` parameter from `applyEdits()` and related APIs; auto-absorb behavior is no longer configurable +- Removed diagnostic warnings for auto-absorbed duplicates from `ApplyResult`; warnings now come only from parser, patcher, or recovery +- Removed legacy hashline block syntax `A-B:`, `A-B:-`, and `^A-B` and replaced edits with `@@ A..B @@` hunks using `+` and `&` body rows - Removed `A:` shorthand syntax; use explicit `A-A:` for single-line anchors - Removed `↑` and `↓` payload sigils; use `|TEXT` for literal rows and `^A-B` for repeating original lines - Removed standalone delete rows; use inline `A-B:-` syntax instead @@ -13,12 +20,18 @@ ### Added +- Added compatibility parsing for apply_patch-style and unified-diff row noise by stripping path noise and converting context/delete body rows into hashline-compatible operations with warnings - Added `A-B:-` inline delete syntax for concrete range anchors - Added `^A-B` repeat payload syntax to emit original file lines inline - Added support for empty anchor blocks to write one blank line at the anchor position ### Changed +- Changed unified-diff compatibility mode to silently drop `-old` rows and convert context rows to `+TEXT` literals with a warning instead of rejecting them +- Changed `ABORT_MARKER` behavior to terminate parsing without surfacing a warning +- Changed numeric ranges to `A..B` form and accepted `@@ A @@` as shorthand for `@@ A..A @@` +- Changed empty hunk behavior so a concrete empty hunk deletes the selected range and `BOF`/`EOF` empty hunks no longer insert a blank line +- Changed parse behavior for `*** Abort` to stop processing without returning a speculative truncation warning - Changed payload row format from three sigils (`|`, `↑`, `↓`) to two (`|`, `^`) - Changed range anchor syntax to require explicit `A-B` form (no single-line shorthand) - Changed error messages to reference new syntax and remove references to removed sigils diff --git a/packages/hashline/README.md b/packages/hashline/README.md index c161f591d..b98cc7e3e 100644 --- a/packages/hashline/README.md +++ b/packages/hashline/README.md @@ -26,7 +26,7 @@ await fs.writeText("hello.ts", before); const tag = snapshots.recordContiguous("hello.ts", 1, before.split("\n"), { fullText: before }); const patcher = new Patcher({ fs, snapshots }); const patch = Patch.parse(String.raw`¶hello.ts#${tag} -1-1: +@@ 1..1 @@ +const greeting = "hello";`); const result = await patcher.apply(patch); @@ -39,19 +39,19 @@ console.log(await fs.readText("hello.ts")); See [`src/prompt.md`](./src/prompt.md) for the user-facing description and [`src/grammar.lark`](./src/grammar.lark) for the formal grammar. -Each hunk starts with a `¶PATH#TAG` header. The tag is a 3-hex opaque pointer -into the `SnapshotStore` that minted it; it is not content-derived and is not -meaningful outside that store. The patcher protects against stale anchors by -resolving the tag, verifying the recorded snapshot lines against live file -content, and refusing or attempting session-aware recovery on mismatch. +Each file section starts with `¶PATH#TAG`. The tag is a 3-hex opaque +pointer into the `SnapshotStore` that minted it; it is not content-derived +and is not meaningful outside that store. The patcher protects against +stale anchors by resolving the tag, verifying the recorded snapshot lines +against live file content, and refusing or attempting session-aware +recovery on mismatch. -Inside a hunk: -- `A-B:` — anchor lines A..B (use `A-A:` for a single line; no shorthand). -- `A-B:-` — delete lines A..B. -- `BOF:` / `EOF:` — virtual anchors at the beginning/end of file. +Inside a section: +- `@@ A..B @@` — open a hunk on lines A..B (use `@@ A,A @@` for a single line; bare `@@ A @@` is also accepted). +- `@@ BOF @@` / `@@ EOF @@` — virtual hunks at the beginning/end of file. - `+TEXT` — literal body row (use `+` alone for a blank line). -- `^A-B` — repeat original file lines A..B inline (`^A-A` for one line). -- Empty body — write one blank line at the anchor/virtual position. +- `&A..B` — repeat original file lines A..B inline (`&A` for one line). +- Empty body — delete the selected range. ## Abstractions diff --git a/packages/hashline/src/apply.ts b/packages/hashline/src/apply.ts index b80bb4cdb..e2ba809bc 100644 --- a/packages/hashline/src/apply.ts +++ b/packages/hashline/src/apply.ts @@ -1,20 +1,9 @@ /** * Apply a parsed list of {@link Edit}s to a text body and return the - * post-edit lines plus any diagnostic warnings. Pure function: no FS, no - * mutation of the input. - * - * The applier normalizes common model boundary mistakes: - * - * - Multi-line replacement-boundary duplicates are auto-absorbed (model - * echoed surrounding context as if it were payload). - * - Single-line structural-boundary duplicates (`}`, `)`, `];`, …) are - * auto-absorbed when delimiter balance suggests the range truncated short. - * - * Diagnostics are returned as `warnings[]` in {@link ApplyResult}; they do - * not abort the apply. + * post-edit lines. Pure function: no FS, no mutation of the input. */ import { cloneCursor } from "./tokenizer"; -import type { Anchor, ApplyOptions, ApplyResult, Cursor, Edit } from "./types"; +import type { Anchor, ApplyResult, Cursor, Edit } from "./types"; type LineOrigin = "original" | "insert" | "replacement"; @@ -27,14 +16,6 @@ interface IndexedEdit { idx: number; } -interface ReplacementGroup { - startIndex: number; - endIndex: number; - sourceLineNum: number; - replacement: string[]; - deletes: DeleteEdit[]; -} - function isReplacementInsert(edit: Edit): edit is InsertEdit & { mode: "replacement" } { return edit.kind === "insert" && edit.mode === "replacement"; } @@ -135,521 +116,6 @@ function insertAtEnd(fileLines: string[], lineOrigins: LineOrigin[], lines: stri return insertIndex + 1; } -/** Bucket edits by the line they target so we can apply each line's group in one splice. */ -function getAnchorTargetLine(edit: AppliedEdit): number | undefined { - if (edit.kind === "delete") return edit.anchor.line; - if (edit.cursor.kind === "before_anchor") return edit.cursor.anchor.line; - return undefined; -} - -function collectAnchorTargetLines(edits: AppliedEdit[]): Set { - const lines = new Set(); - for (const edit of edits) { - const line = getAnchorTargetLine(edit); - if (line !== undefined) lines.add(line); - } - return lines; -} - -function findReplacementGroup(edits: AppliedEdit[], startIndex: number): ReplacementGroup | undefined { - const first = edits[startIndex]; - if (!isReplacementInsert(first) || first.cursor.kind !== "before_anchor") return undefined; - - const sourceLineNum = first.lineNum; - const replacement: string[] = []; - let index = startIndex; - while (index < edits.length) { - const edit = edits[index]; - if (!isReplacementInsert(edit) || edit.lineNum !== sourceLineNum || edit.cursor.kind !== "before_anchor") break; - replacement.push(edit.text); - index++; - } - - const deletes: DeleteEdit[] = []; - while (index < edits.length) { - const edit = edits[index]; - if (edit.kind !== "delete" || edit.lineNum !== sourceLineNum) break; - deletes.push(edit); - index++; - } - if (deletes.length === 0) return undefined; - - const startLine = deletes[0].anchor.line; - for (let offset = 0; offset < deletes.length; offset++) { - if (deletes[offset].anchor.line !== startLine + offset) return undefined; - } - const cursorLine = first.cursor.anchor.line; - if (cursorLine !== startLine) return undefined; - - return { startIndex, endIndex: index - 1, sourceLineNum, replacement, deletes }; -} - -function countMatchingPrefixBlock(fileLines: string[], startLine: number, replacement: string[]): number { - const max = Math.min(replacement.length, startLine - 1); - for (let count = max; count >= 2; count--) { - let matches = true; - for (let offset = 0; offset < count; offset++) { - if (fileLines[startLine - count - 1 + offset] !== replacement[offset]) { - matches = false; - break; - } - } - if (matches) return count; - } - return 0; -} - -function countMatchingSuffixBlock(fileLines: string[], endLine: number, replacement: string[]): number { - const max = Math.min(replacement.length, fileLines.length - endLine); - for (let count = max; count >= 2; count--) { - let matches = true; - for (let offset = 0; offset < count; offset++) { - if (fileLines[endLine + offset] !== replacement[replacement.length - count + offset]) { - matches = false; - break; - } - } - if (matches) return count; - } - return 0; -} - -// Single-line replacement-boundary absorption is limited to structural closing -// delimiters. General one-line context is too easy to delete incorrectly, but -// duplicated `};` / `)` / `]` boundaries often mean a replacement range stopped -// one line early and would otherwise produce a syntax error. -const STRUCTURAL_CLOSING_BOUNDARY_RE = /^\s*[\])}]+[;,]?\s*$/; - -function isStructuralClosingBoundaryLine(line: string): boolean { - return STRUCTURAL_CLOSING_BOUNDARY_RE.test(line); -} - -interface DelimiterBalance { - paren: number; - bracket: number; - brace: number; -} - -/** - * Naive bracket counter — does NOT skip string/template/comment contents. The - * single-line structural absorb relies on this being safe-by-asymmetry: the - * candidate boundary line is constrained by `STRUCTURAL_CLOSING_BOUNDARY_RE` - * to be pure delimiters, so noise in deleted lines or non-boundary kept - * payload tends to push `expected !== kept` and biases the heuristic toward - * NOT absorbing (the safe direction). If we ever extend this to opening - * boundaries or non-structural single lines, swap this for a real tokenizer. - */ -function computeDelimiterBalance(lines: string[]): DelimiterBalance { - const balance: DelimiterBalance = { paren: 0, bracket: 0, brace: 0 }; - for (const line of lines) { - for (const char of line) { - switch (char) { - case "(": - balance.paren++; - break; - case ")": - balance.paren--; - break; - case "[": - balance.bracket++; - break; - case "]": - balance.bracket--; - break; - case "{": - balance.brace++; - break; - case "}": - balance.brace--; - break; - } - } - } - return balance; -} - -function delimiterBalancesEqual(a: DelimiterBalance, b: DelimiterBalance): boolean { - return a.paren === b.paren && a.bracket === b.bracket && a.brace === b.brace; -} - -/** - * Decides whether the structural-boundary candidate should be dropped: the - * `keptPayload` (full payload with the boundary line removed) must restore the - * caller's `expectedBalance`, while the `fullPayload` (boundary line still - * present) must NOT. For replacements `expectedBalance` is the deleted - * region's net delimiter balance; for pure inserts it is zero. - */ -function shouldDropSingleStructuralBoundary( - fullPayload: string[], - keptPayload: string[], - expectedBalance: DelimiterBalance, -): boolean { - return ( - delimiterBalancesEqual(computeDelimiterBalance(keptPayload), expectedBalance) && - !delimiterBalancesEqual(computeDelimiterBalance(fullPayload), expectedBalance) - ); -} - -function countMatchingSingleStructuralPrefixBoundary( - fileLines: string[], - startLine: number, - replacement: string[], - expectedBalance: DelimiterBalance, -): number { - if (replacement.length === 0 || startLine <= 1) return 0; - const line = replacement[0]; - if (!isStructuralClosingBoundaryLine(line)) return 0; - if (fileLines[startLine - 2] !== line) return 0; - return shouldDropSingleStructuralBoundary(replacement, replacement.slice(1), expectedBalance) ? 1 : 0; -} - -function countMatchingSingleStructuralSuffixBoundary( - fileLines: string[], - endLine: number, - replacement: string[], - expectedBalance: DelimiterBalance, -): number { - if (replacement.length === 0 || endLine >= fileLines.length) return 0; - const line = replacement[replacement.length - 1]; - if (!isStructuralClosingBoundaryLine(line)) return 0; - if (fileLines[endLine] !== line) return 0; - return shouldDropSingleStructuralBoundary(replacement, replacement.slice(0, -1), expectedBalance) ? 1 : 0; -} - -/** - * Single-line non-structural boundary duplicate detector for replacement - * groups. Mirrors the same boundary check the pure-insert absorber uses for - * leading/trailing context echoes, but applied to the top/bottom edges of an - * `A-B:` replacement payload. Catches mistakes like - * `103-138:` + `|const X = …` where line 102 already reads `const X = …`. - * - * Gated by `options.autoDropPureInsertDuplicates`: the existing 2+-line block - * absorb already runs unconditionally, and the structural single-line - * absorber is balance-validated; a non-structural single-line duplicate is - * ambiguous (could be an intentional `2-2:foo` over a line that happens to - * sit next to another `foo`), so we only fire when the user has opted in. - */ -function countMatchingSingleNonStructuralPrefixDuplicate( - fileLines: string[], - startLine: number, - replacement: string[], -): number { - if (replacement.length === 0 || startLine <= 1) return 0; - const line = replacement[0]; - if (line.trim().length === 0) return 0; - if (isStructuralClosingBoundaryLine(line)) return 0; - if (fileLines[startLine - 2] !== line) return 0; - return 1; -} - -function countMatchingSingleNonStructuralSuffixDuplicate( - fileLines: string[], - endLine: number, - replacement: string[], -): number { - if (replacement.length === 0 || endLine >= fileLines.length) return 0; - const line = replacement[replacement.length - 1]; - if (line.trim().length === 0) return 0; - if (isStructuralClosingBoundaryLine(line)) return 0; - if (fileLines[endLine] !== line) return 0; - return 1; -} - -function hasExternalTargets(lines: Iterable, externalTargetLines: Set): boolean { - for (const line of lines) { - if (externalTargetLines.has(line)) return true; - } - return false; -} - -function contiguousRange(start: number, count: number): number[] { - return Array.from({ length: count }, (_, offset) => start + offset); -} - -function deleteEditForAutoAbsorbedLine(line: number, sourceLineNum: number, index: number): AppliedEdit { - return { - kind: "delete", - anchor: { line }, - lineNum: sourceLineNum, - index, - }; -} - -interface PureInsertGroup { - startIndex: number; - endIndex: number; - sourceLineNum: number; - cursor: Cursor; - payload: string[]; -} - -function cursorMatches(a: Cursor, b: Cursor): boolean { - if (a.kind !== b.kind) return false; - if (a.kind === "bof" || a.kind === "eof") return true; - if (b.kind === "bof" || b.kind === "eof") return false; - return a.anchor.line === b.anchor.line; -} - -/** - * Collects a run of consecutive `insert` edits that all share the same - * `lineNum` and `cursor`, IFF that run is not immediately followed by a - * `delete` at the same `lineNum` (which would make it a replacement group - * instead). Returns the contiguous payload so we can check it for boundary - * duplicates against the file. - */ -function findPureInsertGroup(edits: AppliedEdit[], startIndex: number): PureInsertGroup | undefined { - const first = edits[startIndex]; - if (first?.kind !== "insert" || isReplacementInsert(first)) return undefined; - - const sourceLineNum = first.lineNum; - const cursor = first.cursor; - const payload: string[] = []; - let index = startIndex; - while (index < edits.length) { - const edit = edits[index]; - if (edit.kind !== "insert" || isReplacementInsert(edit) || edit.lineNum !== sourceLineNum) break; - if (!cursorMatches(edit.cursor, cursor)) break; - payload.push(edit.text); - index++; - } - - // If the run is followed by a delete at the same source lineNum, this is a - // replacement group (handled by absorbReplacement…). Decline. - if (index < edits.length && edits[index].kind === "delete" && edits[index].lineNum === sourceLineNum) { - return undefined; - } - - return { startIndex, endIndex: index - 1, sourceLineNum, cursor, payload }; -} - -/** - * For a pure-insert group, locate the file region adjacent to the insertion - * point. Returns 0-indexed bounds: - * - `aboveEndIdx`: index of the last file line strictly above the insertion - * point (-1 if none). - * - `belowStartIdx`: index of the first file line strictly below the - * insertion point (`fileLines.length` if none). - */ -function pureInsertNeighborhood(cursor: Cursor, fileLines: string[]): { aboveEndIdx: number; belowStartIdx: number } { - if (cursor.kind === "bof") return { aboveEndIdx: -1, belowStartIdx: 0 }; - if (cursor.kind === "eof") return { aboveEndIdx: fileLines.length - 1, belowStartIdx: fileLines.length }; - return { aboveEndIdx: cursor.anchor.line - 2, belowStartIdx: cursor.anchor.line - 1 }; -} - -interface PureInsertAbsorbResult { - keptPayload: string[]; - absorbedLeading: number; - absorbedTrailing: number; - leadingFileRange?: { start: number; end: number }; // 1-indexed inclusive - trailingFileRange?: { start: number; end: number }; // 1-indexed inclusive -} - -/** - * For a pure-insert group, drop only multi-line context echoes that exactly - * duplicate the file lines adjacent to the insertion point. Single-line pure - * insert duplicates are ambiguous (a repeated `}` may be an accidental anchor - * echo or an intentional inserted delimiter), so they are left literal even when - * duplicate absorption is enabled. - */ -function tryAbsorbPureInsertGroup( - group: PureInsertGroup, - fileLines: string[], - allowGenericBoundaryAbsorb: boolean, -): PureInsertAbsorbResult { - const empty: PureInsertAbsorbResult = { keptPayload: group.payload, absorbedLeading: 0, absorbedTrailing: 0 }; - if (group.payload.length === 0) return empty; - - const { aboveEndIdx, belowStartIdx } = pureInsertNeighborhood(group.cursor, fileLines); - - // Leading: payload[0..k-1] vs fileLines[aboveEndIdx-k+1 .. aboveEndIdx]. - let absorbedLeading = 0; - if (allowGenericBoundaryAbsorb) { - const maxLead = Math.min(group.payload.length, aboveEndIdx + 1); - for (let count = maxLead; count >= 2; count--) { - let ok = true; - for (let offset = 0; offset < count; offset++) { - if (group.payload[offset] !== fileLines[aboveEndIdx - count + 1 + offset]) { - ok = false; - break; - } - } - if (ok) { - absorbedLeading = count; - break; - } - } - } - - // Trailing: payload[len-k..len-1] vs fileLines[belowStartIdx..belowStartIdx+k-1]. - // Don't double-count payload lines already absorbed as leading. - let absorbedTrailing = 0; - const remaining = group.payload.length - absorbedLeading; - if (allowGenericBoundaryAbsorb) { - const maxTrail = Math.min(remaining, fileLines.length - belowStartIdx); - for (let count = maxTrail; count >= 2; count--) { - let ok = true; - for (let offset = 0; offset < count; offset++) { - if (group.payload[group.payload.length - count + offset] !== fileLines[belowStartIdx + offset]) { - ok = false; - break; - } - } - if (ok) { - absorbedTrailing = count; - break; - } - } - } - - if (absorbedLeading === 0 && absorbedTrailing === 0) return empty; - - return { - keptPayload: group.payload.slice(absorbedLeading, group.payload.length - absorbedTrailing), - absorbedLeading, - absorbedTrailing, - leadingFileRange: - absorbedLeading > 0 ? { start: aboveEndIdx - absorbedLeading + 2, end: aboveEndIdx + 1 } : undefined, - trailingFileRange: - absorbedTrailing > 0 ? { start: belowStartIdx + 1, end: belowStartIdx + absorbedTrailing } : undefined, - }; -} - -function absorbReplacementBoundaryDuplicates( - edits: AppliedEdit[], - fileLines: string[], - warnings: string[], - options: ApplyOptions, -): AppliedEdit[] { - let nextSyntheticIndex = edits.length; - const absorbed: AppliedEdit[] = []; - - // Anchor targets are stable across the loop because we only ever append - // synthetic deletes (never mutate originals). A line in this set that - // falls outside the current group's range is necessarily owned by another - // op, so absorbing it would silently steal its target. - const allTargetLines = collectAnchorTargetLines(edits); - const emittedAbsorbKeys = new Set(); - - for (let index = 0; index < edits.length; index++) { - const group = findReplacementGroup(edits, index); - if (!group) { - const pureInsert = findPureInsertGroup(edits, index); - if (pureInsert) { - const result = tryAbsorbPureInsertGroup( - pureInsert, - fileLines, - options.autoDropPureInsertDuplicates === true, - ); - if (result.absorbedLeading > 0 || result.absorbedTrailing > 0) { - if (result.leadingFileRange) { - const { start, end } = result.leadingFileRange; - const key = `pure-insert-leading:${start}..${end}`; - if (!emittedAbsorbKeys.has(key)) { - emittedAbsorbKeys.add(key); - warnings.push( - `Auto-dropped ${result.absorbedLeading} duplicate line(s) at the start of insert at line ${pureInsert.sourceLineNum} ` + - `(file lines ${start}..${end} already match the payload's leading lines).`, - ); - } - } - if (result.trailingFileRange) { - const { start, end } = result.trailingFileRange; - const key = `pure-insert-trailing:${start}..${end}`; - if (!emittedAbsorbKeys.has(key)) { - emittedAbsorbKeys.add(key); - warnings.push( - `Auto-dropped ${result.absorbedTrailing} duplicate line(s) at the end of insert at line ${pureInsert.sourceLineNum} ` + - `(file lines ${start}..${end} already match the payload's trailing lines).`, - ); - } - } - for (const text of result.keptPayload) { - absorbed.push({ - kind: "insert", - cursor: cloneCursor(pureInsert.cursor), - text, - lineNum: pureInsert.sourceLineNum, - index: nextSyntheticIndex++, - }); - } - index = pureInsert.endIndex; - continue; - } - for (let groupIndex = pureInsert.startIndex; groupIndex <= pureInsert.endIndex; groupIndex++) { - absorbed.push(edits[groupIndex]); - } - index = pureInsert.endIndex; - continue; - } - absorbed.push(edits[index]); - continue; - } - - const startLine = group.deletes[0].anchor.line; - const endLine = group.deletes[group.deletes.length - 1].anchor.line; - - const deletedBalance = computeDelimiterBalance( - group.deletes.map(deleteEdit => fileLines[deleteEdit.anchor.line - 1] ?? ""), - ); - const optInSingleLineAbsorb = options.autoDropPureInsertDuplicates === true; - const prefixCount = - countMatchingPrefixBlock(fileLines, startLine, group.replacement) || - countMatchingSingleStructuralPrefixBoundary(fileLines, startLine, group.replacement, deletedBalance) || - (optInSingleLineAbsorb - ? countMatchingSingleNonStructuralPrefixDuplicate(fileLines, startLine, group.replacement) - : 0); - const suffixCount = - countMatchingSuffixBlock(fileLines, endLine, group.replacement) || - countMatchingSingleStructuralSuffixBoundary(fileLines, endLine, group.replacement, deletedBalance) || - (optInSingleLineAbsorb - ? countMatchingSingleNonStructuralSuffixDuplicate(fileLines, endLine, group.replacement) - : 0); - const prefixLines = contiguousRange(startLine - prefixCount, prefixCount); - const suffixLines = contiguousRange(endLine + 1, suffixCount); - const safePrefixCount = hasExternalTargets(prefixLines, allTargetLines) ? 0 : prefixCount; - const safeSuffixCount = hasExternalTargets(suffixLines, allTargetLines) ? 0 : suffixCount; - - if (safePrefixCount > 0) { - const absorbStart = startLine - safePrefixCount; - const key = `prefix:${absorbStart}..${startLine - 1}`; - if (!emittedAbsorbKeys.has(key)) { - emittedAbsorbKeys.add(key); - warnings.push( - `Auto-absorbed ${safePrefixCount} duplicate line(s) above replacement ` + - `(file lines ${absorbStart}..${startLine - 1} matched the payload's leading lines; ` + - `widened the deletion to start at file line ${absorbStart} instead of ${startLine}).`, - ); - } - } - if (safeSuffixCount > 0) { - const absorbEnd = endLine + safeSuffixCount; - const key = `suffix:${endLine + 1}..${absorbEnd}`; - if (!emittedAbsorbKeys.has(key)) { - emittedAbsorbKeys.add(key); - warnings.push( - `Auto-absorbed ${safeSuffixCount} duplicate line(s) below replacement ` + - `(file lines ${endLine + 1}..${absorbEnd} matched the payload's trailing lines; ` + - `widened the deletion to end at file line ${absorbEnd} instead of ${endLine}).`, - ); - } - } - - for (const line of contiguousRange(startLine - safePrefixCount, safePrefixCount)) { - absorbed.push(deleteEditForAutoAbsorbedLine(line, group.sourceLineNum, nextSyntheticIndex++)); - } - for (let groupIndex = group.startIndex; groupIndex <= group.endIndex; groupIndex++) { - absorbed.push(edits[groupIndex]); - } - for (const line of contiguousRange(endLine + 1, safeSuffixCount)) { - absorbed.push(deleteEditForAutoAbsorbedLine(line, group.sourceLineNum, nextSyntheticIndex++)); - } - - index = group.endIndex; - } - - return absorbed; -} - function bucketAnchorEditsByLine(edits: IndexedEdit[]): Map { const byLine = new Map(); for (const entry of edits) { @@ -669,26 +135,22 @@ function bucketAnchorEditsByLine(edits: IndexedEdit[]): Map "original"); - const warnings: string[] = []; let firstChangedLine: number | undefined; const trackFirstChanged = (line: number) => { if (firstChangedLine === undefined || line < firstChangedLine) firstChangedLine = line; }; - const expandedEdits = expandRepeatEdits(edits, fileLines); - validateLineBounds(expandedEdits, fileLines); - - const targetEdits = absorbReplacementBoundaryDuplicates(expandedEdits, fileLines, warnings, options); + const targetEdits = expandRepeatEdits(edits, fileLines); + validateLineBounds(targetEdits, fileLines); // Partition edits into BOF, EOF, and anchor-targeted buckets. const bofLines: string[] = []; @@ -751,6 +213,5 @@ export function applyEdits(text: string, edits: Edit[], options: ApplyOptions = return { text: fileLines.join("\n"), firstChangedLine, - ...(warnings.length > 0 ? { warnings } : {}), }; } diff --git a/packages/hashline/src/format.ts b/packages/hashline/src/format.ts index 0dabf6350..6c2077191 100644 --- a/packages/hashline/src/format.ts +++ b/packages/hashline/src/format.ts @@ -4,26 +4,23 @@ * tokenizer, the prompt, and the formal grammar. */ -/** Anchor terminator for every hashline operation block. */ -export const HL_OP_REPLACE = ":"; - -/** Inline-delete suffix for concrete range anchors (`A-B:-`). */ -export const HL_OP_DELETE_SUFFIX = ":-"; +/** File-section header prefix: `¶path#hash`. */ +export const HL_FILE_PREFIX = "¶"; /** Payload sigil for literal body rows. */ export const HL_PAYLOAD_REPLACE = "+"; /** Payload sigil for body rows that repeat original file lines. */ -export const HL_PAYLOAD_REPEAT = "^"; +export const HL_PAYLOAD_REPEAT = "&"; /** All hashline payload sigils, concatenated for fast membership tests. */ export const HL_PAYLOAD_CHARS = `${HL_PAYLOAD_REPLACE}${HL_PAYLOAD_REPEAT}`; -/** Hashline edit file-section header marker. */ -export const HL_FILE_PREFIX = "¶"; - /** Separator between a hashline file path and its opaque snapshot tag. */ export const HL_FILE_HASH_SEP = "#"; +/** Separator between two line numbers in a range, e.g. `5..10`. */ +export const HL_RANGE_SEP = ".."; + /** Separator between a line number and displayed line content in hashline mode. */ export const HL_LINE_BODY_SEP = ":"; @@ -33,12 +30,12 @@ function regexEscape(str: string): string { /** * Decoration prefix that may precede a line number in tool output: - * `>` (context line in grep), `-` (removed line), `*` (match line). - * Any combination, in any order, surrounded by optional whitespace. Output - * formatters emit at most one decoration per line; the parser stays liberal - * because it accepts whatever the model echoes back. + * `*` (match line), `>` (context line in grep). Any combination, in any + * order, surrounded by optional whitespace. Output formatters emit at most + * one decoration per line; the parser stays liberal because it accepts + * whatever the model echoes back. */ -export const HL_ANCHOR_DECORATION_RE_RAW = `\\s*[>\\-*]*\\s*`; +export const HL_ANCHOR_DECORATION_RE_RAW = `\\s*[>*]*\\s*`; /** Capture-group regex source for a decorated bare line-number anchor. */ export const HL_ANCHOR_RE_RAW = `${HL_ANCHOR_DECORATION_RE_RAW}(\\d+)`; @@ -49,9 +46,9 @@ export const HL_LINE_RE_RAW = `[1-9]\\d*`; /** Capture-group form of {@link HL_LINE_RE_RAW}. */ export const HL_LINE_CAPTURE_RE_RAW = `(${HL_LINE_RE_RAW})`; -/** Regex for repeat payload rows (`^A-B`). */ +/** Regex for repeat payload rows (`&A..B`). */ export const HL_PAYLOAD_REPEAT_RE = new RegExp( - `^\\${HL_PAYLOAD_REPEAT}${HL_LINE_CAPTURE_RE_RAW}-${HL_LINE_CAPTURE_RE_RAW}$`, + `^\\${HL_PAYLOAD_REPEAT}${HL_LINE_CAPTURE_RE_RAW},${HL_LINE_CAPTURE_RE_RAW}$`, ); /** Number of hex characters in an opaque snapshot tag. */ diff --git a/packages/hashline/src/grammar.lark b/packages/hashline/src/grammar.lark index 03881ee47..b32f79ef1 100644 --- a/packages/hashline/src/grammar.lark +++ b/packages/hashline/src/grammar.lark @@ -1,22 +1,22 @@ -start: begin_patch hunk+ end_patch +start: begin_patch file_patch+ end_patch begin_patch: "*** Begin Patch" LF -end_patch: "*** End Patch" LF? +end_patch: "*** End Patch" LF? -hunk: update_hunk -update_hunk: "¶" filename ("#" file_hash)? LF block* - -filename: /([^\s#]+)/ +file_patch: file_header hunk+ +file_header: "¶" filename ("#" file_hash)? LF file_hash: /[0-9A-F]{3}/ +filename: /[^\s#]+/ -block: anchor ":" delete_suffix? LF payload* -delete_suffix: "-" -payload: literal_payload | repeat_payload -literal_payload: "+" /[^\n]*/ LF -repeat_payload: "^" range LF - -anchor: range | "BOF" | "EOF" -range: LID "-" LID +hunk: hunk_header op* +hunk_header: anchor LF +op: emit_op | repeat_op +emit_op: "+" /(.*)/ LF +repeat_op: "&" body_range LF +anchor: header_range | "BOF" | "EOF" +header_range: LID (WS LID)? +body_range: LID (".." LID)? LID: /[1-9]\d*/ +WS: /[ \t]+/ %import common.LF diff --git a/packages/hashline/src/input.ts b/packages/hashline/src/input.ts index 6a393993f..e0c2016d6 100644 --- a/packages/hashline/src/input.ts +++ b/packages/hashline/src/input.ts @@ -12,7 +12,7 @@ import { applyEdits } from "./apply"; import { HL_FILE_HASH_SEP, HL_FILE_PREFIX } from "./format"; import { parsePatch, parsePatchStreaming } from "./parser"; import { Tokenizer } from "./tokenizer"; -import type { ApplyOptions, ApplyResult, Edit, SplitOptions } from "./types"; +import type { ApplyResult, Edit, SplitOptions } from "./types"; // Pure classification — single shared tokenizer is safe. const TOKENIZER = new Tokenizer(); @@ -25,8 +25,46 @@ function unquoteHashlinePath(pathText: string): string { return pathText; } +/** + * Strip apply_patch-style noise that models reflexively prepend to the + * path. Examples observed in benchmark traces: + * + * `Update File:foo.ts`, `Update:foo.ts`, `UpdateFile:foo.ts`, + * `Update/File:foo.ts`, `Update-file:foo.ts`, `Update(File):foo.ts`, + * `Update]*(File|to)?[]*:` + * keyword block, case-insensitive. The remaining text is the real path. + */ +const APPLY_PATCH_PATH_NOISE_RE = + /^\*{0,3}\s*(?:(?:update|add|delete|move)[^A-Za-z0-9]*(?:file|to)?[^A-Za-z0-9]*:)?\s*\*{0,3}\s*/i; + +function stripApplyPatchPathNoise(pathText: string): string { + return pathText.replace(APPLY_PATCH_PATH_NOISE_RE, ""); +} + +/** + * Best-effort recovery for `¶`-prefixed lines the strict tokenizer + * rejects. Strips apply_patch keyword noise (`Update File:`, `Update:`, + * etc.) and an extra leading `***` (some models emit a hybrid `¶***foo.ts` + * shape), then expects `PATH(#HASH)?` with no embedded whitespace. + * Returns `null` when no clean path can be salvaged. + */ +function tryParseRecoveryHeader(line: string, cwd?: string): RawSection | null { + if (!line.startsWith(HL_FILE_PREFIX)) return null; + const body = stripApplyPatchPathNoise(line.slice(HL_FILE_PREFIX.length).trim()); + if (body.length === 0) return null; + const match = /^(\S+?)(?:#([0-9A-Fa-f]{3}))?\s*$/.exec(body); + if (match === null) return null; + const path = normalizeHashlinePath(match[1], cwd); + if (path.length === 0) return null; + return match[2] !== undefined ? { path, fileHash: match[2].toUpperCase(), diff: "" } : { path, diff: "" }; +} + function normalizeHashlinePath(rawPath: string, cwd?: string): string { - const unquoted = unquoteHashlinePath(rawPath.trim()); + const unquoted = stripApplyPatchPathNoise(unquoteHashlinePath(rawPath.trim())); if (!cwd || !path.isAbsolute(unquoted)) return unquoted; const relative = path.relative(path.resolve(cwd), path.resolve(unquoted)); const isWithinCwd = relative === "" || (!relative.startsWith("..") && !path.isAbsolute(relative)); @@ -40,9 +78,9 @@ interface RawSection { } /** - * Parse a `¶PATH[#hash]` header line. Returns `null` for lines that do not - * begin with the `¶` prefix; throws the existing "Input header must be …" - * error when a `¶`-prefixed line fails the strict shape (so malformed paths + * Parse a `¶PATH[#hash]` header line. Returns `null` for lines that do + * not start with `¶`. Throws the strict "Input header must be …" error + * when a `¶`-prefixed line fails the strict shape (so malformed paths * surface immediately instead of being silently re-classified as payload). */ function parseHashlineHeaderLine(line: string, cwd?: string): RawSection | null { @@ -51,6 +89,11 @@ function parseHashlineHeaderLine(line: string, cwd?: string): RawSection | null const token = TOKENIZER.tokenize(trimmed); if (token.kind !== "header") { + // Recovery: try to extract a path from the raw line after stripping + // apply_patch noise. This handles `*** Update File:foo.ts#CB5` and + // the half-dozen variants models actually emit. + const recovered = tryParseRecoveryHeader(trimmed, cwd); + if (recovered !== null) return recovered; throw new Error( `Input header must be ${HL_FILE_PREFIX}PATH or ${HL_FILE_PREFIX}PATH${HL_FILE_HASH_SEP}TAG with a 3-hex snapshot tag; got ${JSON.stringify(trimmed)}.`, ); @@ -110,6 +153,15 @@ function splitRawSections(input: string, options: SplitOptions = {}): RawSection const firstLine = lines[0] ?? ""; if (parseHashlineHeaderLine(firstLine, options.cwd) === null) { + // Catch unified-diff hunk-header contamination on the first line so + // the model sees a focused error. + const firstTrimmed = firstLine.trimEnd(); + if (/^@@\s+[-+]?\d+,\d+\s+[-+]?\d+,\d+\s+@@/.test(firstTrimmed)) { + throw new Error( + "unified-diff hunk header (`@@ -N,M +N,M @@`) is not valid in hashline. " + + "File sections start with `¶path#HASH`; hunks are bare `A B` lines.", + ); + } const preview = JSON.stringify(firstLine.slice(0, 120)); throw new Error( `input must begin with "${HL_FILE_PREFIX}PATH${HL_FILE_HASH_SEP}HASH" on the first non-blank line for anchored edits; got: ${preview}. ` + @@ -228,11 +280,11 @@ export class PatchSection { * method directly when you've already validated the file content and * just want the result. */ - applyTo(text: string, options: ApplyOptions = {}): ApplyResult { + applyTo(text: string): ApplyResult { const { edits, warnings } = this.parse(); - const result = applyEdits(text, [...edits], options); - // Preserve parse warnings alongside applier warnings so consumers - // don't need to call `parse()` separately. + const result = applyEdits(text, [...edits]); + // Preserve parse warnings so consumers don't need to call `parse()` + // separately. const merged = warnings.length === 0 ? result.warnings : [...warnings, ...(result.warnings ?? [])]; return merged && merged.length > 0 ? { ...result, warnings: merged } @@ -246,9 +298,9 @@ export class PatchSection { * empty-payload edit. Intended for incremental diff previews; the writer * path should always use {@link applyTo}. */ - applyPartialTo(text: string, options: ApplyOptions = {}): ApplyResult { + applyPartialTo(text: string): ApplyResult { const { edits, warnings } = parsePatchStreaming(this.diff); - const result = applyEdits(text, [...edits], options); + const result = applyEdits(text, [...edits]); const merged = warnings.length === 0 ? result.warnings : [...warnings, ...(result.warnings ?? [])]; return merged && merged.length > 0 ? { ...result, warnings: merged } diff --git a/packages/hashline/src/messages.ts b/packages/hashline/src/messages.ts index 4a9a915a5..d6cf0badf 100644 --- a/packages/hashline/src/messages.ts +++ b/packages/hashline/src/messages.ts @@ -17,79 +17,53 @@ export const END_PATCH_MARKER = "*** End Patch"; /** * Recovery sentinel emitted by an agent loop when a contaminated tool-call * stream is truncated mid-call. Behaves like {@link END_PATCH_MARKER} for - * parsing — terminates the line loop — and additionally surfaces a warning - * so the caller knows to re-issue any remaining edits. + * parsing — terminates the line loop — and does not surface a warning. */ export const ABORT_MARKER = "*** Abort"; -// `ABORT_MARKER` (`*** Abort`) still terminates parsing — see the `abort` -// token in `tokenizer.ts` — but no longer surfaces a warning to the caller. -// The earlier wording ("Tool stream truncated mid-call due to detected -// output corruption") was always speculative: by the time we observe the -// marker the stream is already gone, the warning could not actually -// describe the cause, and downstream consumers were not differentiating it -// from any other warning anyway. /** - * Warning text appended when two consecutive blocks target the exact same - * concrete range. The second block wins; the first block is discarded. + * Warning text appended when two consecutive hunks target the exact same + * concrete range. The second hunk wins; the first is discarded. */ export const REPLACE_PAIR_COALESCED_WARNING = - "Detected two identical-range hashline blocks; kept only the second block. Issue ONE block per range — payload is the final desired content, never both old and new."; + "Detected two identical-range hashline hunks; kept only the second hunk. Issue ONE hunk per range — payload is the final desired content, never both old and new."; /** - * Warning text appended when a bare anchor block (`A-B:` with no payload) - * is followed by an overlapping concrete block. The earlier bare block is - * dropped on the assumption that the model expressed an old/new pair - * across two anchors; only the second block's payload is applied. + * Warning text appended when a bare hunk header (`A B` with no body) + * is followed by an overlapping concrete hunk. The earlier bare hunk is + * dropped on the assumption that the model expressed an old/new pair across + * two hunks; only the second hunk's payload is applied. */ export const REPLACE_PAIR_COALESCED_OVERLAP_WARNING = - "Detected an overlapping bare hashline block immediately followed by a concrete block; dropped the earlier bare block. Issue ONE block per range — payload is the final desired content, never both old and new."; + "Detected an overlapping bare hashline hunk immediately followed by a concrete hunk; dropped the earlier bare hunk. Issue ONE hunk per range — payload is the final desired content, never both old and new."; /** - * Warning text appended when bare body rows (no `+` / `^` prefix) follow a - * concrete anchor and the parser auto-converts them to `+literal` rows - * because no `+`/`^` row was present in the block. Helps the model learn - * the canonical body-row syntax while keeping the patch applying. + * Warning text appended when bare body rows (no `+` / `&` prefix) follow a + * hunk header and the parser auto-converts them to `+literal` rows because + * no `+`/`&` row was present in the hunk. Helps the model learn the + * canonical body-row syntax while keeping the patch applying. */ export const BARE_BODY_AUTO_PIPED_WARNING = - "Auto-prefixed bare body row(s) with `+`. Always start payload rows with `+TEXT` (literal) or `^A-B` (repeat) — pasting raw code as payload is not a portable shape."; + "Auto-prefixed bare body row(s) with `+`. Always start payload rows with `+TEXT` (literal) or `&A..B` (repeat) — pasting raw code as payload is not a portable shape."; /** - * Warning text appended when a lone `-` body row is retroactively converted - * to a `:-` delete on the preceding bare anchor. Models occasionally write - * `A-B:` followed by a `-` row when they meant `A-B:-`. - */ -export const DASH_PAYLOAD_AUTO_DELETE_WARNING = - "Converted a lone `-` body row to a `:-` delete on the preceding anchor. Write `A-B:-` on the anchor line itself to delete the range."; - -/** - * Warning text appended when a single contiguous run of two or more - * single-line empty-body blocks (`A-A:` with no payload) is flushed. - * These commonly indicate the model thought `A-A:` deletes the line; it - * actually replaces with a blank line. Suggest `A-B:-` instead. - */ -export const STACKED_BLANK_REPLACE_WARNING = - "Detected a run of single-line empty-body blocks (`A-A:` with no payload). Each one REPLACES its line with a blank; to delete lines use `A-B:-`."; - -/** - * Warning text emitted when a body row begins with `+^A-B` — the model + * Warning text emitted when a body row begins with `+&A..B` — the model * mistakenly prefixed a repeat row with the `+` literal sigil. We reroute - * the row as a `^A-B` repeat so the patch still applies, then surface this + * the row as a `&A..B` repeat so the patch still applies, then surface this * warning so the model sees the mistake on the next turn. */ export const PLUS_PREFIXED_REPEAT_WARNING = - "A body row started with `+^A-B`. `+` (literal text) and `^A-B` (repeat) are sibling row kinds — a row uses exactly one of them. Treated as `^A-B`; remove the leading `+` next time."; + "A body row started with `+&A..B`. `+` (literal text) and `&A..B` (repeat) are sibling row kinds — a row uses exactly one of them. Treated as `&A..B`; remove the leading `+` next time."; -/** Error text prefix emitted when an anchor line carries inline payload. */ -export const INLINE_PAYLOAD_REJECTED_PREFIX = "Inline payload on the anchor line is rejected."; - -/** Error text emitted when inline delete targets BOF/EOF. */ -export const VIRTUAL_REPLACE_REJECTED_MESSAGE = - "BOF:/EOF: anchors are virtual positions and cannot use `:-`. Use `+TEXT` or `^A-B` body rows to insert at a virtual position."; - -/** Error text emitted when `^A` repeat shorthand is used. */ -export const REPEAT_SHORTHAND_REJECTED_MESSAGE = - "Repeat payload shorthand `^A` is rejected. Use explicit `^A-A` for one line."; +/** + * Warning text emitted when a hunk body contains unified-diff-style rows + * (`-old`, ` context`) and the parser silently converts them: `-` rows are + * dropped (the hunk header's range already deletes those lines), and the + * leading metadata-space on context rows is stripped once unified-diff + * mode is detected. Bare body rows are auto-prefixed with `+` regardless. + */ +export const UNIFIED_DIFF_BODY_AUTO_CONVERT_WARNING = + "Hunk body contained unified-diff-style rows (`-old`, ` context`). The `-` rows were dropped (the hunk header's range already deletes those lines); context rows were treated as `+TEXT` literals. Use `+TEXT` (literal) or `&A..B` (repeat) directly next time."; /** Warning text emitted by `Recovery` when an external write fits a cached snapshot. */ export const RECOVERY_EXTERNAL_WARNING = diff --git a/packages/hashline/src/parser.ts b/packages/hashline/src/parser.ts index b7fc3fbd6..adb25720e 100644 --- a/packages/hashline/src/parser.ts +++ b/packages/hashline/src/parser.ts @@ -5,44 +5,41 @@ * * Lifecycle: * - * 1. Construct one {@link Executor} per hunk (or share one with `reset()`). - * 2. Feed it tokens via {@link Executor.feed}. Block payload rows are - * accumulated across tokens until the next anchor block flushes them. - * 3. Call {@link Executor.end} to flush the trailing pending block and validate - * cross-block invariants (no overlapping deletes, etc.). + * 1. Construct one {@link Executor} per patch (or share one with `reset()`). + * 2. Feed it tokens via {@link Executor.feed}. Hunk body rows accumulate + * until the next hunk header or {@link end} flushes them. + * 3. Call {@link Executor.end} to flush the trailing pending hunk and + * validate cross-hunk invariants (no overlapping deletes, etc.). * * Convenience entry point: {@link parsePatch}. */ import { HL_PAYLOAD_REPEAT, HL_PAYLOAD_REPLACE } from "./format"; import { BARE_BODY_AUTO_PIPED_WARNING, - DASH_PAYLOAD_AUTO_DELETE_WARNING, - INLINE_PAYLOAD_REJECTED_PREFIX, PLUS_PREFIXED_REPEAT_WARNING, REPLACE_PAIR_COALESCED_OVERLAP_WARNING, REPLACE_PAIR_COALESCED_WARNING, - STACKED_BLANK_REPLACE_WARNING, - VIRTUAL_REPLACE_REJECTED_MESSAGE, + UNIFIED_DIFF_BODY_AUTO_CONVERT_WARNING, } from "./messages"; import { type BlockTarget, cloneCursor, type ParsedRange, type Token, Tokenizer } from "./tokenizer"; import type { Anchor, Cursor, Edit } from "./types"; function validateRangeOrder(range: ParsedRange, lineNum: number): void { if (range.end.line < range.start.line) { - throw new Error(`line ${lineNum}: range ${range.start.line}-${range.end.line} ends before it starts.`); + throw new Error(`line ${lineNum}: range ${range.start.line}..${range.end.line} ends before it starts.`); } } /** - * If `text` (the slice after a `+` literal sigil) trims to `^A-B` (or `^A`, - * accepted as `^A-A`), return the parsed range. Otherwise `null`. Used to - * silently reroute `+^A-B` rows as repeats — models reflexively prefix every + * If `text` (the slice after a `+` literal sigil) trims to `&A..B` (or `&A`, + * accepted as `&A,A`), return the parsed range. Otherwise `null`. Used to + * silently reroute `+&A..B` rows as repeats — models reflexively prefix every * body row with `+`, including ones that should be repeats. */ function tryParseLiteralAsRepeat(text: string): ParsedRange | null { const stripped = text.trim(); - if (stripped.length === 0 || stripped.charCodeAt(0) !== 94 /* ^ */) return null; - const match = /^\^([1-9]\d*)(?:-([1-9]\d*))?$/.exec(stripped); + if (stripped.length === 0 || stripped.charCodeAt(0) !== 38 /* & */) return null; + const match = /^&([1-9]\d*)(?:\.\.([1-9]\d*))?$/.exec(stripped); if (match === null) return null; const start = Number.parseInt(match[1], 10); const end = match[2] !== undefined ? Number.parseInt(match[2], 10) : start; @@ -69,18 +66,10 @@ function rangesOverlapBetweenTargets(a: BlockTarget, b: BlockTarget): boolean { * Detect OpenAI-`apply_patch` / unified-diff contamination in a raw line. * Returns the error message to throw, or `null` when the line is clean. * - * We only catch shapes that are unambiguously NOT hashline: - * - `*** Update File:` / `*** Add File:` / `*** Delete File:` / `*** Move to:` sentinels - * - unified-diff hunk headers (`@@`, `@@ -1,3 +1,3 @@`) - * - apply_patch hunk-anchor prefixes `-N:` / `-N-M:` — the bare `-N` form - * (no `:` and no `-M`) is intentionally NOT matched so the existing strict - * "unrecognized hashline block" diagnostic still fires on the legacy - * delete-row shape `-5`. - * - * `+`-prefixed shapes are NOT detected here because `+` is hashline's - * literal payload sigil; `+TEXT` / `+N:` are valid payload rows (or, at - * top level, orphan payloads that fall through to the standard "no - * preceding A-B:" error). + * Hashline's own file-header prefix (`¶path#hash`) sits next to + * apply_patch sentinels (`*** Update File: path`); the latter are caught + * here. Any `@@`-bracketed shape is also caught — hashline hunks are bare + * `A B` lines, never `@@ ... @@`. */ function detectApplyPatchContamination(text: string, _hasPending: boolean): string | null { const trimmed = text.trimStart(); @@ -95,19 +84,21 @@ function detectApplyPatchContamination(text: string, _hasPending: boolean): stri const preview = trimmed.length > 48 ? `${trimmed.slice(0, 48)}…` : trimmed; return ( `apply_patch sentinel ${JSON.stringify(preview)} is not valid in hashline. ` + - `Use \`${"\u00b6"}PATH#HASH\` then \`A-B:\` / \`A-B:-\` / \`BOF:\` / \`EOF:\` blocks; do not wrap edits in another format's envelope.` + "File sections start with `¶path#HASH` (no `Update File:` / `Add File:` keyword). " + + "Hunks are bare `A B` lines with `+TEXT` / `&A..B` body rows." ); } - if (trimmed === "@@" || trimmed.startsWith("@@ ") || trimmed.startsWith("@@\t")) { + if (/^@@\s+[-+]?\d+,\d+\s+[-+]?\d+,\d+\s+@@/.test(trimmed)) { return ( - "unified-diff hunk header (`@@`) is not valid in hashline. " + - "Use a `¶PATH#HASH` header and bare `A-B:` anchor blocks." + "unified-diff hunk header (`@@ -N,M +N,M @@`) is not valid in hashline. " + + "Hashline hunks are bare `A B` lines (or `BOF` / `EOF` keywords)." ); } - if (/^-\d+(-\d+)?:/.test(trimmed)) { + if (trimmed.startsWith("@@")) { + const preview = trimmed.length > 48 ? `${trimmed.slice(0, 48)}…` : trimmed; return ( - "apply_patch line prefix (`-N:` / `-N-M:`) is not valid in hashline. " + - "Drop the `-` prefix; use `A-B:` (replace) or `A-B:-` (delete) on the anchor line itself." + `\`@@\`-bracketed hunk header ${JSON.stringify(preview)} is not valid in hashline. ` + + "Drop the `@@ ... @@` brackets and write the range directly: `5 7` (or `5` for a single line, `BOF` / `EOF` for virtual positions)." ); } return null; @@ -129,13 +120,6 @@ function isSkippableCommentLine(line: string): boolean { return line.trimStart().startsWith("#"); } -function describeTarget(target: BlockTarget): string { - if (target.kind === "bof") return "BOF:"; - if (target.kind === "eof") return "EOF:"; - const { start, end } = target.range; - return `${start.line}-${end.line}:`; -} - interface PendingComment { lineNum: number; text: string; @@ -150,22 +134,29 @@ interface Pending { lineNum: number; payloads: PayloadRow[]; /** - * Bare body rows (no `|`/`^` prefix) buffered while we wait to see - * whether the entire block is uniformly unprefixed. On flush, if every - * row was bare AND no `|`/`^` row was ever observed for this block, we - * auto-pipe the buffered rows and emit a {@link BARE_BODY_AUTO_PIPED_WARNING}. + * Bare body rows (no `+`/`&` prefix) buffered while we wait to see + * whether the entire hunk body is uniformly unprefixed. On flush, if + * every row was bare AND no `+`/`&` row was ever observed for this hunk, + * we auto-prepend `+` and emit a {@link BARE_BODY_AUTO_PIPED_WARNING}. */ pendingRaws: { text: string; lineNum: number }[]; + /** + * Set true the first time a `-` row arrives inside the hunk body. From + * then on we strip one leading space from raw rows (treating them as + * unified-diff context lines) and retroactively strip the same space + * from prior `pendingRaws`/`payloads` literals that began with a space. + */ + unifiedDiffMode: boolean; } /** * Token-driven state machine that turns a stream of {@link Token}s into a * flat list of {@link Edit}s. * - * `feed()` accepts tokens one at a time; block payload rows accumulate until - * the next anchor block or {@link end} flushes them. After `terminated` flips - * true (on `envelope-end` or `abort`) subsequent feeds are silently ignored - * so callers can keep draining their tokenizer. + * `feed()` accepts tokens one at a time; hunk body rows accumulate until + * the next hunk header or {@link end} flushes them. After `terminated` + * flips true (on `envelope-end` or `abort`) subsequent feeds are silently + * ignored so callers can keep draining their tokenizer. */ export class Executor { #edits: Edit[] = []; @@ -174,13 +165,6 @@ export class Executor { #pending: Pending | undefined; #terminated = false; #skippableComments: PendingComment[] = []; - /** - * Length of the current run of consecutive single-line empty-body - * replacements (`A-A:` with no payload). Reset on every non-matching - * flush; surfaces {@link STACKED_BLANK_REPLACE_WARNING} when the run - * reaches two. - */ - #blankSingleRun = 0; #discardPendingSkippableComments(): void { this.#skippableComments = []; @@ -242,44 +226,10 @@ export class Executor { return; case "op-block": this.#discardPendingSkippableComments(); - if (token.deleteSuffix) { - if (token.target.kind !== "range") { - throw new Error(`line ${token.lineNum}: ${VIRTUAL_REPLACE_REJECTED_MESSAGE}`); - } - validateRangeOrder(token.target.range, token.lineNum); - // L5 (delete-suffix variant): if pending is a bare anchor that - // overlaps the new delete range, drop it silently — the model - // expressed `A-B:` then `A-B:-` (the classic before-then-after - // shape) and the actual intent is "delete A-B". - if ( - this.#pending !== undefined && - !pendingHasAnyContent(this.#pending) && - rangesOverlapBetweenTargets(this.#pending.target, token.target) - ) { - this.#pending = undefined; - this.#blankSingleRun = 0; - if (!this.#warnings.includes(REPLACE_PAIR_COALESCED_OVERLAP_WARNING)) { - this.#warnings.push(REPLACE_PAIR_COALESCED_OVERLAP_WARNING); - } - } else { - this.#flushPending(); - } - for (const anchor of expandRange(token.target.range)) { - this.#pushDelete(anchor, token.lineNum); - } - this.#blankSingleRun = 0; - return; - } - if (token.inlineBody !== undefined) { - throw new Error( - `line ${token.lineNum}: ${INLINE_PAYLOAD_REJECTED_PREFIX} ` + - `Write the anchor on its own line (e.g. ${describeTarget(token.target)}), then put the body content on the next line prefixed with ` + - `${HL_PAYLOAD_REPLACE} (literal) or ${HL_PAYLOAD_REPEAT}A-B (repeat). If you pasted "${describeTarget(token.target).slice(0, -1)}CONTENT" from \`read\` output, strip the leading "${describeTarget(token.target).slice(0, -1)}" and prefix the rest with ${HL_PAYLOAD_REPLACE}.`, - ); - } if (token.target.kind === "range") validateRangeOrder(token.target.range, token.lineNum); + if (this.#pending !== undefined && targetsEqualConcreteRange(this.#pending.target, token.target)) { - // Identical-range coalesce: drop the first block. Last-wins. + // Identical-range coalesce: drop the first hunk. Last-wins. this.#pending = undefined; if (!this.#warnings.includes(REPLACE_PAIR_COALESCED_WARNING)) { this.#warnings.push(REPLACE_PAIR_COALESCED_WARNING); @@ -289,9 +239,7 @@ export class Executor { !pendingHasAnyContent(this.#pending) && rangesOverlapBetweenTargets(this.#pending.target, token.target) ) { - // L5 (replace variant): bare pending block overlaps the new - // concrete block; treat as before/after pair, drop the bare - // one. The new block becomes pending. + // Overlapping bare-then-concrete: drop the bare one. this.#pending = undefined; if (!this.#warnings.includes(REPLACE_PAIR_COALESCED_OVERLAP_WARNING)) { this.#warnings.push(REPLACE_PAIR_COALESCED_OVERLAP_WARNING); @@ -299,19 +247,25 @@ export class Executor { } else { this.#flushPending(); } - this.#pending = { target: token.target, lineNum: token.lineNum, payloads: [], pendingRaws: [] }; + this.#pending = { + target: token.target, + lineNum: token.lineNum, + payloads: [], + pendingRaws: [], + unifiedDiffMode: false, + }; return; } } /** - * Flush any open pending block and return the accumulated edits and - * warnings. The executor is single-use; {@link reset} is required for reuse. + * Flush any open pending hunk and return the accumulated edits and + * warnings. The executor is single-use; {@link reset} is required for + * reuse. * - * Throws if two replacement/delete blocks target the same line with - * non-identical ranges. Identical-range blocks in the same hunk are - * coalesced last-wins by `feed()` with a warning, so they never reach the - * validator. + * Throws if two hunks target the same line with non-identical ranges. + * Identical-range hunks in the same patch are coalesced last-wins by + * `feed()` with a warning, so they never reach the validator. */ end(): { edits: Edit[]; warnings: string[] } { this.#consumePendingSkippableComments(); @@ -322,9 +276,9 @@ export class Executor { /** * Streaming-tolerant variant of {@link end}. Identical, except a pending - * block whose payload has not yet accumulated any rows is treated as still - * in flight and dropped instead of flushed (which would otherwise preview a - * destructive bare delete while the model may still be typing payload). + * hunk whose body has not yet accumulated any rows is treated as still + * in flight and dropped instead of flushed (which would otherwise commit + * a destructive delete while the model may still be typing payload). */ endStreaming(): { edits: Edit[]; warnings: string[] } { this.#consumePendingSkippableComments(); @@ -345,14 +299,13 @@ export class Executor { this.#pending = undefined; this.#skippableComments = []; this.#terminated = false; - this.#blankSingleRun = 0; } /** - * Each replacement/delete block contributes a delete edit per line in its - * range; if any line ends up targeted by deletes originating from two - * different source blocks (distinguished by their `lineNum`), the patch is - * internally inconsistent. + * Each hunk contributes a delete edit per line in its range; if any line + * ends up targeted by deletes originating from two different source + * hunks (distinguished by their `lineNum`), the patch is internally + * inconsistent. */ #validateNoOverlappingDeletes(): void { const sourceLinesByAnchor = new Map(); @@ -369,8 +322,8 @@ export class Executor { if (sourceLines.length < 2) continue; const [firstBlock, secondBlock] = [...sourceLines].sort((a, b) => a - b); throw new Error( - `line ${secondBlock}: anchor line ${anchorLine} is already targeted by another op on line ${firstBlock}. ` + - `Issue ONE block per range; payload is only the final desired content, never a before/after pair.`, + `line ${secondBlock}: anchor line ${anchorLine} is already targeted by another hunk on line ${firstBlock}. ` + + `Issue ONE hunk per range; payload is only the final desired content, never a before/after pair.`, ); } } @@ -379,13 +332,13 @@ export class Executor { const pending = this.#pending; if (!pending) { throw new Error( - `line ${lineNum}: payload line has no preceding A-B:, BOF:, or EOF: anchor. ` + + `line ${lineNum}: payload line has no preceding hunk header. ` + `Got ${JSON.stringify(`${HL_PAYLOAD_REPLACE}${text}`)}.`, ); } - // Silent recovery: a body row of `+^A-B` (or `+^A` after L2 shorthand) - // is a repeat row the model mistakenly prefixed with `+`. Reroute as - // a repeat and surface a warning so the model sees the mistake. + // Silent recovery: a body row of `+&A..B` (or `+&A` shorthand) is a + // repeat row the model mistakenly prefixed with `+`. Reroute as a + // repeat and surface a warning so the model sees the mistake. const repeatRange = tryParseLiteralAsRepeat(text); if (repeatRange !== null) { if (!this.#warnings.includes(PLUS_PREFIXED_REPEAT_WARNING)) { @@ -394,11 +347,6 @@ export class Executor { this.#handleRepeatPayload(repeatRange, lineNum); return; } - // L3: a `+literal` row after buffered bare raws means the block is - // NOT uniformly unprefixed — the bare rows were typos. Reject at the - // FIRST bare row's source line so the message points the model at - // what to fix. - this.#rejectBufferedRawsOnMixedBlock(pending); pending.payloads.push({ kind: "literal", text, lineNum }); } @@ -406,91 +354,75 @@ export class Executor { const pending = this.#pending; if (!pending) { throw new Error( - `line ${lineNum}: payload line has no preceding A-B:, BOF:, or EOF: anchor. ` + - `Got ${JSON.stringify(`${HL_PAYLOAD_REPEAT}${range.start.line}-${range.end.line}`)}.`, + `line ${lineNum}: payload line has no preceding hunk header. ` + + `Got ${JSON.stringify(`${HL_PAYLOAD_REPEAT}${range.start.line}..${range.end.line}`)}.`, ); } - // L3: same mixed-block guard as the literal path — see above. - this.#rejectBufferedRawsOnMixedBlock(pending); validateRangeOrder(range, lineNum); pending.payloads.push({ kind: "repeat", range, lineNum }); } - #rejectBufferedRawsOnMixedBlock(pending: Pending): void { - if (pending.pendingRaws.length === 0) return; - const first = pending.pendingRaws[0]; - throw new Error( - `line ${first.lineNum}: payload row in a hashline block must start with ` + - `${HL_PAYLOAD_REPLACE} or ${HL_PAYLOAD_REPEAT}A-B. Got ${JSON.stringify(first.text)}.`, - ); + /** + * Switch the pending hunk into unified-diff mode and retroactively + * strip the leading metadata-space from any literal payloads or + * buffered raws that already arrived. Idempotent. + */ + #enterUnifiedDiffMode(pending: Pending): void { + if (pending.unifiedDiffMode) return; + pending.unifiedDiffMode = true; + for (const row of pending.pendingRaws) { + if (row.text.length > 0 && row.text.charCodeAt(0) === 32) { + row.text = row.text.slice(1); + } + } + for (const payload of pending.payloads) { + if (payload.kind === "literal" && payload.text.length > 0 && payload.text.charCodeAt(0) === 32) { + payload.text = payload.text.slice(1); + } + } } #handleRaw(text: string, lineNum: number): void { - // L8: detect OpenAI-apply_patch / unified-diff contamination first so - // the error message tells the model what format they shipped instead - // of the generic "payload row must start with …" diagnostic. + // Detect OpenAI-apply_patch / unified-diff contamination first so the + // error message names the offending shape instead of the generic + // "payload row must start with …" diagnostic. const contamination = detectApplyPatchContamination(text, this.#pending !== undefined); if (contamination !== null) throw new Error(`line ${lineNum}: ${contamination}`); if (this.#pending) { if (text.trim().length === 0) return; - // L4: a lone `-` row inside a bare pending block is the classic - // "I meant `A-B:-` but typed it on the next line" shape. Convert - // retroactively, emit a warning, and clear pending. - if ( - text.trim() === "-" && - this.#pending.payloads.length === 0 && - this.#pending.pendingRaws.length === 0 && - this.#pending.target.kind === "range" - ) { - const pendingRange = this.#pending.target.range; - const sourceLine = this.#pending.lineNum; - validateRangeOrder(pendingRange, sourceLine); - for (const anchor of expandRange(pendingRange)) { - this.#pushDelete(anchor, sourceLine); + // L9: `-`-prefixed body rows are unified-diff "removed" markers. + // The hunk header's range already deletes those lines, so we + // silently drop them and enter unified-diff mode for subsequent + // rows (which causes leading-space stripping on context lines). + if (text.charCodeAt(0) === 45 /* - */) { + this.#enterUnifiedDiffMode(this.#pending); + if (!this.#warnings.includes(UNIFIED_DIFF_BODY_AUTO_CONVERT_WARNING)) { + this.#warnings.push(UNIFIED_DIFF_BODY_AUTO_CONVERT_WARNING); } - if (!this.#warnings.includes(DASH_PAYLOAD_AUTO_DELETE_WARNING)) { - this.#warnings.push(DASH_PAYLOAD_AUTO_DELETE_WARNING); - } - this.#pending = undefined; - this.#blankSingleRun = 0; return; } - // L3: buffer the bare row. Mixing this with later `|`/`^` rows - // throws via `#rejectBufferedRawsOnMixedBlock`; reject IMMEDIATELY - // when the block already has `|`/`^` rows so the error points at - // the offending bare row, not the (innocent) first `|` row. - if (this.#pending.payloads.length > 0) { - throw new Error( - `line ${lineNum}: payload row in a hashline block must start with ` + - `${HL_PAYLOAD_REPLACE} or ${HL_PAYLOAD_REPEAT}A-B. Got ${JSON.stringify(text)}.`, - ); + // Treat any non-`+`/`&` body row as a literal. When the hunk is + // in unified-diff mode and the row carries the metadata leading + // space, strip ONE space so the actual content lands cleanly. + const literalText = + this.#pending.unifiedDiffMode && text.charCodeAt(0) === 32 /* space */ ? text.slice(1) : text; + if (!this.#warnings.includes(BARE_BODY_AUTO_PIPED_WARNING)) { + this.#warnings.push(BARE_BODY_AUTO_PIPED_WARNING); } - this.#pending.pendingRaws.push({ text, lineNum }); + this.#pending.payloads.push({ kind: "literal", text: literalText, lineNum }); return; } - // Whitespace-only raw lines outside any pending block are silently dropped; - // fully empty lines arrive as `blank` tokens. + // Whitespace-only raw lines outside any pending block are silently + // dropped; fully empty lines arrive as `blank` tokens. if (text.trim().length === 0) return; - const firstChar = text[0]; - if (firstChar === "-" || firstChar === "@" || firstChar === "«" || firstChar === "»") { - if (text.trim() === "-") { - throw new Error( - `line ${lineNum}: a lone "-" is not a valid hashline op. To delete a range, write \`A-B:-\` on the anchor line itself (e.g. \`5-7:-\`).`, - ); - } - throw new Error( - `line ${lineNum}: unrecognized hashline block. Use A-B:, A-B:-, BOF:, or EOF: anchors followed by ` + - `${HL_PAYLOAD_REPLACE}TEXT or ${HL_PAYLOAD_REPEAT}A-B body rows. Got ${JSON.stringify(text)}.`, - ); - } - throw new Error( - `line ${lineNum}: payload line has no preceding A-B:, BOF:, or EOF: anchor. Got ${JSON.stringify(text)}.`, + `line ${lineNum}: payload line has no preceding hunk header. ` + + `Use an \`A B\` (or \`BOF\` / \`EOF\`) line above the body. Got ${JSON.stringify(text)}.`, ); } @@ -532,67 +464,31 @@ export class Executor { const pending = this.#pending; if (!pending) return; - // L3: convert any buffered bare body rows to literal payloads. Mixed + // Convert any buffered bare body rows to literal payloads. Mixed // blocks have already been rejected; we only get here when payloads - // is empty AND pendingRaws holds rows, or when both are empty. - const hadBareBody = pending.pendingRaws.length > 0; - if (hadBareBody) { - for (const raw of pending.pendingRaws) { - pending.payloads.push({ kind: "literal", text: raw.text, lineNum: raw.lineNum }); - } - pending.pendingRaws = []; - if (!this.#warnings.includes(BARE_BODY_AUTO_PIPED_WARNING)) { - this.#warnings.push(BARE_BODY_AUTO_PIPED_WARNING); - } - } - + // `pendingRaws` is kept for type compatibility but no longer used — + // bare rows are now pushed directly into `payloads` as literals at + // arrival time (preserving body-row order). const { target, lineNum, payloads } = pending; if (target.kind === "bof" || target.kind === "eof") { const cursor: Cursor = target.kind === "bof" ? { kind: "bof" } : { kind: "eof" }; - if (payloads.length === 0) { - this.#pushInsert(cursor, "", lineNum); - } else { - for (const payload of payloads) { - this.#emitPayloadRow(cursor, payload, lineNum); - } + for (const payload of payloads) { + this.#emitPayloadRow(cursor, payload, lineNum); } + // Empty body at BOF/EOF is a no-op (nothing to insert). this.#pending = undefined; - this.#blankSingleRun = 0; return; } - // L7 was considered (`^A-B` covering target + literal payload) but - // dropped: the same shape is the canonical "keep line A unchanged, - // insert new content above/below" idiom (e.g. `2-2:\n^2-2\n|NEW`). - // We can't distinguish duplication from intentional pass-through - // from the parse tree alone. - const cursor: Cursor = { kind: "before_anchor", anchor: { ...target.range.start } }; - if (payloads.length === 0) { - this.#pushInsert(cursor, "", lineNum, "replacement"); - } else { - for (const payload of payloads) { - this.#emitPayloadRow(cursor, payload, lineNum, "replacement"); - } + // Empty body = pure delete. Otherwise, emit the body rows as + // replacement payload and delete the original range. + for (const payload of payloads) { + this.#emitPayloadRow(cursor, payload, lineNum, "replacement"); } for (const anchor of expandRange(target.range)) { this.#pushDelete(anchor, lineNum); } - - // L6: track contiguous runs of single-line blank-body replaces. A - // run of two or more is almost always the model mis-using `A-A:` to - // mean "delete this line" (it actually replaces with one blank line). - const isBlankSingleReplace = - target.range.start.line === target.range.end.line && payloads.length === 0 && !hadBareBody; - if (isBlankSingleReplace) { - this.#blankSingleRun++; - if (this.#blankSingleRun >= 2 && !this.#warnings.includes(STACKED_BLANK_REPLACE_WARNING)) { - this.#warnings.push(STACKED_BLANK_REPLACE_WARNING); - } - } else { - this.#blankSingleRun = 0; - } - this.#pending = undefined; } } @@ -623,14 +519,14 @@ export function parsePatch(diff: string): { edits: Edit[]; warnings: string[] } * parsed successfully when the diff is still being typed: * * - per-token feed errors stop the drain but preserve the edits already - * collected (the trailing block is malformed mid-stream — wait for the next - * chunk), - * - the trailing pending block is dropped if it has no payload yet (avoids a - * destructive bare-delete preview while payload may still be coming). + * collected (the trailing hunk is malformed mid-stream — wait for the + * next chunk), + * - the trailing pending hunk is dropped if it has no payload yet (avoids + * a destructive bare-delete preview while payload may still be coming). * - * Throws only on the cross-block overlap validator, which catches conflicting - * shapes (two replacements/deletes hitting the same anchor). Streaming preview - * callers should treat any throw here as "no preview this tick". + * Throws only on the cross-hunk overlap validator, which catches conflicting + * shapes (two hunks hitting the same anchor). Streaming preview callers + * should treat any throw here as "no preview this tick". */ export function parsePatchStreaming(diff: string): { edits: Edit[]; warnings: string[] } { const tokenizer = new Tokenizer(); diff --git a/packages/hashline/src/patcher.ts b/packages/hashline/src/patcher.ts index a9ebc21f3..6946b5130 100644 --- a/packages/hashline/src/patcher.ts +++ b/packages/hashline/src/patcher.ts @@ -1,8 +1,8 @@ /** * High-level patch orchestrator. Reads each section's target file via the * configured {@link Filesystem}, strips BOM and normalizes line endings, - * validates the section snapshot tag (with optional {@link Recovery}), applies - * the edits, and writes the result back through the same {@link Filesystem}. + * validates the section snapshot tag (with {@link Recovery}), applies the + * result back through the same {@link Filesystem}. * * Two layers: * @@ -31,18 +31,13 @@ import { MismatchError } from "./mismatch"; import { detectLineEnding, type LineEnding, normalizeToLF, restoreLineEndings, stripBom } from "./normalize"; import { Recovery, type RecoveryResult } from "./recovery"; import type { Snapshot, SnapshotStore } from "./snapshots"; -import type { ApplyOptions, ApplyResult, Edit } from "./types"; +import type { ApplyResult, Edit } from "./types"; export interface PatcherOptions { /** Storage backend used for all reads and writes. */ fs: Filesystem; /** Snapshot store that minted and resolves hashline section tags. Required. */ snapshots: SnapshotStore; - /** - * Optional default {@link ApplyOptions} forwarded to every section. - * Per-call overrides win on a key-by-key basis. - */ - applyOptions?: ApplyOptions; } /** Per-section result returned by {@link Patcher.apply} / {@link Patcher.commit}. */ @@ -122,6 +117,24 @@ function recoveryToApplyResult(result: RecoveryResult): ApplyResult { }; } +/** + * Decide whether `snapshot` proves the live file is byte-for-byte the read + * the model authored against. Two shapes: + * - Full-text snapshot: cheap string equality. + * - Sparse snapshot (e.g. selector reads, search hits): every anchor line + * must be in the snapshot AND every recorded line must match the live + * file. Without this branch, sparse reads can't short-circuit and fall + * through to recovery, which declines them as "patcher-owned direct + * apply" — yielding a spurious MismatchError on unchanged files. + */ +function snapshotProvesUnchanged(snapshot: Snapshot, currentText: string, section: PatchSection): boolean { + if (snapshot.fullText !== undefined) return snapshot.fullText === currentText; + for (const lineNumber of section.collectAnchorLines()) { + if (snapshot.get(lineNumber) === undefined) return false; + } + return snapshot.matchesLiveFile(currentText.split("\n")); +} + function mergeWarnings(...sources: ReadonlyArray): string[] { const out: string[] = []; for (const source of sources) { @@ -144,14 +157,6 @@ function assertUniqueCanonicalPaths(prepared: readonly PreparedSection[]): void } } -function snapshotMatchesCurrent(snapshot: Snapshot, currentText: string, anchorLines: readonly number[]): boolean { - if (snapshot.fullText !== undefined) return snapshot.fullText === currentText; - for (const lineNumber of anchorLines) { - if (snapshot.get(lineNumber) === undefined) return false; - } - return snapshot.matchesLiveFile(currentText.split("\n")); -} - /** * High-level patcher. Wires a {@link Filesystem} and a required * {@link SnapshotStore} together with the parsing + applying core. @@ -162,7 +167,6 @@ export class Patcher { readonly fs: Filesystem; readonly snapshots: SnapshotStore; readonly recovery: Recovery; - readonly applyOptions: ApplyOptions; constructor(options: PatcherOptions) { if (!options.snapshots) { @@ -171,7 +175,6 @@ export class Patcher { this.fs = options.fs; this.snapshots = options.snapshots; this.recovery = new Recovery(options.snapshots); - this.applyOptions = options.applyOptions ?? {}; } /** @@ -180,19 +183,17 @@ export class Patcher { * multi-section batch is naturally all-or-nothing. Returns one * {@link PatchSectionResult} per section in the original patch order. */ - async apply(patch: Patch, options: ApplyOptions = {}): Promise { - const merged: ApplyOptions = { ...this.applyOptions, ...options }; - + async apply(patch: Patch): Promise { // Single-section fast path. if (patch.sections.length === 1) { - const prepared = await this.prepare(patch.sections[0], merged); + const prepared = await this.prepare(patch.sections[0]); return { sections: [await this.commit(prepared)] }; } // Prepare every section first so any failure (stale hash, missing // file, parse error, in-memory no-op) surfaces before any write. const prepared: PreparedSection[] = []; - for (const section of patch.sections) prepared.push(await this.prepare(section, merged)); + for (const section of patch.sections) prepared.push(await this.prepare(section)); assertUniqueCanonicalPaths(prepared); for (const entry of prepared) { if (entry.isNoop) { @@ -209,10 +210,9 @@ export class Patcher { * Run the preflight pass only: read, parse, validate, apply-in-memory. * No writes hit the filesystem. Use for CI checks and dry runs. */ - async preflight(patch: Patch, options: ApplyOptions = {}): Promise { - const merged: ApplyOptions = { ...this.applyOptions, ...options }; + async preflight(patch: Patch): Promise { const prepared: PreparedSection[] = []; - for (const section of patch.sections) prepared.push(await this.prepare(section, merged)); + for (const section of patch.sections) prepared.push(await this.prepare(section)); assertUniqueCanonicalPaths(prepared); for (const entry of prepared) { if (entry.isNoop) { @@ -230,8 +230,7 @@ export class Patcher { * Throws on parse error, missing-file-for-anchored-edit, or unrecovered * tag mismatch ({@link MismatchError}). */ - async prepare(section: PatchSection, options: ApplyOptions = {}): Promise { - const applyOptions: ApplyOptions = { ...this.applyOptions, ...options }; + async prepare(section: PatchSection): Promise { const { edits, warnings: parseWarnings } = section.parse(); assertSectionHashAllowed(section.path, section.fileHash, edits); @@ -252,7 +251,6 @@ export class Patcher { exists, normalized, edits, - applyOptions, }); return new PreparedSection( @@ -335,25 +333,21 @@ export class Patcher { exists: boolean; normalized: string; edits: readonly Edit[]; - applyOptions: ApplyOptions; }): ApplyResult { - const { section, canonicalPath, exists, normalized, edits, applyOptions } = args; + const { section, canonicalPath, exists, normalized, edits } = args; const expected = exists ? section.fileHash : undefined; - if (expected === undefined) return applyEdits(normalized, [...edits], applyOptions); + if (expected === undefined) return applyEdits(normalized, [...edits]); const snapshot = this.snapshots.byHash(canonicalPath, expected); - const anchorLines = section.collectAnchorLines(); - if (snapshot && snapshotMatchesCurrent(snapshot, normalized, anchorLines)) { - return applyEdits(normalized, [...edits], applyOptions); + if (snapshot && snapshotProvesUnchanged(snapshot, normalized, section)) { + return applyEdits(normalized, [...edits]); } - if (snapshot) { const recovered = this.recovery.tryRecover({ path: canonicalPath, currentText: normalized, fileHash: expected, edits, - options: applyOptions, }); if (recovered) return recoveryToApplyResult(recovered); } diff --git a/packages/hashline/src/prompt.md b/packages/hashline/src/prompt.md index e0c3116e5..2eb7ff8d9 100644 --- a/packages/hashline/src/prompt.md +++ b/packages/hashline/src/prompt.md @@ -1,12 +1,12 @@ -Your patch language selects ranges of file lines and rewrites them. The body rows below an anchor describe the new content of the selected range. +Your patch language selects ranges of file lines and rewrites them. Each hunk picks a range and lists its new content; an empty body deletes the range. Every body row is **exactly one** of two kinds: +TEXT add a new literal line `TEXT` (verbatim, leading whitespace included) - ^A-B keep original lines A..B as-is + &A..B keep original lines A..B as-is -`+` and `^` are siblings, not stackable. Never write `+^…`. A row starts with one of them, never both. +`+` and `&` are siblings, not stackable. Never write `+&…`. A row starts with one of them, never both. @@ -18,13 +18,13 @@ This is the original file (the exact shape `read` returns): 3:} ``` -To add a null check between the signature and the return, select lines 1..3 and rewrite: +To add a null check between the signature and the return, open a hunk on lines 1..3 and list its new content: ``` ¶greet.ts#0A3 -1-3: -^1-1 +1 3 +&1 + if (!name) return "Hello, stranger!"; -^2-3 +&2..3 ``` The body says: keep line 1, then add the new literal line, then keep lines 2..3. Result: @@ -38,31 +38,33 @@ The body says: keep line 1, then add the new literal line, then keep lines 2..3. ``` -A-B: select lines A..B; the body rows below describe their new content -A-B:- delete lines A..B (no body) -BOF: virtual position before line 1; body rows insert there -EOF: virtual position after the last line; body rows insert there +A B select lines A..B; the body rows below describe their new content + (empty body = delete the range) +A select single line A (shorthand for `A A`) +BOF virtual position before line 1; body rows insert there +EOF virtual position after the last line; body rows insert there ``` -`A-A:` for one line is preferred over the bare shorthand `A:`. `BOF:` / `EOF:` take no range. + +A hunk header is **just the anchor on its own line** — no `@@`, no brackets, no prefix.
-Every section starts with `¶PATH#HASH`. `HASH` is the snapshot tag from your latest `read`/`search` of that file. It is required whenever a block uses a line-number anchor (`A-B:` or `A-B:-`). Hashless `¶PATH` is only valid for new-file creation or BOF/EOF-only patches. +Every file section starts with `¶PATH#HASH`. `HASH` is the snapshot tag from your latest `read`/`search` of that file. It is required whenever a hunk uses a numeric anchor. Hashless `¶PATH` is only valid for new-file creation or BOF/EOF-only patches.
-- Anchors are line **numbers**, never line **content**. `read` shows each file row as `LINE:TEXT`; for a patch the anchor is `4-4:` and the body is `+TEXT` (or `^4-4` to keep it). -- Each range may appear in only ONE block per patch. -- Line numbers refer to the ORIGINAL file and stay valid for the whole patch — they do not shift as your blocks land. -- `A-B:` with no body replaces the range with ONE blank line. To **delete** the lines entirely, use `A-B:-`. -- If you want to replace lines A..B with completely new content, just list the new content; do not write `^A-B`. +- Anchors are line **numbers**, never line **content**. `read` shows each file row as `LINE:TEXT`; for a patch the hunk header is `4` (or `4 4`) and the body is `+TEXT` (or `&4` to keep it). +- Each range may appear in only ONE hunk per patch. +- Line numbers refer to the ORIGINAL file and stay valid for the whole patch — they do not shift as your hunks land. +- An empty body **deletes** the selected range entirely. To replace lines A..B with completely new content, list the new content under the hunk header (do not write `&A..B` for the lines you are replacing). +- `@@` is NOT a hashline construct. Do not wrap headers in `@@ ... @@` — write the anchor bare. # Replace line 1 of `greet.ts#0A3` with two new lines. ``` ¶greet.ts#0A3 -1-1: +1 +const X = "b"; +export const Y = X; ``` @@ -70,33 +72,33 @@ Every section starts with `¶PATH#HASH`. `HASH` is the snapshot tag from your la # Delete lines 2..3 of `greet.ts#0A3`. ``` ¶greet.ts#0A3 -2-3:- +2 3 ``` # Prepend a header. ``` ¶greet.ts#0A3 -BOF: +BOF +// generated header ``` -# WRONG — two blocks expressing old → new. Rejected as overlap. -1-1: -1-1:- +# WRONG — do not include old lines. +2 3 +- print "hello" ++ print "hi" -# WRONG — `A-A:` with no body REPLACES with a blank line. Use `A-A:-` to delete. -2-2: -2-2:- +# WRONG — do not include context lines. +2 3 + fn hi(): ++ print "hi" -# WRONG — `read`-output rows pasted as body. Body rows need `+` or `^`. -2-3: - return `Hello, ${name}!`; -} +# WRONG — no `@@` brackets in hashline. +@@ 2..3 @@ ++ print "hi" # RIGHT — same intent, well-formed. -2-3: -+ return `Hello, ${name}!`; -+} +2 3 ++ print "hi" diff --git a/packages/hashline/src/recovery.ts b/packages/hashline/src/recovery.ts index 87602ec0b..229a3a771 100644 --- a/packages/hashline/src/recovery.ts +++ b/packages/hashline/src/recovery.ts @@ -12,7 +12,7 @@ import * as Diff from "diff"; import { applyEdits } from "./apply"; import { RECOVERY_EXTERNAL_WARNING, RECOVERY_SESSION_CHAIN_WARNING, RECOVERY_SESSION_REPLAY_WARNING } from "./messages"; import type { Snapshot, SnapshotStore } from "./snapshots"; -import type { Anchor, ApplyOptions, ApplyResult, Edit } from "./types"; +import type { Anchor, ApplyResult, Edit } from "./types"; // Section tags are line-precise; never let Diff.applyPatch slide a hunk // onto a duplicate closer 100+ lines away. If snapshot replay does not @@ -24,7 +24,6 @@ export interface RecoveryArgs { currentText: string; fileHash: string; edits: readonly Edit[]; - options?: ApplyOptions; } export interface RecoveryResult { @@ -40,12 +39,11 @@ function applyEditsToSnapshot( previousText: string, currentText: string, edits: readonly Edit[], - options: ApplyOptions, recoveryWarning: string, ): RecoveryResult | null { let applied: ApplyResult; try { - applied = applyEdits(previousText, [...edits], options); + applied = applyEdits(previousText, [...edits]); } catch { return null; } @@ -107,7 +105,6 @@ function replaySessionChainOnCurrent( previousText: string, currentText: string, edits: readonly Edit[], - options: ApplyOptions, ): RecoveryResult | null { // Two guards narrow the corruption window. Neither alone is sufficient, // and even together they don't fully prove correctness — replay is the @@ -126,7 +123,7 @@ function replaySessionChainOnCurrent( if (!verifyAnchorContent(previousText, currentText, edits)) return null; let applied: ApplyResult; try { - applied = applyEdits(currentText, [...edits], options); + applied = applyEdits(currentText, [...edits]); } catch { return null; } @@ -208,7 +205,7 @@ export class Recovery { * caller should then surface a {@link MismatchError}. */ tryRecover(args: RecoveryArgs): RecoveryResult | null { - const { path, currentText, fileHash, edits, options = {} } = args; + const { path, currentText, fileHash, edits } = args; const head = this.store.head(path); const snapshot = this.store.byHash(path, fileHash); if (!snapshot || !snapshotHasEntries(snapshot)) return null; @@ -218,19 +215,19 @@ export class Recovery { const isSessionChain = !isHead; if (snapshot.fullText !== undefined) { - const merged = applyEditsToSnapshot(snapshot.fullText, currentText, edits, options, recoveryWarning); + const merged = applyEditsToSnapshot(snapshot.fullText, currentText, edits, recoveryWarning); if (merged !== null) return merged; // Session-chain fallback: the 3-way merge on the snapshot refused. // Replay onto current is gated by line-count equality AND // anchor-content alignment — see `replaySessionChainOnCurrent` // for why both guards together still don't fully prove correctness. - if (isSessionChain) return replaySessionChainOnCurrent(snapshot.fullText, currentText, edits, options); + if (isSessionChain) return replaySessionChainOnCurrent(snapshot.fullText, currentText, edits); return null; } if (!sparseSnapshotCoversAnchors(snapshot, edits)) return null; if (sparseSnapshotMatchesCurrent(currentText, snapshot)) return null; const overlayText = buildSparseOverlayText(currentText, snapshot); - return applyEditsToSnapshot(overlayText, currentText, edits, options, recoveryWarning); + return applyEditsToSnapshot(overlayText, currentText, edits, recoveryWarning); } } diff --git a/packages/hashline/src/tokenizer.ts b/packages/hashline/src/tokenizer.ts index 702ca1167..f9ce36101 100644 --- a/packages/hashline/src/tokenizer.ts +++ b/packages/hashline/src/tokenizer.ts @@ -7,10 +7,16 @@ * line number so downstream consumers (parser, validators, error messages) * can refer back to the input precisely. * - * The tokenizer is intentionally permissive about decorations and prefixes - * the model may echo back from `read`/`search` output — leading `*`/`>`/`-` - * markers, CR-terminated lines, leading whitespace before line numbers, and - * so on are all stripped before anchor classification. + * Format shape: + * ``` + * *** path/to/file.ts#0A3 + * @@ 5,7 @@ + * +literal new line + * &3,4 + * ``` + * Each `***` line opens a new file section; each `@@ A,B @@` line opens a + * new hunk whose body (zero or more `+`/`&` rows) replaces the selected + * range. Empty body = delete the selected range. */ import { @@ -18,8 +24,6 @@ import { HL_FILE_HASH_LENGTH, HL_FILE_HASH_SEP, HL_FILE_PREFIX, - HL_OP_DELETE_SUFFIX, - HL_OP_REPLACE, HL_PAYLOAD_REPEAT, HL_PAYLOAD_REPLACE, } from "./format"; @@ -33,16 +37,19 @@ const CHAR_NINE = 57; const CHAR_HASH = 35; const CHAR_TAB = 9; const CHAR_SPACE = 32; +const CHAR_DOT = 46; const CHAR_HYPHEN = 45; +const CHAR_ELLIPSIS = 0x2026; const CHAR_UPPER_A = 65; const CHAR_UPPER_F = 70; const CHAR_LOWER_A = 97; const CHAR_LOWER_F = 102; -const CHAR_PILCROW = HL_FILE_PREFIX.charCodeAt(0); -const CHAR_OP_REPLACE = HL_OP_REPLACE.charCodeAt(0); const CHAR_PAYLOAD_REPLACE = HL_PAYLOAD_REPLACE.charCodeAt(0); const CHAR_PAYLOAD_REPEAT = HL_PAYLOAD_REPEAT.charCodeAt(0); +const FILE_PREFIX_LENGTH = HL_FILE_PREFIX.length; +const BOF_ANCHOR = "BOF"; +const EOF_ANCHOR = "EOF"; function isDigitCode(code: number): boolean { return code >= CHAR_ZERO && code <= CHAR_NINE; @@ -52,14 +59,6 @@ function isNonZeroDigitCode(code: number): boolean { return code > CHAR_ZERO && code <= CHAR_NINE; } -function isDecorationCode(code: number): boolean { - // `*` (grep match marker) and `>` (grep context marker). We intentionally - // do NOT include `-` here: a leading `-` is the unified-diff "removed - // line" prefix and the OpenAI apply_patch hunk-line prefix, both of - // which we want to reject loudly rather than silently swallow. - return code === 42 || code === 62; -} - function isHexDigitCode(code: number): boolean { return ( isDigitCode(code) || @@ -68,12 +67,19 @@ function isHexDigitCode(code: number): boolean { ); } +function isWhitespaceCode(code: number): boolean { + return code === CHAR_SPACE || (code >= CHAR_TAB && code <= CHAR_CARRIAGE_RETURN); +} + function skipWhitespace(line: string, index: number, end = line.length): number { - return end - line.slice(index, end).trimStart().length; + while (index < end && isWhitespaceCode(line.charCodeAt(index))) index++; + return index; } function trimEndIndex(line: string): number { - return line.trimEnd().length; + let end = line.length; + while (end > 0 && isWhitespaceCode(line.charCodeAt(end - 1))) end--; + return end; } function isEmptyLine(line: string): boolean { @@ -81,17 +87,14 @@ function isEmptyLine(line: string): boolean { } function markerLineEquals(line: string, marker: string): boolean { - return line.trimEnd() === marker; + const end = trimEndIndex(line); + return end === marker.length && line.startsWith(marker); } /** * Split a hashline diff into individual lines without losing the trailing * empty line that callers may rely on for explicit blank payloads. CRLF pairs * are normalized to a single line break. - * - * This mirrors the line-splitting performed by {@link Tokenizer}'s streaming - * drain loop and is kept for non-streaming callers that prefer a single-shot - * split. */ export function splitHashlineLines(text: string): string[] { if (text.length === 0) return [""]; @@ -118,14 +121,6 @@ export function cloneCursor(cursor: Cursor): Cursor { if (cursor.kind === "before_anchor") return { kind: "before_anchor", anchor: { ...cursor.anchor } }; return cursor; } -// Leniently accept anchors copied from read/search output: -// - optional leading line-marker decoration (`*`, `>`) -// - the required bare line number / BOF / EOF anchor -function skipDecoratedAnchorPrefix(line: string, end = trimEndIndex(line)): number { - let index = skipWhitespace(line, 0, end); - while (index < end && isDecorationCode(line.charCodeAt(index))) index++; - return skipWhitespace(line, index, end); -} interface NumberScan { line: number; @@ -149,7 +144,7 @@ function scanLineNumber(line: string, index: number, end: number): NumberScan | /** Parse a bare line-number anchor. Throws on malformed input. */ export function parseLid(raw: string, lineNum: number): Anchor { const end = trimEndIndex(raw); - const numberStart = skipDecoratedAnchorPrefix(raw, end); + const numberStart = skipWhitespace(raw, 0, end); const number = scanLineNumber(raw, numberStart, end); if (number === null || skipWhitespace(raw, number.nextIndex, end) !== end) { throw new Error( @@ -165,40 +160,67 @@ interface RangeScan { nextIndex: number; } -function scanRange(line: string, end = trimEndIndex(line)): RangeScan | null { - const numberStart = skipDecoratedAnchorPrefix(line, end); +/** + * Scan a numeric range for a hunk header. Canonical form is `A B` (two + * numbers separated by whitespace); models also reflexively emit `A-B`, + * `A..B`, and `A…B` (unicode ellipsis), so we accept any of those as the + * range separator. Bare `A` is the single-line shorthand for `A A`. + * Repeat-row bodies (`&A..B`) keep their own parser; see + * {@link tryParseRepeatPayload}. + */ +function scanHeaderRange(line: string, index = 0, end = trimEndIndex(line)): RangeScan | null { + const numberStart = skipWhitespace(line, index, end); const start = scanLineNumber(line, numberStart, end); if (start === null) return null; - // Canonical form is `A-B` (and `A-A` for one line). The bare single-line - // shorthand `A` is also accepted because models that learned the format - // from `read` output — which renders each file row as `LINE:content` — - // frequently reproduce that shape as an anchor. We treat `A` as `A-A` - // here so `tryParseBlockOp` can still raise its strict inline-payload - // diagnostic on `A:content` lines. - if (start.nextIndex < end && line.charCodeAt(start.nextIndex) === CHAR_HYPHEN) { - const endNumber = scanLineNumber(line, start.nextIndex + 1, end); + const afterFirst = scanRangeSeparator(line, start.nextIndex, end); + if (afterFirst !== null) { + const endNumber = scanLineNumber(line, afterFirst, end); if (endNumber === null) return null; return { range: { start: { line: start.line }, end: { line: endNumber.line } }, nextIndex: skipWhitespace(line, endNumber.nextIndex, end), }; } - if (start.nextIndex < end && line.charCodeAt(start.nextIndex) === CHAR_OP_REPLACE) { - return { - range: { start: { line: start.line }, end: { line: start.line } }, - nextIndex: start.nextIndex, - }; - } - return null; + // Shorthand: bare `A` treated as `A..A`. Trailing non-whitespace past + // `cursor` signals a malformed header (caller verifies). + return { + range: { start: { line: start.line }, end: { line: start.line } }, + nextIndex: skipWhitespace(line, start.nextIndex, end), + }; } -function startsWithWord(line: string, index: number, end: number, word: string): boolean { - if (index + word.length > end) return false; - for (let offset = 0; offset < word.length; offset++) { - if (line.charCodeAt(index + offset) !== word.charCodeAt(offset)) return false; +/** + * Consume an optional range separator (whitespace, `-`, `..`, or `…`) + * after the first number in a header. Returns the index of the second + * number, or `null` when the next non-whitespace char isn't a digit + * (i.e. we're looking at a single-line shorthand). + */ +function scanRangeSeparator(line: string, index: number, end: number): number | null { + let cursor = index; + let consumedSeparator = false; + while (cursor < end) { + const code = line.charCodeAt(cursor); + if (isWhitespaceCode(code)) { + cursor++; + consumedSeparator = true; + continue; + } + if (code === CHAR_HYPHEN || code === CHAR_ELLIPSIS) { + cursor++; + consumedSeparator = true; + continue; + } + if (code === CHAR_DOT && cursor + 1 < end && line.charCodeAt(cursor + 1) === CHAR_DOT) { + cursor += 2; + consumedSeparator = true; + continue; + } + break; } - return true; + if (!consumedSeparator) return null; + if (cursor >= end || !isNonZeroDigitCode(line.charCodeAt(cursor))) return null; + return cursor; } export type BlockTarget = { kind: "range"; range: ParsedRange } | { kind: "bof" } | { kind: "eof" }; @@ -208,103 +230,86 @@ interface TargetScan { nextIndex: number; } -function scanBlockTarget(line: string, end = trimEndIndex(line)): TargetScan | null { - const targetStart = skipDecoratedAnchorPrefix(line, end); - if (startsWithWord(line, targetStart, end, "BOF")) { - const nextIndex = skipBofEofRangeSuffix(line, targetStart + 3, end); - return { target: { kind: "bof" }, nextIndex }; +/** + * Scan the anchor portion of a hunk header. Accepts `BOF`, `EOF`, `A B` + * (range), or `A` (single-line shorthand for `A A`). + */ +function scanHunkAnchor(line: string, start: number, end: number): TargetScan | null { + const cursor = skipWhitespace(line, start, end); + if (line.startsWith(BOF_ANCHOR, cursor)) { + return { target: { kind: "bof" }, nextIndex: skipWhitespace(line, cursor + BOF_ANCHOR.length, end) }; } - if (startsWithWord(line, targetStart, end, "EOF")) { - const nextIndex = skipBofEofRangeSuffix(line, targetStart + 3, end); - return { target: { kind: "eof" }, nextIndex }; + if (line.startsWith(EOF_ANCHOR, cursor)) { + return { target: { kind: "eof" }, nextIndex: skipWhitespace(line, cursor + EOF_ANCHOR.length, end) }; } - - const range = scanRange(line, end); - return range === null ? null : { target: { kind: "range", range: range.range }, nextIndex: range.nextIndex }; + const range = scanHeaderRange(line, cursor, end); + if (range === null) return null; + return { target: { kind: "range", range: range.range }, nextIndex: range.nextIndex }; } -// Models sometimes write `BOF-BOF:`, `EOF-EOF:`, or even `BOF-EOF:` by analogy -// with the numeric `A-B:` form. The range portion carries no information for -// virtual anchors (they do not span lines), so we just consume and discard it. -function skipBofEofRangeSuffix(line: string, index: number, end: number): number { - const cursor = skipWhitespace(line, index, end); - if (cursor >= end || line.charCodeAt(cursor) !== CHAR_HYPHEN) return cursor; - const afterHyphen = skipWhitespace(line, cursor + 1, end); - if (startsWithWord(line, afterHyphen, end, "BOF") || startsWithWord(line, afterHyphen, end, "EOF")) { - return skipWhitespace(line, afterHyphen + 3, end); - } - return cursor; -} - -interface ParsedBlockOp { +interface ParsedHunkHeader { target: BlockTarget; - inlineBody: string | undefined; - deleteSuffix: boolean; } -function tryParseBlockOp(line: string): ParsedBlockOp | null { +/** + * Parse a bare hunk-header line: `A B` (range), `A` (single-line shorthand + * for `A A`), or the keywords `BOF` / `EOF`. Returns `null` for lines that + * do not match the shape. + */ +function tryParseHunkHeader(line: string): ParsedHunkHeader | null { const end = trimEndIndex(line); - const target = scanBlockTarget(line, end); - if (target === null) return null; - - const opIndex = skipWhitespace(line, target.nextIndex, end); - if (opIndex >= end || line.charCodeAt(opIndex) !== CHAR_OP_REPLACE) return null; - - if ( - opIndex === target.nextIndex && - line.startsWith(HL_OP_DELETE_SUFFIX, opIndex) && - opIndex + HL_OP_DELETE_SUFFIX.length === end - ) { - return { target: target.target, inlineBody: undefined, deleteSuffix: true }; - } - - const inlineStart = opIndex + HL_OP_REPLACE.length; - return { - target: target.target, - inlineBody: skipWhitespace(line, inlineStart, end) === end ? undefined : line.slice(inlineStart, end), - deleteSuffix: false, - }; + const start = skipWhitespace(line, 0, end); + if (start >= end) return null; + const scan = scanHunkAnchor(line, start, end); + if (scan === null) return null; + if (scan.nextIndex !== end) return null; + return { target: scan.target }; } +/** + * Parse a `&A,B` repeat payload row (or `&A` shorthand for `&A,A`). Returns + * `null` when the line does not match. + */ function tryParseRepeatPayload(line: string): ParsedRange | null { const end = trimEndIndex(line); if (line.length === 0 || line.charCodeAt(0) !== CHAR_PAYLOAD_REPEAT) return null; const start = scanLineNumber(line, 1, end); if (start === null) return null; - // Canonical form is `^A-B`; the explicit `^A-A` is preferred. The bare - // single-line `^A` is also accepted as `^A-A` because the strict form - // adds friction without disambiguating anything. if (start.nextIndex === end) { + // `&A` shorthand → `&A,A`. return { start: { line: start.line }, end: { line: start.line } }; } - if (start.nextIndex >= end || line.charCodeAt(start.nextIndex) !== CHAR_HYPHEN) return null; + if ( + start.nextIndex + 1 >= end || + line.charCodeAt(start.nextIndex) !== CHAR_DOT || + line.charCodeAt(start.nextIndex + 1) !== CHAR_DOT + ) + return null; - const finish = scanLineNumber(line, start.nextIndex + 1, end); + const finish = scanLineNumber(line, start.nextIndex + 2, end); if (finish === null) return null; if (skipWhitespace(line, finish.nextIndex, end) !== end) return null; return { start: { line: start.line }, end: { line: finish.line } }; } /** - * Strict header scan: `¶+` prefix, optional whitespace, path body that - * excludes whitespace, `#`, and `¶`, optional three-hex hash suffix, - * optional trailing whitespace. Returns `null` when any byte deviates from - * the shape. + * Parse a `¶PATH[#hash]` file-header line. Returns `null` for lines that + * do not start with the file prefix or that fail the strict shape. + * + * `*** Begin Patch` / `*** End Patch` / `*** Abort` markers are matched + * earlier in {@link classifyLine}, so envelope markers never reach here. */ function tryParseHeader(line: string): { path: string; fileHash?: string } | null { + if (!line.startsWith(HL_FILE_PREFIX)) return null; const end = trimEndIndex(line); - if (end === 0 || line.charCodeAt(0) !== CHAR_PILCROW) return null; - - let index = 0; - while (index < end && line.charCodeAt(index) === CHAR_PILCROW) index++; - index = skipWhitespace(line, index, end); + let index = FILE_PREFIX_LENGTH; if (index >= end) return null; const pathStart = index; while (index < end) { const code = line.charCodeAt(index); - if (code === CHAR_HASH || code === CHAR_PILCROW || code === CHAR_SPACE || code === CHAR_TAB) break; + if (code === CHAR_HASH || code === CHAR_SPACE || code === CHAR_TAB) break; index++; } if (index === pathStart) return null; @@ -339,7 +344,7 @@ export type Token = | (TokenBase & { kind: "envelope-end" }) | (TokenBase & { kind: "abort" }) | (TokenBase & { kind: "header"; path: string; fileHash?: string }) - | (TokenBase & { kind: "op-block"; target: BlockTarget; inlineBody: string | undefined; deleteSuffix: boolean }) + | (TokenBase & { kind: "op-block"; target: BlockTarget }) | (TokenBase & { kind: "payload-literal"; text: string }) | (TokenBase & { kind: "payload-repeat"; range: ParsedRange }) | (TokenBase & { kind: "raw"; text: string }); @@ -350,7 +355,9 @@ function classifyLine(line: string, lineNum: number): Token { if (markerLineEquals(line, END_PATCH_MARKER)) return { kind: "envelope-end", lineNum }; if (markerLineEquals(line, ABORT_MARKER)) return { kind: "abort", lineNum }; - if (line.charCodeAt(0) === CHAR_PILCROW) { + const firstCode = line.charCodeAt(0); + + if (line.startsWith(HL_FILE_PREFIX)) { const header = tryParseHeader(line); if (header !== null) { return header.fileHash !== undefined @@ -359,7 +366,16 @@ function classifyLine(line: string, lineNum: number): Token { } } - const firstCode = line.charCodeAt(0); + // Hunk header lines start with a digit (range / single-line) or the + // keyword `BOF` / `EOF`. `@@`-bracketed forms are intentionally NOT + // accepted here — they fall through to `raw` and the parser rejects + // them as apply_patch contamination. + const isHunkLead = isNonZeroDigitCode(firstCode) || line.startsWith(BOF_ANCHOR) || line.startsWith(EOF_ANCHOR); + if (isHunkLead) { + const hunk = tryParseHunkHeader(line); + if (hunk !== null) return { kind: "op-block", lineNum, target: hunk.target }; + } + if (firstCode === CHAR_PAYLOAD_REPLACE) { return { kind: "payload-literal", lineNum, text: line.slice(1) }; } @@ -368,17 +384,6 @@ function classifyLine(line: string, lineNum: number): Token { if (range !== null) return { kind: "payload-repeat", lineNum, range }; } - const op = tryParseBlockOp(line); - if (op !== null) { - return { - kind: "op-block", - lineNum, - target: op.target, - inlineBody: op.inlineBody, - deleteSuffix: op.deleteSuffix, - }; - } - return { kind: "raw", lineNum, text: line }; } @@ -447,7 +452,7 @@ export class Tokenizer { } isOp(line: string): boolean { - return tryParseBlockOp(line) !== null; + return tryParseHunkHeader(line) !== null; } isHeader(line: string): boolean { diff --git a/packages/hashline/src/types.ts b/packages/hashline/src/types.ts index 0262abb26..57b2ed90a 100644 --- a/packages/hashline/src/types.ts +++ b/packages/hashline/src/types.ts @@ -44,20 +44,10 @@ export interface ApplyResult { text: string; /** First line number (1-indexed) that changed, or `undefined` for a no-op apply. */ firstChangedLine?: number; - /** Diagnostic warnings collected by the applier (auto-absorb, boundary checks, …). */ + /** Diagnostic warnings collected by the parser, patcher, or recovery. */ warnings?: string[]; } -/** Optional knobs forwarded to {@link Edit} application. */ -export interface ApplyOptions { - /** - * When `true`, pure-insert and single-line replacement-boundary duplicates - * are dropped opportunistically. Default `false`: only multi-line block - * duplicates and structural-boundary single lines are absorbed. - */ - autoDropPureInsertDuplicates?: boolean; -} - /** A parsed `[A..B]` line range. */ export interface ParsedRange { start: Anchor; diff --git a/packages/hashline/test/format-v2.test.ts b/packages/hashline/test/format-v2.test.ts index b78036e9e..739e08e1f 100644 --- a/packages/hashline/test/format-v2.test.ts +++ b/packages/hashline/test/format-v2.test.ts @@ -8,79 +8,75 @@ function applyPatch(text: string, diff: string): string { describe("hashline format v2", () => { it("emits literal and repeat body rows in textual order", () => { const text = "a\nb\nc"; - const diff = ["2-2:", "+before", "^1-2", "+after"].join("\n"); + const diff = ["2 2", "+before", "&1..2", "+after"].join("\n"); expect(applyPatch(text, diff)).toBe("a\nbefore\na\nb\nafter\nc"); }); it("repeats a single source line with explicit A-A syntax", () => { const text = "a\nb\nc"; - const diff = ["2-2:", "^3-3"].join("\n"); + const diff = ["2 2", "&3..3"].join("\n"); expect(applyPatch(text, diff)).toBe("a\nc\nc"); }); it("keeps the file unchanged when a repeat covers the anchored range", () => { const text = "a\nb\nc\nd"; - const diff = ["2-3:", "^2-3"].join("\n"); + const diff = ["2 3", "&2..3"].join("\n"); expect(applyPatch(text, diff)).toBe(text); }); - it("deletes a concrete range with inline delete", () => { + it("deletes a concrete range via an empty hunk body", () => { const text = "a\nb\nc\nd"; - - expect(applyPatch(text, "2-3:-")).toBe("a\nd"); + expect(applyPatch(text, "2 3")).toBe("a\nd"); }); - it("rejects body rows after inline delete", () => { - expect(() => parsePatch("2-2:-\n+x")).toThrow(/payload line has no preceding/); - }); - - it("treats an empty concrete block as a blank-line replacement", () => { + it("empty body at a concrete range deletes the range (no blank-line insertion)", () => { const text = "a\nb\nc"; - - expect(applyPatch(text, "2-2:")).toBe("a\n\nc"); + expect(applyPatch(text, "2 2")).toBe("a\nc"); }); - it("treats empty BOF and EOF blocks as one blank-line insert", () => { + it("empty body at BOF/EOF is a no-op (nothing inserted)", () => { const text = "a\nb"; - - expect(applyPatch(text, "BOF:")).toBe("\na\nb"); - expect(applyPatch(text, "EOF:")).toBe("a\nb\n"); + expect(applyPatch(text, "BOF")).toBe(text); + expect(applyPatch(text, "EOF")).toBe(text); }); it("accepts `^A` repeat shorthand as `^A-A`", () => { const text = "a\nb\nc"; // `^A` mirrors `^A-A`; we use it to keep line 2 unchanged while // also targeting it. - expect(applyPatch(text, "2-2:\n^2")).toBe(text); + expect(applyPatch(text, "2 2\n&2")).toBe(text); }); it("auto-pipes bare body rows (legacy sigils flow through as literal text)", () => { // `↑`/`↓` are no longer reserved sigils; bare body rows are // auto-prefixed with `|` as plain literal text. const text = "a\nb\nc"; - expect(applyPatch(text, "2-2:\n↑x")).toBe("a\n↑x\nc"); - expect(applyPatch(text, "2-2:\n↓x")).toBe("a\n↓x\nc"); + expect(applyPatch(text, "2 2\n↑x")).toBe("a\n↑x\nc"); + expect(applyPatch(text, "2 2\n↓x")).toBe("a\n↓x\nc"); // And the warning is surfaced. - const { warnings } = parsePatch("2-2:\n↑x"); + const { warnings } = parsePatch("2 2\n↑x"); expect(warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); - it("rejects removed standalone delete rows through the normal op diagnostic", () => { - expect(() => parsePatch("-5")).toThrow(/unrecognized hashline block/); - expect(() => parsePatch("-5..7")).toThrow(/unrecognized hashline block/); + it("accepts `-A` and `-A..B` as standalone delete ops", () => { + // `-A..B` (and `-A` shorthand) on its own line is the canonical + // delete op in the new grammar. + const text = "a\nb\nc\nd\ne\nf\ng"; + expect(applyPatch(text, "5 5")).toBe("a\nb\nc\nd\nf\ng"); + expect(applyPatch(text, "5 7")).toBe("a\nb\nc\nd"); }); it("validates repeat ranges against file bounds", () => { - const edits = parsePatch("1-1:\n^4-4").edits; + const edits = parsePatch("1 1\n&4..4").edits; expect(() => applyEdits("a\nb", edits)).toThrow(/Line 4 does not exist/); }); it("does not flush a streaming pending empty block", () => { - const result = parsePatchStreaming("5-5:\n"); + const result = parsePatchStreaming("5 5\n"); expect(result.edits).toEqual([]); }); diff --git a/packages/hashline/test/leniency.test.ts b/packages/hashline/test/leniency.test.ts index 94c2b5941..bafe7b71e 100644 --- a/packages/hashline/test/leniency.test.ts +++ b/packages/hashline/test/leniency.test.ts @@ -7,31 +7,47 @@ function applyPatch(text: string, diff: string): string { const FILE = "a\nb\nc\nd\ne"; -describe("hashline leniency L1 — bare `A:` shorthand", () => { - it("treats `A:` as `A-A:`", () => { - expect(applyPatch(FILE, "2:\n+B")).toBe("a\nB\nc\nd\ne"); +describe("hashline core — hunk header shorthand", () => { + it("accepts `A` as `A..A` (single-line shorthand)", () => { + expect(applyPatch(FILE, "2\n+B")).toBe("a\nB\nc\nd\ne"); }); - it("treats `A:-` as `A-A:-`", () => { - expect(applyPatch(FILE, "2:-")).toBe("a\nc\nd\ne"); + it("an empty `A..A` deletes the line", () => { + expect(applyPatch(FILE, "2 2")).toBe("a\nc\nd\ne"); }); - it("preserves the inline-payload rejection on `A:content`", () => { - expect(() => parsePatch("2:hello")).toThrow(/Inline payload on the anchor line is rejected/); + it("accepts hyphen as a range separator (`A-B`)", () => { + // Models reflexively type `301-314` when copying a `read` range. + expect(applyPatch(FILE, "2-3\n+X")).toBe("a\nX\nd\ne"); }); - it("still rejects `LINE:content` rows pasted in the middle of a payload", () => { - // First block is fine; the second line looks like another op-block - // with inline payload (after L1 it parses as `3-3:` with body - // ` ddd`), which triggers the inline-payload diagnostic. - expect(() => parsePatch("2-2:\n+first\n3: ddd")).toThrow(/Inline payload on the anchor line is rejected/); + it("accepts `..` as a range separator (`A..B`)", () => { + expect(applyPatch(FILE, "2..3\n+X")).toBe("a\nX\nd\ne"); + }); + + it("accepts unicode ellipsis as a range separator (`A…B`)", () => { + expect(applyPatch(FILE, "2\u20263\n+X")).toBe("a\nX\nd\ne"); + }); + + it("tolerates whitespace around the separator (`A - B`)", () => { + expect(applyPatch(FILE, "2 - 3\n+X")).toBe("a\nX\nd\ne"); + }); + + it("rejects `LINE=content` rows pasted from old format as orphan payload", () => { + expect(() => parsePatch("2=hello")).toThrow(/payload line has no preceding hunk header/); + }); + + it("auto-pipes mid-payload bare rows in a mixed block (was previously rejected)", () => { + const result = parsePatch("2 2\n+first\n3= ddd"); + expect(applyEdits(FILE, result.edits).text).toBe("a\nfirst\n3= ddd\nc\nd\ne"); + expect(result.warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); }); describe("hashline leniency L2 — bare `^A` repeat shorthand", () => { it("treats `^A` as `^A-A`", () => { // `^2-2` keeps the original line 2 between the inserted rows. - expect(applyPatch(FILE, "2-2:\n+ABOVE\n^2\n+BELOW")).toBe("a\nABOVE\nb\nBELOW\nc\nd\ne"); + expect(applyPatch(FILE, "2 2\n+ABOVE\n&2\n+BELOW")).toBe("a\nABOVE\nb\nBELOW\nc\nd\ne"); }); it("auto-pipes `^A-` (malformed range) as literal text via L3", () => { @@ -39,224 +55,197 @@ describe("hashline leniency L2 — bare `^A` repeat shorthand", () => { // tokenizer classifies it as raw; L3's uniformly-bare auto-pipe // then folds it back into the block as a literal. The model sees // the warning and can re-issue with a well-formed repeat. - const result = parsePatch("2-2:\n^2-"); - expect(applyEdits(FILE, result.edits).text).toBe("a\n^2-\nc\nd\ne"); + const result = parsePatch("2 2\n&2-"); + expect(applyEdits(FILE, result.edits).text).toBe("a\n&2-\nc\nd\ne"); expect(result.warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); }); describe("hashline leniency L3 — auto-pipe uniformly bare bodies", () => { it("accepts a block whose body is uniformly unprefixed", () => { - const result = parsePatch("2-2:\n hello\n world"); + const result = parsePatch("2 2\n hello\n world"); expect(applyEdits(FILE, result.edits).text).toBe("a\n hello\n world\nc\nd\ne"); expect(result.warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); - it("rejects literal-then-bare mixed blocks at the bare row's line", () => { - expect(() => parsePatch("2-2:\n+first\nsecond")).toThrow(/line 3: payload row in a hashline block/); + it("auto-pipes a bare row after a `+` row (was previously rejected)", () => { + const result = parsePatch("2 2\n+first\nsecond"); + expect(applyEdits(FILE, result.edits).text).toBe("a\nfirst\nsecond\nc\nd\ne"); + expect(result.warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); - it("rejects bare-then-literal mixed blocks at the bare row's line", () => { - // The bare `first` is buffered on line 2; on line 3 the `|second` - // arrives and triggers a retro-rejection pointed at line 2. - expect(() => parsePatch("2-2:\nfirst\n+second")).toThrow(/line 2: payload row in a hashline block/); + it("auto-pipes a bare row before a `+` row (was previously rejected)", () => { + // `first` is buffered. When `+second` arrives, we auto-pipe both rows. + const result = parsePatch("2 2\nfirst\n+second"); + expect(applyEdits(FILE, result.edits).text).toBe("a\nfirst\nsecond\nc\nd\ne"); + expect(result.warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); it("does NOT auto-pipe across block boundaries", () => { - // `2-2:` accumulates `foo` as a bare row; `4-4:` flushes the first + // `2 2` accumulates `foo` as a bare row; `4 4` flushes the first // block (auto-pipe fires) and starts a new pending. The second // block's `bar` row is also bare → second auto-pipe. - const result = parsePatch("2-2:\nfoo\n4-4:\nbar"); + const result = parsePatch("2 2\nfoo\n4 4\nbar"); expect(applyEdits(FILE, result.edits).text).toBe("a\nfoo\nc\nbar\ne"); }); }); -describe("hashline leniency L4 — lone `-` body row as delete", () => { - it("retroactively converts a lone `-` row to a `:-` delete", () => { - const result = parsePatch("2-3:\n-"); - expect(applyEdits(FILE, result.edits).text).toBe("a\nd\ne"); - expect(result.warnings.some(w => /Converted a lone `-` body row/.test(w))).toBe(true); +describe("hashline leniency L9 — unified-diff body conversion", () => { + it("drops `-`-prefixed body rows (already deleted by the hunk range)", () => { + // Classic apply_patch / unified-diff shape: -old / +new pair. + // Model expects the `-` row to mark line for deletion; hashline's + // `A..B` already deletes the range, so we drop the `-` row + // and keep the `+` row. + const result = parsePatch("2 2\n-original line\n+replacement"); + expect(applyEdits(FILE, result.edits).text).toBe("a\nreplacement\nc\nd\ne"); + expect(result.warnings.some(w => /Hunk body contained unified-diff-style rows/.test(w))).toBe(true); }); - it("does NOT fire when the block already has `|` rows", () => { - // L4 only triggers on a totally-bare pending block; once a literal - // row arrives, the `-` is a regular mixed-block bare row → reject. - expect(() => parsePatch("2-2:\n+X\n-")).toThrow(/payload row in a hashline block must start with/); + it("strips the unified-diff metadata-space from context rows once a `-` row is seen", () => { + // Body has a context row ` keep this`, a `-old` row, a `+new` row. + // Result: lines 2..3 replaced with [keep this, new]. + const text = "a\nb\nc\nd"; + const result = parsePatch("2 3\n keep this\n-original\n+new"); + expect(applyEdits(text, result.edits).text).toBe("a\nkeep this\nnew\nd"); }); - it("does NOT fire when the block already has bare raw rows", () => { - // `foo` is bare and buffered. The `-` row arrives next; pendingRaws - // is non-empty so L4 does not retro-convert. The block is uniformly - // bare so it auto-pipes both rows (so `-` becomes literal text). - const result = parsePatch("2-2:\nfoo\n-"); - expect(applyEdits(FILE, result.edits).text).toBe("a\nfoo\n-\nc\nd\ne"); + it("retroactively strips the metadata-space from context rows that arrived BEFORE the `-` row", () => { + // Streaming order: context first, then `-`. The `-` is what tells us + // we are in unified-diff mode; we must go back and strip the space + // from the context row. + const text = "a\nb\nc\nd"; + const result = parsePatch("2 3\n keep this\n+new\n-original"); + expect(applyEdits(text, result.edits).text).toBe("a\nkeep this\nnew\nd"); }); }); describe("hashline leniency L5 — overlapping bare/concrete coalesce", () => { - it("coalesces `A-B:` + `A-B:-` (identical-range before-then-delete) into a delete", () => { - const result = parsePatch("2-3:\n2-3:-"); - expect(applyEdits(FILE, result.edits).text).toBe("a\nd\ne"); - expect(result.warnings.some(w => /overlapping bare hashline block/.test(w))).toBe(true); + it("coalesces two identical-range hunks (last-wins)", () => { + // Two `2 3` hunks back-to-back. The first has no body, the + // second has a payload. We drop the first and emit only the second. + const result = parsePatch("2 3\n2 3\n+X"); + expect(applyEdits(FILE, result.edits).text).toBe("a\nX\nd\ne"); + expect(result.warnings.some(w => /identical-range hashline hunks/.test(w))).toBe(true); }); - it("coalesces an overlapping (not identical) bare anchor followed by `:-`", () => { - // Bare `2-3:` overlaps with the later `3-4:-`. Drop the bare - // pending (which would have replaced 2-3 with a blank), emit only - // the deletes for 3-4. - const result = parsePatch("2-3:\n3-4:-"); - expect(applyEdits(FILE, result.edits).text).toBe("a\nb\ne"); - expect(result.warnings.some(w => /overlapping bare hashline block/.test(w))).toBe(true); - }); - - it("coalesces an overlapping bare anchor followed by a concrete replace", () => { - // Bare `2-3:` overlaps with the concrete `3-4:` (has payload). - // Drop the bare pending, keep the concrete one. - const result = parsePatch("2-3:\n3-4:\n+NEW"); + it("coalesces an overlapping bare hunk followed by a concrete hunk", () => { + // Bare `2 3` overlaps with the concrete `3 4`. Drop the + // bare pending; keep the concrete one. + const result = parsePatch("2 3\n3 4\n+NEW"); expect(applyEdits(FILE, result.edits).text).toBe("a\nb\nNEW\ne"); - expect(result.warnings.some(w => /overlapping bare hashline block/.test(w))).toBe(true); + expect(result.warnings.some(w => /overlapping bare hashline hunk/.test(w))).toBe(true); }); it("still rejects two concrete overlapping replaces", () => { - // Both pending blocks have payload → no L5 short-circuit. The + // Both pending hunks have payload → no L5 short-circuit. The // post-hoc validator catches the line-3 collision. - expect(() => parsePatch("2-3:\n+X\n+Y\n3-4:\n+Z")).toThrow(/anchor line 3 is already targeted by another op/); + expect(() => parsePatch("2 3\n+X\n+Y\n3 4\n+Z")).toThrow(/anchor line 3 is already targeted by another hunk/); }); }); -describe("hashline leniency L6 — stacked blank-body `A-A:` warning", () => { - it("warns once when two or more consecutive `A-A:` blocks have empty bodies", () => { - const result = parsePatch("2-2:\n3-3:"); - // Both lines became blank. - expect(applyEdits(FILE, result.edits).text).toBe("a\n\n\nd\ne"); - expect(result.warnings.some(w => /run of single-line empty-body blocks/.test(w))).toBe(true); +describe("hashline — apply_patch / unified-diff contamination", () => { + it("rejects `*** Update File:` sentinels as contamination", () => { + expect(() => parsePatch("*** Update File: a.ts\n2 2\n+X")).toThrow(/apply_patch sentinel/); }); - it("does NOT warn when only one blank `A-A:` block exists", () => { - const result = parsePatch("2-2:"); - expect(result.warnings.some(w => /run of single-line empty-body blocks/.test(w))).toBe(false); + it("rejects `*** Add File:` sentinels as contamination", () => { + expect(() => parsePatch("*** Add File: a.ts\n2 2\n+X")).toThrow(/apply_patch sentinel/); }); - it("does NOT warn when blank blocks are interleaved with non-blank blocks", () => { - const result = parsePatch("2-2:\n3-3:\n+X"); - // Run interrupted on the second block; counter resets. - expect(result.warnings.some(w => /run of single-line empty-body blocks/.test(w))).toBe(false); - }); -}); - -describe("hashline leniency L8 — apply_patch / unified-diff contamination", () => { - it("rejects `*** Update File:` sentinels at top level", () => { - expect(() => parsePatch("*** Update File: a.ts\n2-2:\n+X")).toThrow(/apply_patch sentinel/); - }); - - it("rejects `*** Add File:` sentinels", () => { - expect(() => parsePatch("*** Add File: a.ts\n2-2:\n+X")).toThrow(/apply_patch sentinel/); - }); - - it("rejects unified-diff hunk headers (`@@`)", () => { - expect(() => parsePatch("@@ -1,3 +1,3 @@\n2-2:\n+X")).toThrow(/unified-diff hunk header/); - expect(() => parsePatch("@@\n2-2:\n+X")).toThrow(/unified-diff hunk header/); - }); - - it("rejects `-N-M:` / `-N:` apply_patch hunk anchors", () => { - // `+`-prefixed shapes are no longer flagged because `+` is now the - // canonical payload sigil; `+2-2:` tokenizes as a literal payload - // row containing `2-2:` and either lands inside a pending block or - // throws the standard orphan-payload error. - expect(() => parsePatch("-2-3:\n+X")).toThrow(/apply_patch line prefix/); - expect(() => parsePatch("-2:\n+X")).toThrow(/apply_patch line prefix/); + it("rejects unified-diff hunk headers (`-N,M +N,M`) as contamination", () => { + expect(() => parsePatch("@@ -1,3 +1,3 @@\n2 2\n+X")).toThrow(/unified-diff hunk header/); }); it("treats top-level `+TEXT` as an orphan literal payload", () => { - // `+` is the payload sigil — at top level (no pending anchor) this - // surfaces the standard "no preceding A-B:" error rather than the - // apply_patch-specific one, so the model still gets a clear pointer - // to add an anchor above the body row. - expect(() => parsePatch("+ const X = 1;\n2-2:")).toThrow( - /payload line has no preceding A-B:, BOF:, or EOF: anchor/, - ); - }); - - it("keeps `-N` bare delete rejection on the legacy unrecognized-block diagnostic", () => { - // `-5` and `-5..7` are the legacy "removed standalone delete row" - // shapes — they still throw with the existing unrecognized-block - // diagnostic, not the apply_patch one. - expect(() => parsePatch("-5")).toThrow(/unrecognized hashline block/); - expect(() => parsePatch("-5..7")).toThrow(/unrecognized hashline block/); - }); - - it("gives a focused message for a lone `-` outside any pending block", () => { - expect(() => parsePatch("-")).toThrow(/a lone "-" is not a valid hashline op/); + expect(() => parsePatch("+ const X = 1;\n2 2")).toThrow(/payload line has no preceding hunk header/); }); }); describe("hashline leniency — composite scenarios from the benchmark dumps", () => { - it("recovers GLM's `LINE:` paste + bare body (chat-simple.ts shape)", () => { + it("recovers GLM's `LINE=`-shaped paste + bare body (chat-simple.ts shape)", () => { const text = "aaa\nbbb\nccc\nddd"; - // Authored: bare `2:` anchor followed by a uniformly-bare body - // pasted from `read` output. L1 promotes `2:` to `2-2:`; L3 + // Authored: bare `2 2` anchor followed by a uniformly-bare body + // pasted from `read` output. L1 promotes `2 2` to `2 2`; L3 // auto-pipes the bare body rows. - const result = parsePatch("2:\n NEW_LINE_ONE\n NEW_LINE_TWO"); + const result = parsePatch("2 2\n NEW_LINE_ONE\n NEW_LINE_TWO"); expect(applyEdits(text, result.edits).text).toBe("aaa\n NEW_LINE_ONE\n NEW_LINE_TWO\nccc\nddd"); expect(result.warnings.some(w => /Auto-prefixed bare body row/.test(w))).toBe(true); }); - it("recovers gpt-5-spark's identical-range `89-90: ⏎ 89-90:-` shape", () => { + it("two back-to-back identical-range hunks coalesce last-wins", () => { const text = "aaa\nbbb\nccc\nddd"; - // Identical-range before-then-delete: should delete lines 2-3. - const result = parsePatch("2-3:\n2-3:-"); + // Two `2 3` hunks; the first has no body, the second is the + // "real" deletion. The first should be dropped via the identical- + // range coalesce, leaving the deletion to fire. + const result = parsePatch("2 3\n2 3"); expect(applyEdits(text, result.edits).text).toBe("aaa\nddd"); expect(result.warnings.length).toBeGreaterThan(0); }); - it("recovers gpt-5-spark's `+^A-B` shape (model prefixed a repeat with +)", () => { + it("recovers gpt-5-spark's `+&A..B` shape (model prefixed a repeat with +)", () => { const text = "aaa\nbbb\nccc"; - // Authored: `2-2: +NEW +^2-2`. The second body row is a repeat row + // Authored: `2-2: +NEW +&2..2`. The second body row is a repeat row // the model mistakenly prefixed with `+`. It should be silently // rerouted as `^2-2` so the patch effectively inserts NEW above // the original line 2, with a warning. - const result = parsePatch("2-2:\n+NEW\n+^2-2"); + const result = parsePatch("2 2\n+NEW\n+&2..2"); expect(applyEdits(text, result.edits).text).toBe("aaa\nNEW\nbbb\nccc"); - expect(result.warnings.some(w => /A body row started with `\+\^A-B`/.test(w))).toBe(true); + expect(result.warnings.some(w => /A body row started with `\+&A\.\.B`/.test(w))).toBe(true); }); - it("accepts `+^A-B` with leading whitespace inside the literal text", () => { + it("accepts `+&A..B` with leading whitespace inside the literal text", () => { // gpt-5-spark / chat-simple.ts shape: `+ ^85-85` — the model // added indentation between `+` and `^A-B`. We trim before checking. const text = "aaa\nbbb\nccc"; - const result = parsePatch("2-2:\n+NEW\n+ ^2-2"); + const result = parsePatch("2 2\n+NEW\n+ &2..2"); expect(applyEdits(text, result.edits).text).toBe("aaa\nNEW\nbbb\nccc"); - expect(result.warnings.some(w => /A body row started with `\+\^A-B`/.test(w))).toBe(true); + expect(result.warnings.some(w => /A body row started with `\+&A\.\.B`/.test(w))).toBe(true); }); it("accepts `+^A` shorthand (single line)", () => { const text = "aaa\nbbb\nccc"; - const result = parsePatch("2-2:\n+NEW\n+^2"); + const result = parsePatch("2 2\n+NEW\n+&2"); expect(applyEdits(text, result.edits).text).toBe("aaa\nNEW\nbbb\nccc"); - expect(result.warnings.some(w => /A body row started with `\+\^A-B`/.test(w))).toBe(true); + expect(result.warnings.some(w => /A body row started with `\+&A\.\.B`/.test(w))).toBe(true); }); it("does NOT misclassify `+^literal-text` (not a valid repeat shape)", () => { - // `+^hello` is just a literal payload row whose text is `^hello`. + // `+&hello` is just a literal payload row whose text is `^hello`. // No range follows the `^`, so it's not a repeat — emit the literal // as-is, no warning. const text = "aaa\nbbb\nccc"; - const result = parsePatch("2-2:\n+^hello"); - expect(applyEdits(text, result.edits).text).toBe("aaa\n^hello\nccc"); - expect(result.warnings.some(w => /A body row started with `\+\^A-B`/.test(w))).toBe(false); + const result = parsePatch("2 2\n+&hello"); + expect(applyEdits(text, result.edits).text).toBe("aaa\n&hello\nccc"); + expect(result.warnings.some(w => /A body row started with `\+&A\.\.B`/.test(w))).toBe(false); }); }); describe("hashline leniency — BOF/EOF range suffix", () => { - it("accepts `BOF-BOF:` as `BOF:`", () => { - expect(applyPatch(FILE, "BOF-BOF:\n+HEAD")).toBe("HEAD\na\nb\nc\nd\ne"); + it("accepts `BOF..BOF=` as `BOF`", () => { + expect(applyPatch(FILE, "BOF\n+HEAD")).toBe("HEAD\na\nb\nc\nd\ne"); }); - it("accepts `EOF-EOF:` as `EOF:`", () => { - expect(applyPatch(FILE, "EOF-EOF:\n+TAIL")).toBe("a\nb\nc\nd\ne\nTAIL"); + it("accepts `EOF..EOF=` as `EOF`", () => { + expect(applyPatch(FILE, "EOF\n+TAIL")).toBe("a\nb\nc\nd\ne\nTAIL"); }); - it("accepts `BOF-EOF:` (degenerate but harmless)", () => { - expect(applyPatch(FILE, "BOF-EOF:\n+HEAD")).toBe("HEAD\na\nb\nc\nd\ne"); + it("accepts `BOF..EOF=` (degenerate but harmless)", () => { + expect(applyPatch(FILE, "BOF\n+HEAD")).toBe("HEAD\na\nb\nc\nd\ne"); + }); +}); + +describe("hashline apply — duplicate boundary payloads", () => { + it("keeps replacement boundary echoes literal", () => { + const text = ["// one", "// two", "old();"].join("\n"); + const diff = "3 3\n+// one\n+// two\n+new();"; + + expect(applyPatch(text, diff)).toBe(["// one", "// two", "// one", "// two", "new();"].join("\n")); + }); + + it("keeps pure-insert context echoes literal", () => { + const text = ["aaa", "bbb", "ccc"].join("\n"); + const diff = "EOF\n+bbb\n+ccc\n+NEW"; + + expect(applyPatch(text, diff)).toBe("aaa\nbbb\nccc\nbbb\nccc\nNEW"); }); }); diff --git a/packages/hashline/test/patcher.test.ts b/packages/hashline/test/patcher.test.ts index 7139099b3..2deff54c9 100644 --- a/packages/hashline/test/patcher.test.ts +++ b/packages/hashline/test/patcher.test.ts @@ -17,7 +17,7 @@ describe("Patcher snapshot tag integrity", () => { const tag = snapshots.recordContiguous(PATH, 1, ["before", ""], { fullText: "before\n" }); const patcher = new Patcher({ fs, snapshots }); - const result = await patcher.apply(Patch.parse(`¶${PATH}#${tag}\n1-1:\n+after`)); + const result = await patcher.apply(Patch.parse(`¶${PATH}#${tag}\n1 1\n+after`)); expect(result.sections[0]?.op).toBe("update"); expect(result.sections[0]?.fileHash).toMatch(/^[0-9A-F]{3}$/); @@ -26,7 +26,7 @@ describe("Patcher snapshot tag integrity", () => { }); it("normalizes lowercase section tags while parsing", () => { - const section = Patch.parseSingle(`¶${PATH}#0a3\n1-1:\n+after`); + const section = Patch.parseSingle(`¶${PATH}#0a3\n1 1\n+after`); expect(section.fileHash).toBe("0A3"); }); @@ -42,7 +42,7 @@ describe("Patcher snapshot tag integrity", () => { snapshots.recordContiguous(PATH, 1, [`unrelated ${index}`]); } const patcher = new Patcher({ fs, snapshots }); - const patch = Patch.parse(`¶${PATH}#${staleTag}\n1-1:\n|changed`); + const patch = Patch.parse(`¶${PATH}#${staleTag}\n1 1\n|changed`); await expect(patcher.apply(patch)).rejects.toBeInstanceOf(MismatchError); expect(fs.get(PATH)).toBe("target\n"); diff --git a/packages/hashline/test/recovery-session-chain.test.ts b/packages/hashline/test/recovery-session-chain.test.ts index e5093bcae..df3b8ccbf 100644 --- a/packages/hashline/test/recovery-session-chain.test.ts +++ b/packages/hashline/test/recovery-session-chain.test.ts @@ -33,7 +33,7 @@ describe("Recovery — session-chain replay anchor-content gate", () => { // rewrote. Replaying onto current would overwrite "L5-CHANGED" with // payload the model authored against the stale "L5". That is // corruption, not recovery. - const { edits } = parsePatch("5-5:\n|L5-MODEL"); + const { edits } = parsePatch("5 5\n|L5-MODEL"); const recovered = new Recovery(store).tryRecover({ path: PATH, @@ -51,7 +51,7 @@ describe("Recovery — session-chain replay anchor-content gate", () => { // merge fails (patch context includes the rewritten line 5), but the // replay fallback is safe because the model's anchor still names the // same logical content. - const { edits } = parsePatch("3-3:\n|L3-MODEL"); + const { edits } = parsePatch("3 3\n|L3-MODEL"); const recovered = new Recovery(store).tryRecover({ path: PATH,