- Added three new benchmark run reports for GPT-5.3 Codex model with different edit variants (hashline, patch, replace).
- Recorded performance metrics across 60 tasks per variant with success rates ranging from 80-83.3%.
- Fixed incorrect path resolution for CLI module by using Bun.fileURLToPath() to convert import.meta.resolve() result to a proper file path.
- Reordered constant declarations to ensure TMP_DIR is initialized before CLI_PATH which depends on it.
- Updated scripts/repro-stuck.ts to use consistent CLI path resolution with Bun.fileURLToPath().
- Exported renderPromptTemplate function from config/prompt-templates module for programmatic prompt template rendering.
- Exported computeLineHash function from patch/hashline module for external patch utility access.
- Added ./cli export path in package.json for direct CLI module access.
- Changed hashline display format separator from pipe to two spaces for improved readability.
- Removed `lines` and `hashes` parameters from read tool in favor of automatic file display mode resolution.
- Added `resolveFileDisplayMode` utility to centralize file display mode configuration logic.
- Integrated file display mode settings into grep and read tools for consistent output formatting.
- Consolidated tool parameter types to use schema-derived types via Typebox `Static` utility.
- Updated parseLineRef to handle both legacy pipe-separator and new two-space hashline formats.
- Removed outdated benchmark reports from previous test runs across multiple models and edit variants.
- Added new benchmark reports for Claude Haiku 4.5, Claude Sonnet 4.5, Deepseek V3.2, Devstral Medium, Gemini 3 Flash, GLM-4.5-Air, GPT-5.1-Codex-Mini, GPT-5.2-Codex, Grok-4-1-Fast, Grok-4-Fast-Non-Reasoning, Grok-Code-Fast-1, Kimi K2.5, Minimax M2.1, Qwen Turbo, and Zai-GLM-4.7 models across hashline, patch, and replace edit variants.
- Updated benchmark test results with improved task success rates and edit success metrics across all tested models.
- Fixed resource leak in browser query handler by properly disposing owned proxy elements for non-winning candidates.
- Fixed script evaluation to support async functions and await expressions in browser evaluate operations.
- Removed unused maxLines parameter from buildMutationPreviewAgainstOriginal function and simplified preview generation logic.
- [ { "text": "Updated runner tool arguments from 'read,edit,write,ls' to 'read,edit,write' with '--no-skills' flag in single task execution.", "user_visible": false }, { "text": "Updated runner tool arguments from 'read,edit,write,ls' to 'read,edit,write' with '--no-skills' flag in batch execution.", "user_visible": false } ], "issue_refs">.
- Added new `replace` hashline edit operation for substring-style fuzzy text replacement without line references, with optional `all` flag for replace-all behavior.
- Added `noopEdits` array to `applyHashlineEdits` return value to report edits that produced no changes, including edit index, location, and current content for diagnostics.
- Added validation to detect and reject hashline edits using wrong-format fields (`old_text`/`new_text` from replace mode, `diff` from patch mode) with helpful error messages.
- Renamed hashline edit operation keys from `single`/`range`/`insertAfter` to `set_line`/`replace_lines`/`insert_after` for clearer semantics.
- Enhanced error messages for wrong-format hashline edits to guide users toward correct operation syntax.
- Added noopEdits array to applyHashlineEdits return type to track edits that produce no changes.
- Added validation to reject edits with wrong-format fields (old_text/new_text from replace mode, diff from patch mode) that indicate model confusion.
- Added additionalProperties tolerance to hashline edit schemas to allow flexible field handling.
- Added deduplication logic to remove duplicate edits targeting the same line(s) with identical destination content.
- Improved error handling for missing end fields in edit ranges by returning single-line specs instead of requiring both start and end.
- Enhanced no-op error recovery guidance in prompts with detailed instructions to re-read file and function context after consecutive no-op errors.
- Added `noOpRetryLimit` and `mutationScopeWindow` configuration parameters to benchmark config.
- Implemented zero-tool-call retry logic in single task runner to automatically retry when model produces no tool calls, up to the configured limit.
- Implemented zero-tool-call retry logic in batched task runner to automatically retry when model produces no tool calls, up to the configured limit.
- Removed unused `_formatEditArgs` function from report module.
- Removed preventable failure diagnostics reporting from benchmark report generation.
- Renamed hashline edit operation types from replaceLine/replaceLines to single/range for improved clarity.
- Renamed content field to replacement in hashline edit operations to better reflect its purpose.
- Enhanced no-op edit diagnostics to perform line-by-line comparisons and distinguish between identical and normalized replacements.
- Improved error messages for hashline edits to clarify differences between literal identical content and whitespace-normalized content.
- Removed timeout-retries configuration option from react-edit-benchmark CLI and simplified retry logic.
- Updated all test cases and documentation to reflect renamed hashline edit operation types and fields.
- Removed insertBefore and substr hashline edit operations, simplifying the edit API to support only replaceLine, replaceLines, and insertAfter operations.
- Updated hashline edit schema and type definitions to remove insertBefore and substr operation types from the ParsedRefs union and edit validation logic.
- Removed insertBefore and substr test cases from hashline test suite, including tests for insert-before functionality, anchor echo stripping, and substring matching.
- Updated benchmark runner to refactor edit operations from src/dst format to discriminated union types (replaceLine, replaceLines, insertAfter) and adjusted insertAfter line references.
- Added remaps property to HashlineMismatchError to store a map of stale references to their corrected versions with quick fix suggestions.
- Added diagnostic output for no-op edits that displays target lines from failed edit operations to help users understand why edits had no effect.
- Added computeLineHash and parseLineRef helper functions to public exports for use in diagnostic and validation logic.
- Enhanced error messages in HashlineMismatchError with quick fix suggestions showing stale references and their corrected versions.
- Updated hashline edit tool documentation with critical guidelines on preserving formatting, recovery procedures, and pre-submission verification checklist.
- Redesigned hashline edit API to use discriminated union variants (replaceLine, replaceLines, insertAfter, insertBefore, substr) instead of nested src/dst structure.
- Changed hash algorithm from xxHash64 hex to xxHash32 base36 encoding and increased hash length from 2 to 3 characters.
- Added substr edit variant to match and replace content by unique substring when line hashes are unavailable.
- Implemented line relocation heuristics to resolve stale line references when hash uniquely identifies a moved line.
- Moved substr needle search from edit application phase to pre-validation phase for earlier error detection.
- Migrated hashline edit type definitions from manual TypeScript interfaces to schema-derived types using Static<typeof schema> pattern.
- Added mutation preview hints to error messages when edits fail with 'No changes made' errors, showing line numbers, hashes, and added/removed lines.
- Changed formatter to pin JavaScript fixtures to the flow parser to avoid parser-dependent formatting drift.
- Added whitespace preservation logic to treat whitespace-only differences as passing in file verification.
- Removed substring source specification kind from hashline edits, requiring users to use line-hash references instead.
- Simplified SrcSpec type to use generic type parameter for line references, removing substring variant from union.
- Updated react-edit-benchmark to use --max-tasks option for deterministic task sampling and changed default batch sizes to 1.
- Migrated models export from TypeScript module to JSON format, changing the public API from importing MODELS from './models.generated' to importing from './models.json' with JSON import assertion.
- Updated @anthropic-ai/sdk dependency from ^0.72.1 to ^0.74.0.
- Simplified model generation script by replacing 49 lines of TypeScript code generation with direct JSON serialization.
- Updated @types/bun devDependency from ^1.3.8 to ^1.3.9 across all packages.
- Removed models.generated.ts exclusion from biome.json linting configuration.
- Added guided mode support with --guided and --no-guided command-line options to enable contextual editing hints.
- Added --max-attempts option (default: 2, range: 1-5) to enable retry logic for failed benchmark tasks.
- Added --require-read-tool-call option to enforce read tool invocation before edit operations.
- Implemented retry mechanism in task runner that re-attempts failed tasks up to maxAttempts with guided context on subsequent attempts.
- Extended TaskMetadata to capture file location details including filePath, fileName, lineNumber, originalSnippet, and mutatedSnippet.
- Changed success metric from patchApplied to editSucceeded to better reflect successful edit tool invocations.
- Fixed line hash computation to normalize whitespace before hashing, ensuring consistent hashes for semantically identical lines.
- Removed index parameter from hash computation to prevent hash variance based on line position.
- Changed `HashlineEdit.src` from string format (e.g., `"5:ab"`, `"5:ab..9:ef"`) to structured `SrcSpec` object with discriminated union types (`{ kind: "single", ref: "..." }`, `{ kind: "range", start: "...", end: "..." }`, `{ kind: "insertAfter", after: "..." }`, `{ kind: "insertBefore", before: "..." }`, `{ kind: "substring", needle: "..." }`).
- Refactored `parseSrc` function to `parseSrcSpec` to handle structured `SrcSpec` objects instead of string parsing, replacing string validation and delimiter parsing logic with explicit case handling for each operation kind.
- Updated tool schema in `index.ts` to define `srcSpecSchema` as a discriminated union type with five variants, replacing the previous simple string-based `src` parameter definition.
- Added substring-based source matching for hashline edits to support matching lines by content when hash references are unavailable.
- Added automatic detection and repair of single-line merges where models incorrectly combine multiple lines into one.
- Added normalization of Unicode-confusable hyphens to ASCII hyphens to handle model-generated variations.
- Added heuristics to restore indentation and preserve wrapped line formatting during edits.
- Relaxed comma validation in src to allow commas while rejecting inputs with multiple line references.
- Enhanced parseLineRef to accept shorter hash prefixes instead of requiring exact hash matches.
- Improved error messages for hash mismatches to provide more actionable guidance.
- Replaced HashlineEdit API from old/new/after fields to src/dst fields with unified range syntax supporting single lines, ranges, insert-after, and insert-before operations.
- Added parseSrc() function to parse structured line references from src strings supporting formats like '5:ab', '5:ab..9:ef', '5:ab..', and '..5:ab'.
- Added heuristics to strip anchor line echoes and range boundary echoes from model-generated replacement content.
- Implemented preserveWhitespaceOnlyLinesLoose() for improved whitespace preservation using loose matching strategy when replacement line counts don't match.
- Added comprehensive benchmark reports for Claude Sonnet 4.5, Gemini 2.5 Flash Lite, and GPT-5.1 Codex Mini models evaluating hashline edit performance across 60 tasks.
- Added MCPRequestOptions interface with signal property for request cancellation via AbortSignal.
- Added abort signal support to MCP tool execution enabling request cancellation via Escape-to-interrupt or other abort mechanisms.
- Enhanced MCP request handling with abort signal propagation through HTTP, SSE, and stdio transports with proper cleanup.
- Improved stdio transport request handling to use Promise.withResolvers for cleaner async flow and better abort signal integration.
- Updated HTTP transport to combine operation abort signals with timeout signals using AbortSignal.any() for unified cancellation.
- Modified SSE response parsing to support abort signals and distinguish between timeout and user-initiated cancellation.
- Added automatic stripping of `LINE:HASH|` display prefixes and unified-diff `+` markers from replacement content in hashline edits.
- Enhanced replace edits to preserve original whitespace on lines where only whitespace differs, preventing spurious formatting diffs when models reformat code.
- Added helper functions `stripNewLinePrefixes()`, `preserveWhitespaceOnlyLines()`, and `equalsIgnoringWhitespace()` to improve edit robustness.
- Optimized line hash computation by precomputing hex lookup table and using bitwise operations instead of string formatting.
- Simplified conditional logic in read tool for hash and line number configuration.
- Refactored react-edit-benchmark model/provider parsing to use concise ternary operator.
- Added hashline edit mode for line-addressed edits using hash-verified line references with xxHash64 integrity verification.
- Replaced `edit.patchMode` boolean setting with `edit.mode` enum supporting 'replace', 'patch', and 'hashline' modes.
- Added `readHashLines` configuration setting to include line hashes in read output for hashline edit mode.
- Implemented `computeLineHash`, `formatHashLines`, `parseLineRef`, `validateLineRef`, and `applyHashlineEdits` utility functions for hashline operations.
- Updated read tool to support optional `hashes` parameter and prioritize hash lines over line numbers when both are requested.
- Changed `getEditModelVariants()` return type from `Record<string, 'patch' | 'replace'>` to `Record<string, EditMode | null>` and removed hardcoded model-specific defaults.
- Replaced ASCII ellipsis characters (three dots '...') with Unicode ellipsis character ('...') throughout the codebase for improved typography.
- Adjusted string truncation logic to account for single-character Unicode ellipsis instead of three-character ASCII ellipsis, reducing reserved space from 3 to 1 character in truncation calculations.
- Updated truncation offsets in multiple files (session-manager, agent, executor, footer) to preserve 2 additional characters before ellipsis due to more compact Unicode representation.
- Migrated from @types/node to @types/bun across all packages to align with Bun runtime.
- Removed @types/node and bun-types from root and package-specific devDependencies.
- Updated tsconfig.base.json to remove 'node' from types array and 'DOM.AsyncIterable' from lib, keeping only Bun-compatible configurations.
- Simplified ReadableStream type annotation in openai-codex-stream.test.ts by removing explicit Uint8Array generic parameter.
- Migrated environment variable access from direct process.env to centralized getEnv() utility function across all packages.
- Renamed environment variable prefix from OMP_ to PI_ throughout codebase (e.g., OMP_CODING_AGENT_DIR -> PI_CODING_AGENT_DIR).
- Removed automatic environment variable migration from PI_ to OMP_ prefixes via migrate-env.ts module.
- Removed env setting from configuration schema and applyEnvironmentVariables() method from settings.
- Updated CI/CD build configuration to use PI_COMPILED flag instead of OMP_COMPILED.
- Changed venvPath property in PythonRuntime from nullable (string | null) to optional (string | undefined).
- Updated @anthropic-ai/sdk from ^0.71.2 to ^0.72.1 in packages/ai.
- Updated @aws-sdk/client-bedrock-runtime from ^3.975.0 to ^3.982.0 in packages/ai.
- Updated @google/genai from ^1.38.0 to ^1.39.0 in packages/ai.
- Updated @smithy/node-http-handler from ^4.4.8 to ^4.4.9 in packages/ai.
- Updated openai from ^6.16.0 to ^6.17.0 in packages/ai.
- Removed proxy-agent and undici dependencies from packages/ai.
- Extracted manual space padding logic into a reusable `padding()` utility function across all UI components and utilities.
- Optimized padding operations by introducing a pre-allocated 512-space buffer in the `padding()` function to reduce repeated string allocations.
- Updated all imports across 30 files to use the new `padding()` function from pi-tui instead of inline `' '.repeat()` calls.
- Renamed local variables from `padding` to `pad`, `padSize`, `indent`, or `linePad` to avoid naming conflicts with the imported `padding()` function.
- Exported the `padding()` utility function from the tui package index for public use.
- Updated @bufbuild/protoc-gen-es from ^2.10.2 to ^2.11.0.
- Updated @types/bun from ^1.2.18 to ^1.3.7.
- Updated prettier from ^3.8.0 to ^3.8.1.
- Updated @sinclair/typebox from ^0.34.46 to ^0.34.48 across multiple packages.
- Updated @bufbuild/protobuf from ^2.10.2 to ^2.11.0 in packages/ai.
- Updated multiple dependencies including chalk, diff, file-type, zod, winston, recharts, and mime-types to their latest patch and minor versions.
- Replaced external JSON/JSONL parsing libraries with Bun's built-in JSON5 and JSONL APIs across all packages.
- Removed dependencies: json5, ndjson, get-east-asian-width, json-stringify-safe, split2, through2, and @types/ndjson.
- Replaced custom text width and ANSI wrapping implementations with Bun.stringWidth() and Bun.wrapAnsi() APIs.
- Updated Bun type definitions from ^1.2.17 to ^1.2.18 and bun-types from ^1.3.5 to ^1.3.7.
- Refactored JSONL parsing throughout codebase to use Bun.JSONL.parse() and Bun.JSONL.parseChunk() with improved buffer management and error handling.
- Updated test assertions to use Bun.stringWidth() for dynamic width calculations instead of hardcoded values.
- Removed Prettier configuration files (.prettierignore and .prettierrc) and migrated formatting to Biome.
- Updated Biome configuration from version 2.3.11 to 2.3.12 and changed arrowParentheses rule from 'always' to 'asNeeded'.
- Pinned @biomejs/biome dependency to exact version 2.3.12 in package.json and bun.lock.
- Applied consistent arrow function formatting across 489 files by removing unnecessary parentheses around single parameters.
- Removed blank lines after comment blocks and reorganized imports for consistency across the codebase.
- Converted named imports from node modules (fs, path, os) to namespace imports across all packages.
- Extended extension loader error handling with isEacces and hasFsCode type guards.
- Converted readdirSync, readFileSync, and statSync to async readdir, readFile, stat across skills and agent discovery.
- Made scanDirectoryForSkills async and refactored custom directory scanning to use Promise.all for concurrent processing.
- Updated agent discovery to use fs/promises for async file reading and refactored helper patterns.
- Added AgentParsingError exception class for better error handling during agent parsing.
- Added filesystem error type guards (isEnoent, isEacces, isPerm, etc.) to pi-utils for safe error checking.
- Added color manipulation utilities to pi-utils for accessibility features.
- Added color-blind mode setting to settings manager.
- Migrated plugins, settings, and config modules from sync to async file operations.
- Updated error handling to use new pi-utils type guards for type-safe checking.
- Removed WASM generation script; use Bun `wasm?raw` loader for imports.
- Added bunfig.toml with loaders for `.md`, `.py`, and `.wasm?raw` text imports.
- Added types/assets/index.d.ts for global TypeScript module declarations.
- Unified TypeScript configuration with tsgo-based checking across monorepo.
- Removed build and WASM steps from install and publish pipelines.
- Added multi-cell Python execution with sequential processing in persistent kernel.
- Changed Python tool API to use cells array instead of single code parameter.
- Renamed workdir parameter to cwd across Bash and Python tools for consistency.
- Fixed indentation adjustment logic for mixed indentation levels in patch tool.
- Fixed OAuth callback flows for GitHub Copilot and Google Gemini to handle cancellation and manual input errors.
- Hardened database file permissions to 0o600 and directory creation to 0o700 to prevent credential leakage.
- Fixed cache invalidation for streaming edits and file existence checks for prompt templates.
- Fixed bash output streaming to prevent premature closure and LSP client request handling for aborted signals.
- Created new @oh-my-pi/pi-utils workspace package with shared utilities for logging, process management, stream handling, and temporary directory management.
- Migrated all packages to use centralized logger from @oh-my-pi/pi-utils instead of local winston implementations.
- Replaced custom process spawning and stream reading implementations with standardized cspawn and readLines utilities across all modules.
- Converted synchronous file operations and process spawning to async patterns using Bun shell syntax and fs/promises.
- Added streaming edit abort functionality with configurable setting to abort on patch preview failures.
- Updated test framework from vitest to bun:test across all test suites.
- Reformatted code indentation and spacing across multiple TypeScript files.
- Standardized export statements to single-line format in various modules.
- Fixed async/await precedence and corrected indentation in benchmark files.
- Added benchmark report files for GPT-5.1-codex-mini model performance.
- Added normative patch generation for canonicalizing edit tool output.
- Implemented tool call argument rewriting for session history persistence.
- Enhanced patch applicator to support normalized patch input processing.
- Updated agent and AI types to include normative input parameters.
- Removed tar-stream dependency and replaced with Bun.Archive for benchmark package.
- Enhanced patch applicator with fallback variant generation and improved fuzzy matching.
- Added support for ellipsis placeholders, top-of-file anchors, and comment-prefix normalization.
- Renamed operation and moveTo parameters to op and rename across patch tool interfaces.
- Added batch processing system to benchmark runner for parallel task execution.
- Added comprehensive benchmark reports for Claude Haiku and GPT-5.1-codex-mini models.
- Added new @oh-my-pi/react-edit-benchmark package for React source code mutation testing.
- Added --no-title flag to disable automatic session title generation in coding-agent.
- Changed default behavior of read tool to omit line numbers by default.
- Added environment variable support for edit tool configuration and fuzzy matching.
- Updated test files to use import.meta.dir for ES module compatibility.
- Added Prettier configuration and CI improvements for private package handling.