- Added retry mechanism for benchmark tasks with separate system and retry prompt templates to improve edit success rates.
- Introduced autocorrect tracking metrics including autocorrect-free success rate and edit autocorrect counts in task and benchmark summaries.
- Refactored prompt building into modular functions (buildBenchmarkSystemPrompt, buildInitialBenchmarkPrompt, buildRetryBenchmarkPrompt) with BenchmarkPromptDelivery type for distinguishing initial and follow-up messages.
- Added session management with cache-keyed provider session IDs using xxHash64 and centralized RPC argument building via prepareBenchmarkSessionSetup.
- Deleted tsconfig.build.json and tsconfig.check.json files that only extended the base tsconfig.json without adding configuration.
- Updated check:ts script to remove the unused packages/stats/tsconfig.client.json reference.
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Removed static thinking mode constant exports (THINKING_LEVELS, ALL_THINKING_LEVELS, ALL_THINKING_MODES, THINKING_MODE_DESCRIPTIONS, THINKING_MODE_LABELS) in favor of dynamic function-based API.
- Renamed formatThinking() to getThinkingMetadata() with return type changed from string to structured ThinkingMetadata object containing value, label, and description.
- Renamed getAvailableThinkingLevel() to getAvailableThinkingLevels() and getAvailableThinkingEffort() to getAvailableThinkingEfforts() with added default parameters for runtime flexibility.
- Updated all consumer modules to use new function-based API instead of static constants, enabling dynamic thinking mode configuration.
- Updated option interfaces, param builders, and stream functions to use unified `reasoning` field.
- Added `resolveOpenAiReasoningEffort()` to centralize xhigh clamping logic.
- Replaced type casts with `castApi()` helper and fixed test option names.
- Updated coding-agent and benchmark references for consistency.
- Extracted thinking module with ThinkingEffort, ThinkingLevel, and ThinkingMode types to centralize reasoning configuration across packages.
- Migrated ThinkingLevel type from pi-agent-core to pi-ai package with new validation functions parseThinkingLevel() and getAvailableThinkingLevel().
- Consolidated thinking level constants and descriptions into reusable exports (ALL_THINKING_LEVELS, THINKING_MODE_DESCRIPTIONS) for consistent UI display.
- Removed local thinking-effort-label utility and replaced formatThinkingEffortLabel() with centralized formatThinking() function from pi-ai.
- Refactored thinking mode handling to distinguish ThinkingSelector (user-facing with 'off' option) from ThinkingEffort (provider-level).
- Added sync-exports.ts script to enforce canonical package.json field ordering and auto-generate exports from filesystem structure.
- Expanded subpath exports across 11 packages (ai, agent, coding-agent, natives, stats, utils, tui, swarm-extension) to enable flexible module imports.
- Reorganized package.json field ordering to follow standard conventions with metadata before main/exports and files/exports at end.
- Consolidated wildcard export patterns to simplify import paths while maintaining backward compatibility with explicit subpath exports.
- Changed hashline format separator from pipe (|) to colon (:) for improved readability across all tools and output formats.
- Refactored hashline edit API with operation-based structure: renamed delete->rm, rename->mv, set->target/new_content, and added explicit op field for operation types.
- Updated hashline hash encoding from 4-character base36 to 2-character hexadecimal for more compact representation.
- Replaced anchor terminology with tags throughout hashline documentation and API for clearer semantics.
- Redesigned hashline edit API with new operation names (set, set_range, insert) and structured body parameter accepting string arrays for multiline edits.
- Changed hashline reference format from LINE:HASH to LINE#ID throughout tools and documentation for improved clarity.
- Enhanced insert operation to support optional before/after anchors enabling flexible insertion positioning and boundary echo stripping.
- Made hashline autocorrect heuristics conditional on PI_HL_AUTOCORRECT environment variable for controlled behavior.
- Added benchmark reports for claude-haiku-4-5 and GPT-5.2-Codex models demonstrating hashline edit variant performance.
- Added provider failure detection and exponential backoff retry logic to handle authentication and authorization errors in benchmark tasks.
- Implemented HashlineMismatchError behavior in coding-agent to fail on stale hash references instead of silently relocating edits.
- Simplified hashline validation by removing automatic line relocation logic and hash tracking infrastructure.
- Added benchmark report for claude-sonnet-4-6 model showing 85% task success rate with detailed failure analysis and performance metrics.
- Added `--no-rules` CLI flag to coding-agent to disable rules discovery and loading.
- Added `rules` option to CreateAgentSessionOptions to allow custom rules configuration.
- Added `sessionDir` option to RpcClientOptions and implemented Symbol.dispose() for resource cleanup.
- Removed tarball-based task loading; migrated to directory-based fixtures with required inputDir and expectedDir properties.
- Refactored runner to use RpcClient resource management with `using` statement and simplified fixture handling.
- Consolidated type definitions and removed tarball.ts module in favor of streamlined task interface.
- Replaced regex-based code mutation with Babel AST.
- Introduced a wider range of precise, structural bug types.
- Integrated code formatting into generated benchmark fixtures.
- Updated dependencies with Babel and `regexp-tree`.
- Clarified hashline anchor format to emphasize `LINE:HASH` only without content suffix.
- Added explicit guidance on `set_line` with empty text behavior versus `replace_lines` for deletion.
- Added instruction to prefer `insert_after` over line replacement when adding fields/arguments/imports.
- Added common failure patterns section documenting anchor copying errors and wide range pitfalls.
- Updated all example tool calls to use concrete `LINE:HASH` format instead of template placeholders.
- Added issue templates (bug report, feature request, question) with provider/platform dropdowns
- Added PR template and security policy
- Added /triage Claude command for automated issue classification
- Updated LICENSE to include Can Boluk copyright
- Updated author/contributors across all package.json files
- Added homepage, bugs, keywords, engines to all packages
- Created platform:wsl label for WSL-specific issues
- Added three new benchmark run reports for GPT-5.3 Codex model with different edit variants (hashline, patch, replace).
- Recorded performance metrics across 60 tasks per variant with success rates ranging from 80-83.3%.
- Fixed incorrect path resolution for CLI module by using Bun.fileURLToPath() to convert import.meta.resolve() result to a proper file path.
- Reordered constant declarations to ensure TMP_DIR is initialized before CLI_PATH which depends on it.
- Updated scripts/repro-stuck.ts to use consistent CLI path resolution with Bun.fileURLToPath().
- Exported renderPromptTemplate function from config/prompt-templates module for programmatic prompt template rendering.
- Exported computeLineHash function from patch/hashline module for external patch utility access.
- Added ./cli export path in package.json for direct CLI module access.
- Changed hashline display format separator from pipe to two spaces for improved readability.
- Removed `lines` and `hashes` parameters from read tool in favor of automatic file display mode resolution.
- Added `resolveFileDisplayMode` utility to centralize file display mode configuration logic.
- Integrated file display mode settings into grep and read tools for consistent output formatting.
- Consolidated tool parameter types to use schema-derived types via Typebox `Static` utility.
- Updated parseLineRef to handle both legacy pipe-separator and new two-space hashline formats.
- Removed outdated benchmark reports from previous test runs across multiple models and edit variants.
- Added new benchmark reports for Claude Haiku 4.5, Claude Sonnet 4.5, Deepseek V3.2, Devstral Medium, Gemini 3 Flash, GLM-4.5-Air, GPT-5.1-Codex-Mini, GPT-5.2-Codex, Grok-4-1-Fast, Grok-4-Fast-Non-Reasoning, Grok-Code-Fast-1, Kimi K2.5, Minimax M2.1, Qwen Turbo, and Zai-GLM-4.7 models across hashline, patch, and replace edit variants.
- Updated benchmark test results with improved task success rates and edit success metrics across all tested models.
- Fixed resource leak in browser query handler by properly disposing owned proxy elements for non-winning candidates.
- Fixed script evaluation to support async functions and await expressions in browser evaluate operations.
- Removed unused maxLines parameter from buildMutationPreviewAgainstOriginal function and simplified preview generation logic.
- [ { "text": "Updated runner tool arguments from 'read,edit,write,ls' to 'read,edit,write' with '--no-skills' flag in single task execution.", "user_visible": false }, { "text": "Updated runner tool arguments from 'read,edit,write,ls' to 'read,edit,write' with '--no-skills' flag in batch execution.", "user_visible": false } ], "issue_refs">.
- Added new `replace` hashline edit operation for substring-style fuzzy text replacement without line references, with optional `all` flag for replace-all behavior.
- Added `noopEdits` array to `applyHashlineEdits` return value to report edits that produced no changes, including edit index, location, and current content for diagnostics.
- Added validation to detect and reject hashline edits using wrong-format fields (`old_text`/`new_text` from replace mode, `diff` from patch mode) with helpful error messages.
- Renamed hashline edit operation keys from `single`/`range`/`insertAfter` to `set_line`/`replace_lines`/`insert_after` for clearer semantics.
- Enhanced error messages for wrong-format hashline edits to guide users toward correct operation syntax.
- Added noopEdits array to applyHashlineEdits return type to track edits that produce no changes.
- Added validation to reject edits with wrong-format fields (old_text/new_text from replace mode, diff from patch mode) that indicate model confusion.
- Added additionalProperties tolerance to hashline edit schemas to allow flexible field handling.
- Added deduplication logic to remove duplicate edits targeting the same line(s) with identical destination content.
- Improved error handling for missing end fields in edit ranges by returning single-line specs instead of requiring both start and end.
- Enhanced no-op error recovery guidance in prompts with detailed instructions to re-read file and function context after consecutive no-op errors.
- Added `noOpRetryLimit` and `mutationScopeWindow` configuration parameters to benchmark config.
- Implemented zero-tool-call retry logic in single task runner to automatically retry when model produces no tool calls, up to the configured limit.
- Implemented zero-tool-call retry logic in batched task runner to automatically retry when model produces no tool calls, up to the configured limit.
- Removed unused `_formatEditArgs` function from report module.
- Removed preventable failure diagnostics reporting from benchmark report generation.
- Renamed hashline edit operation types from replaceLine/replaceLines to single/range for improved clarity.
- Renamed content field to replacement in hashline edit operations to better reflect its purpose.
- Enhanced no-op edit diagnostics to perform line-by-line comparisons and distinguish between identical and normalized replacements.
- Improved error messages for hashline edits to clarify differences between literal identical content and whitespace-normalized content.
- Removed timeout-retries configuration option from react-edit-benchmark CLI and simplified retry logic.
- Updated all test cases and documentation to reflect renamed hashline edit operation types and fields.
- Removed insertBefore and substr hashline edit operations, simplifying the edit API to support only replaceLine, replaceLines, and insertAfter operations.
- Updated hashline edit schema and type definitions to remove insertBefore and substr operation types from the ParsedRefs union and edit validation logic.
- Removed insertBefore and substr test cases from hashline test suite, including tests for insert-before functionality, anchor echo stripping, and substring matching.
- Updated benchmark runner to refactor edit operations from src/dst format to discriminated union types (replaceLine, replaceLines, insertAfter) and adjusted insertAfter line references.
- Added remaps property to HashlineMismatchError to store a map of stale references to their corrected versions with quick fix suggestions.
- Added diagnostic output for no-op edits that displays target lines from failed edit operations to help users understand why edits had no effect.
- Added computeLineHash and parseLineRef helper functions to public exports for use in diagnostic and validation logic.
- Enhanced error messages in HashlineMismatchError with quick fix suggestions showing stale references and their corrected versions.
- Updated hashline edit tool documentation with critical guidelines on preserving formatting, recovery procedures, and pre-submission verification checklist.
- Redesigned hashline edit API to use discriminated union variants (replaceLine, replaceLines, insertAfter, insertBefore, substr) instead of nested src/dst structure.
- Changed hash algorithm from xxHash64 hex to xxHash32 base36 encoding and increased hash length from 2 to 3 characters.
- Added substr edit variant to match and replace content by unique substring when line hashes are unavailable.
- Implemented line relocation heuristics to resolve stale line references when hash uniquely identifies a moved line.
- Moved substr needle search from edit application phase to pre-validation phase for earlier error detection.
- Migrated hashline edit type definitions from manual TypeScript interfaces to schema-derived types using Static<typeof schema> pattern.
- Added mutation preview hints to error messages when edits fail with 'No changes made' errors, showing line numbers, hashes, and added/removed lines.
- Changed formatter to pin JavaScript fixtures to the flow parser to avoid parser-dependent formatting drift.
- Added whitespace preservation logic to treat whitespace-only differences as passing in file verification.
- Removed substring source specification kind from hashline edits, requiring users to use line-hash references instead.
- Simplified SrcSpec type to use generic type parameter for line references, removing substring variant from union.
- Updated react-edit-benchmark to use --max-tasks option for deterministic task sampling and changed default batch sizes to 1.
- Migrated models export from TypeScript module to JSON format, changing the public API from importing MODELS from './models.generated' to importing from './models.json' with JSON import assertion.
- Updated @anthropic-ai/sdk dependency from ^0.72.1 to ^0.74.0.
- Simplified model generation script by replacing 49 lines of TypeScript code generation with direct JSON serialization.
- Updated @types/bun devDependency from ^1.3.8 to ^1.3.9 across all packages.
- Removed models.generated.ts exclusion from biome.json linting configuration.
- Added guided mode support with --guided and --no-guided command-line options to enable contextual editing hints.
- Added --max-attempts option (default: 2, range: 1-5) to enable retry logic for failed benchmark tasks.
- Added --require-read-tool-call option to enforce read tool invocation before edit operations.
- Implemented retry mechanism in task runner that re-attempts failed tasks up to maxAttempts with guided context on subsequent attempts.
- Extended TaskMetadata to capture file location details including filePath, fileName, lineNumber, originalSnippet, and mutatedSnippet.
- Changed success metric from patchApplied to editSucceeded to better reflect successful edit tool invocations.
- Fixed line hash computation to normalize whitespace before hashing, ensuring consistent hashes for semantically identical lines.
- Removed index parameter from hash computation to prevent hash variance based on line position.
- Changed `HashlineEdit.src` from string format (e.g., `"5:ab"`, `"5:ab..9:ef"`) to structured `SrcSpec` object with discriminated union types (`{ kind: "single", ref: "..." }`, `{ kind: "range", start: "...", end: "..." }`, `{ kind: "insertAfter", after: "..." }`, `{ kind: "insertBefore", before: "..." }`, `{ kind: "substring", needle: "..." }`).
- Refactored `parseSrc` function to `parseSrcSpec` to handle structured `SrcSpec` objects instead of string parsing, replacing string validation and delimiter parsing logic with explicit case handling for each operation kind.
- Updated tool schema in `index.ts` to define `srcSpecSchema` as a discriminated union type with five variants, replacing the previous simple string-based `src` parameter definition.
- Added substring-based source matching for hashline edits to support matching lines by content when hash references are unavailable.
- Added automatic detection and repair of single-line merges where models incorrectly combine multiple lines into one.
- Added normalization of Unicode-confusable hyphens to ASCII hyphens to handle model-generated variations.
- Added heuristics to restore indentation and preserve wrapped line formatting during edits.
- Relaxed comma validation in src to allow commas while rejecting inputs with multiple line references.
- Enhanced parseLineRef to accept shorter hash prefixes instead of requiring exact hash matches.
- Improved error messages for hash mismatches to provide more actionable guidance.
- Replaced HashlineEdit API from old/new/after fields to src/dst fields with unified range syntax supporting single lines, ranges, insert-after, and insert-before operations.
- Added parseSrc() function to parse structured line references from src strings supporting formats like '5:ab', '5:ab..9:ef', '5:ab..', and '..5:ab'.
- Added heuristics to strip anchor line echoes and range boundary echoes from model-generated replacement content.
- Implemented preserveWhitespaceOnlyLinesLoose() for improved whitespace preservation using loose matching strategy when replacement line counts don't match.
- Added comprehensive benchmark reports for Claude Sonnet 4.5, Gemini 2.5 Flash Lite, and GPT-5.1 Codex Mini models evaluating hashline edit performance across 60 tasks.
- Added MCPRequestOptions interface with signal property for request cancellation via AbortSignal.
- Added abort signal support to MCP tool execution enabling request cancellation via Escape-to-interrupt or other abort mechanisms.
- Enhanced MCP request handling with abort signal propagation through HTTP, SSE, and stdio transports with proper cleanup.
- Improved stdio transport request handling to use Promise.withResolvers for cleaner async flow and better abort signal integration.
- Updated HTTP transport to combine operation abort signals with timeout signals using AbortSignal.any() for unified cancellation.
- Modified SSE response parsing to support abort signals and distinguish between timeout and user-initiated cancellation.
- Added automatic stripping of `LINE:HASH|` display prefixes and unified-diff `+` markers from replacement content in hashline edits.
- Enhanced replace edits to preserve original whitespace on lines where only whitespace differs, preventing spurious formatting diffs when models reformat code.
- Added helper functions `stripNewLinePrefixes()`, `preserveWhitespaceOnlyLines()`, and `equalsIgnoringWhitespace()` to improve edit robustness.
- Optimized line hash computation by precomputing hex lookup table and using bitwise operations instead of string formatting.
- Simplified conditional logic in read tool for hash and line number configuration.
- Refactored react-edit-benchmark model/provider parsing to use concise ternary operator.
- Added hashline edit mode for line-addressed edits using hash-verified line references with xxHash64 integrity verification.
- Replaced `edit.patchMode` boolean setting with `edit.mode` enum supporting 'replace', 'patch', and 'hashline' modes.
- Added `readHashLines` configuration setting to include line hashes in read output for hashline edit mode.
- Implemented `computeLineHash`, `formatHashLines`, `parseLineRef`, `validateLineRef`, and `applyHashlineEdits` utility functions for hashline operations.
- Updated read tool to support optional `hashes` parameter and prioritize hash lines over line numbers when both are requested.
- Changed `getEditModelVariants()` return type from `Record<string, 'patch' | 'replace'>` to `Record<string, EditMode | null>` and removed hardcoded model-specific defaults.