- Updated parseChunkUsage to subtract prompt_tokens_details.cache_write_tokens from prompt token input so OpenRouter write tokens are not misclassified as billable input.
- Set cacheWrite and total token counts to include cache-write usage, while preserving cache-read behavior from existing cached_tokens handling.
- Added OpenRouter attribution tests verifying cacheWrite and cacheRead totals for write-heavy and cache-warm prompts.
- Allowed bare `LidA..LidB` to recover a missing-range-delete typo and accepted `|` as a legacy range replacement separator while validating ranges and replacement text.
- Enabled indented hashline statements to parse as replacement edits and added a preflight pass that validates all atom sections before any file write occurs.
- Updated hash-mismatch messaging, prompt wording, and tests to reflect hash-only rebase candidates and the new range/section behaviors.
- Added atom edit support for Lid ranges, before-anchor inserts, and no-op Lid=TEXT success handling.
- Expanded hashline recovery to scan shifted hashes, match unique alternates, and emit ±5 anchor-shift hints.
- Added edit failure categorization with per-category counts, percentages, and detailed report lines.
- Added regression coverage for shifted-hash recovery, range continuation, cursor shorthand, and split-file atom ops.
- Added compact Lark grammar processing and applied it to OpenAI custom-format tools before conversion.
- Reworked atom mode into `---PATH` compact commands with new grammar, parser, and rm/mv file operations.
- Updated `hline`/`href`/`hrefr` helper behavior and hashline mismatch guidance using shared anchor state.
- Standardized path formatting with `formatPathRelativeToCwd` across LSP, prompts, and edit/search/write tools.
- Added benchmark run-path handling, including `.gitignore` runs mapping, absolute reports, and safer snapshot output.
- Added tests for compact grammar payloads, atom parsing/execution, renderer streaming, and path-list outputs.
- Updated package-manager config to use the hoisted linker and relaxed the @types/bun version range.
- Regenerated bun.lock with updated dependency pins, including postcss, safe-buffer, and string_decoder, plus added jszip nested lock entries.
- Updated the check:tools script to run biome checks without the removed tsgo typecheck step.
- Removed `pi-natives` chunk language classifier modules and all core chunk subsystems (kind, state, render, edit, resolve).
- Removed chunk-mode CLI/read/edit entrypoints, including `read` command and chunk mode registration/prompt tooling.
- Removed chunk selectors from `read` and `grep` tools, switching behavior to raw/L-range handling.
- Fixed poll wait parsing to keep defaulting to `30s` when the provided value is empty.
- Detected transport failures by checking failed run errors for "Timeout exhausted".
- Counted those failures in benchmark summaries and included them in ghost-like run classification, reducing effective run totals accordingly.
- Updated benchmark reporting to display excluded transport-failure runs when present.
- Added `CodeFrameMarker` and `formatCodeFrameLine()` to centralize code-frame gutter formatting.
- Extended diff rendering to preserve `|` and `│` separators, aligning gutter markers and line numbers.
- Reworked AST, grep, hashline, Vim, and diff renderers to use shared line formatting with computed `lineNumberWidth`.
- Updated atom editing flow and tests, including `resolveAtomEntryPaths` migration and new loc-based/edge-case coverage.
- Adjusted benchmark runner early-stop configuration by passing `buildEarlyStop` through prompt collection.
- Expanded hashline and chunk bigram tables to 647 entries and moved chunk checksums to a 40-item namespace.
- Changed hashline anchors from `LINE#ID`/`:` to concatenated `LINEID`\t forms across parsing and tool outputs.
- Removed line-number padding and routed diff/read/grep/renderer output through raw numbers, tabs, and `toDisplayLine` formatting.
- Renamed atom ops to `pre`/`post`, removed `ins`, and updated schemas, prompts, and tests for new insertion behavior.
- Added a --no-early-stop-on-match CLI option, passed into benchmark configuration as earlyStopOnMatch.
- Added early-stop support that verified expected files after mutation-tool completion and aborted the prompt loop on a match.
- Captured an earlyStopped flag in task results and emitted an early_stop event when match-based termination occurred.
- Added raw read output propagation so read/archive commands bypass anchors, line numbers, and chunk formatting.
- Fixed atom/hashline editing by tightening hashline prefixes and applying grouped anchor edits in stable order.
- Hardened chunk parsing and path handling by using `bigram_end` checks and resolving edit paths via shared `args.path` fallback.
- Updated path-related behavior for chunk, replace, patch, and hashline previews to honor optional edit paths.
- Removed mode-level param validators and replaced them with runtime handling, then removed obsolete validation tests.
- Relaxed `--edit-variant` parsing in `packages/typescript-edit-benchmark/src/index.ts` to accept any string.
- Adjusted token delta calculation in `runner.ts` to subtract estimated system-prompt overhead per assistant turn.
- Captured initial system-prompt tokens in `runSingleTask` and added rough `estimateTokens` helper for corrected accounting.
- Reframed `scripts/rate-edit-tool.py` prompts to target edit-tool behavior and constraints for each fixture type.
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
- Added `toolStrictMode` support with `all_strict`/`none`/`mixed` options to OpenAI compatibility.
- Fixed OpenAI-completion strict-mode flows by capturing failed HTTP responses and retrying once as non-strict.
- Fixed completion error reporting by surfacing captured status, headers, and JSON `type`/`param`/`code` details.
- Improved strict-schema enforcement with WeakMap memoization and circular-schema detection in sanitization.
- Fixed OpenRouter provider lookup by resolving fallback model IDs for suffix and date variants in registry resolution.
- Refactored benchmark tooling and added async RPC error-window tracking for scheduled run execution.
- Added `vim` as an edit variant in benchmark CLI/config and rating script coverage.
- Expanded benchmark execution so `vim` is treated as a mutation tool for retries, stats, and edit intent checks.
- Adjusted `TaskTool` output schema precedence so explicit params override agent frontmatter.
- Fixed `TaskTool` success counting by excluding aborted tasks from success totals.
- Improved validation guidance in `SubmitResultTool`/`TodoWriteTool` for clearer recovery when payloads are missing or invalid.
- Added background command PID regression coverage in `executeBash` to confirm a real, terminateable PID is returned.
- Migrated hash API calls from Bun.hash.xxHash64() to Bun.hash() across TypeScript packages for simplified hash generation.
- Consolidated mermaid cache failure tracking by replacing separate failed Set with null values in cache Map.
- Refactored schema-based child extraction in Rust by introducing promotion_fields tracking and schema_wrapper_child() helper function.
- Updated TypeScript configuration files with reformatted arrays and added compiler options for consistency.
- Added PI_STRICT_EDIT_MODE environment variable to control model-specific edit mode defaults.
- Wrapped model-specific edit mode logic behind PI_STRICT_EDIT_MODE condition for conditional behavior.
- Replaced Bun.env direct access with $env utility for consistent environment variable handling.
- Updated rate-edit-tool and typescript-edit-benchmark to set PI_STRICT_EDIT_MODE in test environments.
- Migrated all package tsconfig files to extend tsconfig.workspace.json for unified TypeScript configuration across monorepo.
- Consolidated build and check scripts across 10+ packages to use biome for linting/formatting with separate type checking via tsgo.
- Renamed build scripts from build:native and build:binary to build for simplified command naming across packages/natives and packages/coding-agent.
- Refactored CI workflow to invoke bun tasks instead of inline shell scripts, reducing workflow complexity by 40+ lines.
- Removed sync-exports.ts and repro-stuck.ts scripts; deleted path aliases from tsconfig.base.json in favor of workspace-based configuration.
- Updated turbo.json with new task definitions (check:types, lint, fmt, fix) and removed build:native/embed:native tasks.
- Added comprehensive code-editing tool evaluation framework with `rate-edit-tool.py` supporting multi-model benchmarking across TypeScript, Rust, Python, and Markdown.
- Enhanced chunk edit error messages to display fresh chunk context with resolved selectors and anchors for improved debugging.
- Added `--no-lsp` flag to benchmark RPC arguments for TypeScript edit task evaluation.
- Improved chunk body boundary calculation to correctly include closing line indentation in epilogue.
- Added comprehensive benchmark results dataset (`all_models_results.json`) with performance metrics for 6 AI models.
- Enhanced chunk-edit documentation with clarified `@body` selector behavior and append/prepend examples.
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).
- Excluded ghost runs (failed runs with zero activity) from benchmark statistics to prevent skewed averages.
- Introduced isGhostRun() helper function to identify and filter runs with no tokens or tool calls.
- Updated report generation and task summarization to use nonGhostRuns for all metric calculations.
- Fixed indentation and formatting across multiple files for consistency.
- Reformatted code blocks to use tabs instead of spaces and improved line breaks for readability.
- Added connection timeout configuration to benchmark runner for early abort on no events.
- Implemented two-phase timeout strategy in benchmark runner with connection and activity phases.
- Migrated react-edit-benchmark package to typescript-edit-benchmark with pi-mono source repository.
- Added InProcessClient implementation to eliminate subprocess spawning overhead in benchmark runs.
- Extended benchmark configuration with chunk edit variant, retry limits, and conversation dump support.
- Refactored runner.ts to support both RPC and in-process client modes with improved error telemetry.
- Cleaned up 129 benchmark report files from react-edit-benchmark/runs directory.