- Relaxed `--edit-variant` parsing in `packages/typescript-edit-benchmark/src/index.ts` to accept any string.
- Adjusted token delta calculation in `runner.ts` to subtract estimated system-prompt overhead per assistant turn.
- Captured initial system-prompt tokens in `runSingleTask` and added rough `estimateTokens` helper for corrected accounting.
- Reframed `scripts/rate-edit-tool.py` prompts to target edit-tool behavior and constraints for each fixture type.
- Added a Go fixture and registered it as `main.go` in the reference files, fixture list, and descriptions.
- Changed prompt construction and workspace setup to exercise every fixture in a single session instead of one file at a time.
- Updated run identifiers and output filenames, and set run results to report a fixture scope of `all`.
- Added a rerun flow that synthesizes oracle findings from existing `review_*.md` files without re-running fixtures.
- Refactored oracle input handling to source `(model, fixture, path)` tuples, and persisted oracle prompt and synthesis outputs in the results directory.
- Removed `OPENROUTER_API_KEY` bootstrapping and passed environment, and now write oracle errors to `oracle_error.txt` on failure.
- Added `vim` as an edit variant in benchmark CLI/config and rating script coverage.
- Expanded benchmark execution so `vim` is treated as a mutation tool for retries, stats, and edit intent checks.
- Adjusted `TaskTool` output schema precedence so explicit params override agent frontmatter.
- Fixed `TaskTool` success counting by excluding aborted tasks from success totals.
- Improved validation guidance in `SubmitResultTool`/`TodoWriteTool` for clearer recovery when payloads are missing or invalid.
- Added background command PID regression coverage in `executeBash` to confirm a real, terminateable PID is returned.
- Restructured blank-line cleanup logic in chunk editor to correctly handle artifacts from deletions before and after delimiters.
- Refactored rate-edit-tool.py to support per-fixture model runs instead of single model-wide runs, enabling parallel evaluation across multiple code fixtures.
- Extracted fixture metadata into FIXTURES tuple and introduced build_fixture_prompt() helper to customize prompts per fixture.
- Updated ModelRunRecorder and ProgressPrinter to track run_id (model + fixture) separately from model name, enabling independent progress tracking.
- Modified materialize_workspace() to optionally create single-fixture workspaces and updated run_model_sync() to generate fixture-scoped workspace paths and result files.
- Enhanced chunk-edit tool documentation with clarifications on @decl region behavior and guidance for attribute/decorator modifications.
- Improved chunk edit implementation to handle @body region scanning and markdown blank-line preservation with additional test coverage.
- Added oracle model review synthesis to rate-edit-tool.py for aggregating findings across multiple model reviews.
- Updated test infrastructure to run TypeScript and Rust tests in parallel using cargo nextest with improved logging control.
- Updated TypeScript dependencies including @typescript/native-preview and typescript to latest versions.
- Added PI_STRICT_EDIT_MODE environment variable to control model-specific edit mode defaults.
- Wrapped model-specific edit mode logic behind PI_STRICT_EDIT_MODE condition for conditional behavior.
- Replaced Bun.env direct access with $env utility for consistent environment variable handling.
- Updated rate-edit-tool and typescript-edit-benchmark to set PI_STRICT_EDIT_MODE in test environments.
- Added bounds clamping to prologue and epilogue byte calculations to prevent out-of-range boundary violations.
- Extended chunk boundaries for indent-based languages when epilogue exceeds calculated range with trailing newline.
- Added comprehensive test coverage for Python chunk editing operations including body/head replacement and indentation preservation.
- Extracted working directory formatting logic into reusable utility function and applied tab sanitization to bash command previews.
- Added Auto QA tool (`report_tool_issue`) for automated tracking of unexpected tool behavior with environment variable and setting support.
- Added Python tool environment warmup on first execution to ensure prelude helpers are available before use.
- Fixed Python prelude introspection to respect execution timeout and signal options, preventing hangs.
- Refactored prelude documentation caching and loading logic into reusable helper functions with test environment awareness.
- Enhanced kernel introspection with optional timeout and signal parameters for better execution control.
- Added system prompt guidance to encourage agents to report tool issues via Auto QA when available.
- Added auto-retry event tracking and message-end event handling to improve agent completion detection.
- Implemented is_effectively_complete() method to detect agent completion based on review sections, todo state, and quiet period.
- Enhanced wait_for_settle() logic to handle auto-retry delays and graceful timeout recovery instead of immediate failure.
- Added token usage tracking from partial message updates to capture intermediate token counts.
- Refactored note_tool_end() to only update activity on error, removing redundant success case.
- Removed last_activity updates from todo reminder and auto-clear handlers to simplify state management.
- Added comprehensive code-editing tool evaluation framework with `rate-edit-tool.py` supporting multi-model benchmarking across TypeScript, Rust, Python, and Markdown.
- Enhanced chunk edit error messages to display fresh chunk context with resolved selectors and anchors for improved debugging.
- Added `--no-lsp` flag to benchmark RPC arguments for TypeScript edit task evaluation.
- Improved chunk body boundary calculation to correctly include closing line indentation in epilogue.
- Added comprehensive benchmark results dataset (`all_models_results.json`) with performance metrics for 6 AI models.
- Enhanced chunk-edit documentation with clarified `@body` selector behavior and append/prepend examples.