- Added PI_STRICT_EDIT_MODE environment variable to control model-specific edit mode defaults.
- Wrapped model-specific edit mode logic behind PI_STRICT_EDIT_MODE condition for conditional behavior.
- Replaced Bun.env direct access with $env utility for consistent environment variable handling.
- Updated rate-edit-tool and typescript-edit-benchmark to set PI_STRICT_EDIT_MODE in test environments.
- Added bounds clamping to prologue and epilogue byte calculations to prevent out-of-range boundary violations.
- Extended chunk boundaries for indent-based languages when epilogue exceeds calculated range with trailing newline.
- Added comprehensive test coverage for Python chunk editing operations including body/head replacement and indentation preservation.
- Extracted working directory formatting logic into reusable utility function and applied tab sanitization to bash command previews.
- Added Auto QA tool (`report_tool_issue`) for automated tracking of unexpected tool behavior with environment variable and setting support.
- Added Python tool environment warmup on first execution to ensure prelude helpers are available before use.
- Fixed Python prelude introspection to respect execution timeout and signal options, preventing hangs.
- Refactored prelude documentation caching and loading logic into reusable helper functions with test environment awareness.
- Enhanced kernel introspection with optional timeout and signal parameters for better execution control.
- Added system prompt guidance to encourage agents to report tool issues via Auto QA when available.
- Added auto-retry event tracking and message-end event handling to improve agent completion detection.
- Implemented is_effectively_complete() method to detect agent completion based on review sections, todo state, and quiet period.
- Enhanced wait_for_settle() logic to handle auto-retry delays and graceful timeout recovery instead of immediate failure.
- Added token usage tracking from partial message updates to capture intermediate token counts.
- Refactored note_tool_end() to only update activity on error, removing redundant success case.
- Removed last_activity updates from todo reminder and auto-clear handlers to simplify state management.
- Added comprehensive code-editing tool evaluation framework with `rate-edit-tool.py` supporting multi-model benchmarking across TypeScript, Rust, Python, and Markdown.
- Enhanced chunk edit error messages to display fresh chunk context with resolved selectors and anchors for improved debugging.
- Added `--no-lsp` flag to benchmark RPC arguments for TypeScript edit task evaluation.
- Improved chunk body boundary calculation to correctly include closing line indentation in epilogue.
- Added comprehensive benchmark results dataset (`all_models_results.json`) with performance metrics for 6 AI models.
- Enhanced chunk-edit documentation with clarified `@body` selector behavior and append/prepend examples.