- Removed outdated benchmark reports from previous test runs across multiple models and edit variants.
- Added new benchmark reports for Claude Haiku 4.5, Claude Sonnet 4.5, Deepseek V3.2, Devstral Medium, Gemini 3 Flash, GLM-4.5-Air, GPT-5.1-Codex-Mini, GPT-5.2-Codex, Grok-4-1-Fast, Grok-4-Fast-Non-Reasoning, Grok-Code-Fast-1, Kimi K2.5, Minimax M2.1, Qwen Turbo, and Zai-GLM-4.7 models across hashline, patch, and replace edit variants.
- Updated benchmark test results with improved task success rates and edit success metrics across all tested models.