17 Commits

Author SHA1 Message Date
can1357 55f5ebec49 chore: reformat 2026-07-15 00:08:50 +02:00
can1357 084791dade chore(scripts): removed todo workflow from scripts/rate-edit-tool
- Removed todo workflow event handling from `scripts/rate-edit-tool.py` by deleting `Todo*` event plumbing and ops.
- Pruned Go/Markdown fixture coverage from prompts and fixture maps, keeping TypeScript/Rust/Python only.
- Dropped todo-based completion gating in `is_effectively_complete` and removed todo columns from progress output.
2026-05-02 06:59:01 +02:00
can1357 2826ec2ea0 refactor(scripts): restructured edit-mode fallbacks to ignore strict mode
- Removed PI_STRICT_EDIT_MODE gating from edit-mode resolution so model fallbacks now always apply.
- Stopped injecting PI_STRICT_EDIT_MODE in edit-benchmark.py and rate-edit-tool.py execution environments.
- Removed PI_STRICT_EDIT_MODE from environment-variable documentation and strict-mode test coverage.
2026-05-02 04:34:39 +02:00
can1357 a188651f63 fix(coding-agent/session): skipped empty thinking entries when formatting session dumps
- Updated `formatSessionDumpText` to skip `thinking` entries with empty or whitespace-only content.
- Prevented empty `<thinking>` sections from being emitted in session dump output.
2026-04-29 20:43:29 +02:00
can1357 5ea1d55e56 feat: removed chunk-mode modules and read/edit entrypoints from pi-natives
- Removed `pi-natives` chunk language classifier modules and all core chunk subsystems (kind, state, render, edit, resolve).
- Removed chunk-mode CLI/read/edit entrypoints, including `read` command and chunk mode registration/prompt tooling.
- Removed chunk selectors from `read` and `grep` tools, switching behavior to raw/L-range handling.
- Fixed poll wait parsing to keep defaulting to `30s` when the provided value is empty.
2026-04-26 08:19:02 +02:00
can1357 60b2fbf1f9 chore(benchmark): cleaned benchmark edit parsing and token accounting
- Relaxed `--edit-variant` parsing in `packages/typescript-edit-benchmark/src/index.ts` to accept any string.
- Adjusted token delta calculation in `runner.ts` to subtract estimated system-prompt overhead per assistant turn.
- Captured initial system-prompt tokens in `runSingleTask` and added rough `estimateTokens` helper for corrected accounting.
- Reframed `scripts/rate-edit-tool.py` prompts to target edit-tool behavior and constraints for each fixture type.
2026-04-26 00:29:17 +02:00
can1357 d82377cda8 revert: "read-to-open"
This reverts commit c48d2e6080.
2026-04-24 22:36:44 +02:00
can1357 66a89b343c feat(scripts): expanded rate-edit suite with Go and all-fixture sessions
- Added a Go fixture and registered it as `main.go` in the reference files, fixture list, and descriptions.
- Changed prompt construction and workspace setup to exercise every fixture in a single session instead of one file at a time.
- Updated run identifiers and output filenames, and set run results to report a fixture scope of `all`.
2026-04-24 19:32:39 +02:00
can1357 ed241ab698 feat(scripts): added oracle rerun mode and persisted synthesis artifacts
- Added a rerun flow that synthesizes oracle findings from existing `review_*.md` files without re-running fixtures.
- Refactored oracle input handling to source `(model, fixture, path)` tuples, and persisted oracle prompt and synthesis outputs in the results directory.
- Removed `OPENROUTER_API_KEY` bootstrapping and passed environment, and now write oracle errors to `oracle_error.txt` on failure.
2026-04-24 19:28:31 +02:00
can1357 c62ab2d53e chore(benchmarks-misc-fixes): cleaned benchmark task/pid validation
- Added `vim` as an edit variant in benchmark CLI/config and rating script coverage.
- Expanded benchmark execution so `vim` is treated as a mutation tool for retries, stats, and edit intent checks.
- Adjusted `TaskTool` output schema precedence so explicit params override agent frontmatter.
- Fixed `TaskTool` success counting by excluding aborted tasks from success totals.
- Improved validation guidance in `SubmitResultTool`/`TodoWriteTool` for clearer recovery when payloads are missing or invalid.
- Added background command PID regression coverage in `executeBash` to confirm a real, terminateable PID is returned.
2026-04-13 12:27:10 +02:00
can1357 91c9e1fd0b refactor: restructured blank-line cleanup and fixture-scoped model evaluation
- Restructured blank-line cleanup logic in chunk editor to correctly handle artifacts from deletions before and after delimiters.
- Refactored rate-edit-tool.py to support per-fixture model runs instead of single model-wide runs, enabling parallel evaluation across multiple code fixtures.
- Extracted fixture metadata into FIXTURES tuple and introduced build_fixture_prompt() helper to customize prompts per fixture.
- Updated ModelRunRecorder and ProgressPrinter to track run_id (model + fixture) separately from model name, enabling independent progress tracking.
- Modified materialize_workspace() to optionally create single-fixture workspaces and updated run_model_sync() to generate fixture-scoped workspace paths and result files.
2026-04-10 19:02:25 +02:00
can1357 bb550efc6b chore: cleaned up documentation, test infrastructure, and dependencies across tooling
- Enhanced chunk-edit tool documentation with clarifications on @decl region behavior and guidance for attribute/decorator modifications.
- Improved chunk edit implementation to handle @body region scanning and markdown blank-line preservation with additional test coverage.
- Added oracle model review synthesis to rate-edit-tool.py for aggregating findings across multiple model reviews.
- Updated test infrastructure to run TypeScript and Rust tests in parallel using cargo nextest with improved logging control.
- Updated TypeScript dependencies including @typescript/native-preview and typescript to latest versions.
2026-04-10 15:53:41 +02:00
can1357 f49098eb76 config: configured PI_STRICT_EDIT_MODE for model-specific edit behavior
- Added PI_STRICT_EDIT_MODE environment variable to control model-specific edit mode defaults.
- Wrapped model-specific edit mode logic behind PI_STRICT_EDIT_MODE condition for conditional behavior.
- Replaced Bun.env direct access with $env utility for consistent environment variable handling.
- Updated rate-edit-tool and typescript-edit-benchmark to set PI_STRICT_EDIT_MODE in test environments.
2026-04-08 21:30:52 +02:00
can1357 26d7783370 fix(chunk): corrected chunk boundary calculations to prevent out-of-range violations
- Added bounds clamping to prologue and epilogue byte calculations to prevent out-of-range boundary violations.
- Extended chunk boundaries for indent-based languages when epilogue exceeds calculated range with trailing newline.
- Added comprehensive test coverage for Python chunk editing operations including body/head replacement and indentation preservation.
- Extracted working directory formatting logic into reusable utility function and applied tab sanitization to bash command previews.
2026-04-08 11:50:42 +02:00
can1357 4c03bad90d feat(coding-agent): added Auto QA tool and Python environment warmup for tool reliability
- Added Auto QA tool (`report_tool_issue`) for automated tracking of unexpected tool behavior with environment variable and setting support.
- Added Python tool environment warmup on first execution to ensure prelude helpers are available before use.
- Fixed Python prelude introspection to respect execution timeout and signal options, preventing hangs.
- Refactored prelude documentation caching and loading logic into reusable helper functions with test environment awareness.
- Enhanced kernel introspection with optional timeout and signal parameters for better execution control.
- Added system prompt guidance to encourage agents to report tool issues via Auto QA when available.
2026-04-08 11:49:13 +02:00
can1357 0c351738fa feat: added auto-retry tracking and completion detection for agent settlement
- Added auto-retry event tracking and message-end event handling to improve agent completion detection.
- Implemented is_effectively_complete() method to detect agent completion based on review sections, todo state, and quiet period.
- Enhanced wait_for_settle() logic to handle auto-retry delays and graceful timeout recovery instead of immediate failure.
- Added token usage tracking from partial message updates to capture intermediate token counts.
- Refactored note_tool_end() to only update activity on error, removing redundant success case.
- Removed last_activity updates from todo reminder and auto-clear handlers to simplify state management.
2026-04-08 08:59:04 +02:00
can1357 53c78766f3 feat: introduced code-editing tool evaluation framework with multi-model benchmarking
- Added comprehensive code-editing tool evaluation framework with `rate-edit-tool.py` supporting multi-model benchmarking across TypeScript, Rust, Python, and Markdown.
- Enhanced chunk edit error messages to display fresh chunk context with resolved selectors and anchors for improved debugging.
- Added `--no-lsp` flag to benchmark RPC arguments for TypeScript edit task evaluation.
- Improved chunk body boundary calculation to correctly include closing line indentation in epilogue.
- Added comprehensive benchmark results dataset (`all_models_results.json`) with performance metrics for 6 AI models.
- Enhanced chunk-edit documentation with clarified `@body` selector behavior and append/prepend examples.
2026-04-08 07:11:53 +02:00