12 Commits

Author SHA1 Message Date
can1357 9856f904d7 feat(typescript-edit-benchmark): introduced empirical edit mutation planning
- Add new structural, multi-edit, and block-level mutation classes with updated category mappings.
- Introduce hunk extraction, placement, rendering, and solver utilities along with unit tests.
- Implement size-based mutation planning, prompt validation logic, and new prompt markdown templates.
- Update benchmark generation scripts and package configurations to support empirical edit shape statistics.
2026-07-30 07:45:01 +02:00
can1357 0856055dfe feat(harbor-manager): unified benchmark normalization and reporting
- Implemented a unified benchmark normalization layer to handle metrics and artifacts across harbor, edit, and snapcompact benchmarks.
- Integrated the TypeScript edit benchmark directly into the manager, migrating logic and deprecating the standalone package.
- Updated the API and database schema to support standardized benchmark configurations, metrics, and trace-based reporting.
- Enhanced the UI to visualize comparative performance metrics, including pass rate, cost, and latency deltas for benchmark runs.
2026-07-13 04:21:54 +02:00
can1357 18f3386aed feat(typescript-edit-benchmark): tracked reasoning tokens
- Added reasoning token counts to benchmark results and reporting.
- Updated session statistics and task summaries to include reasoning metrics.
- Updated report generation to display reasoning token breakdown.
2026-06-28 07:27:02 +02:00
can1357 5bce7ed6df feat: added advisory transcript formatting and one-shot benchmark metrics
- Introduced advisory note output as `<advisory>` tags with optional severity and guidance.
- Updated session transcript formatting to `### Session update` and inline watched role labels.
- Added shared `escapeXmlText` utility and escaped XML-sensitive text in advisor outputs.
- Added one-shot success run token metrics and one-shot statistics reporting.
2026-06-16 18:34:50 +02:00
can1357 9d457f73d9 test: migrated test imports to package subpath exports
- Replaced relative `../src` imports with `@oh-my-pi/pi-ai` and `@oh-my-pi/pi-agent-core` subpaths.
2026-06-08 19:03:55 +02:00
can1357 a1ba50b4da feat(benchmark): added median, p1, and p99 token distribution stats
- Added `percentile` and `summarizeTokenDistribution` helpers to runner.
- Extended `BenchmarkSummary` with `medianTokensPerTask`, `p1TokensPerTask`, and `p99TokensPerTask`.
- Updated live progress output and markdown report table to show distribution columns.
- Added unit tests covering percentile interpolation and summary fields.
2026-05-31 01:14:52 +02:00
can1357 2ef4eb7931 refactor(typescript-edit-benchmark): restructured line pairing to use diff-based matching
- Replaced line-by-line comparison with diff-based pairing to correctly handle insertions and deletions that shift line indices.
- Added logic to preserve trailing newline semantics from the actual content.
- Added test case verifying that indent-only changes are normalized even when earlier insertions shift line positions.
2026-05-27 13:34:19 +02:00
can1357 30793c1655 refactor: restructured hashline to use file-level hash validation with colon separators
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.
2026-05-26 13:25:22 +02:00
can1357 975941aba4 chore: remove garbage tests 2026-05-12 04:09:33 +02:00
can1357 47dab9ee97 feat: added atom-mode Lid range edits and hashline shifted-hash recovery
- Added atom edit support for Lid ranges, before-anchor inserts, and no-op Lid=TEXT success handling.
- Expanded hashline recovery to scan shifted hashes, match unique alternates, and emit ±5 anchor-shift hints.
- Added edit failure categorization with per-category counts, percentages, and detailed report lines.
- Added regression coverage for shifted-hash recovery, range continuation, cursor shorthand, and split-file atom ops.
2026-04-29 23:51:22 +02:00
can1357 c6a11079f5 feat: added compact atom-mode parser and execution support
- Added compact Lark grammar processing and applied it to OpenAI custom-format tools before conversion.
- Reworked atom mode into `---PATH` compact commands with new grammar, parser, and rm/mv file operations.
- Updated `hline`/`href`/`hrefr` helper behavior and hashline mismatch guidance using shared anchor state.
- Standardized path formatting with `formatPathRelativeToCwd` across LSP, prompts, and edit/search/write tools.
- Added benchmark run-path handling, including `.gitignore` runs mapping, absolute reports, and safer snapshot output.
- Added tests for compact grammar payloads, atom parsing/execution, renderer streaming, and path-list outputs.
2026-04-29 01:16:33 +02:00
can1357 d1561beb99 refactor(edit-benchmark): migrated to typescript with in-process client support
- Migrated react-edit-benchmark package to typescript-edit-benchmark with pi-mono source repository.
- Added InProcessClient implementation to eliminate subprocess spawning overhead in benchmark runs.
- Extended benchmark configuration with chunk edit variant, retry limits, and conversation dump support.
- Refactored runner.ts to support both RPC and in-process client modes with improved error telemetry.
- Cleaned up 129 benchmark report files from react-edit-benchmark/runs directory.
2026-04-06 18:39:04 +02:00