Commit Graph

26 Commits

Author SHA1 Message Date
can1357 fbe051bcbd fix(ai): fixed OpenRouter cache write attribution in usage parsing
- Updated parseChunkUsage to subtract prompt_tokens_details.cache_write_tokens from prompt token input so OpenRouter write tokens are not misclassified as billable input.
- Set cacheWrite and total token counts to include cache-write usage, while preserving cache-read behavior from existing cached_tokens handling.
- Added OpenRouter attribution tests verifying cacheWrite and cacheRead totals for write-heavy and cache-warm prompts.
2026-04-30 02:24:57 +02:00
can1357 863560ffb6 fix(coding-agent/edit): added atom range-repair and multi-section preflight checks
- Allowed bare `LidA..LidB` to recover a missing-range-delete typo and accepted `|` as a legacy range replacement separator while validating ranges and replacement text.
- Enabled indented hashline statements to parse as replacement edits and added a preflight pass that validates all atom sections before any file write occurs.
- Updated hash-mismatch messaging, prompt wording, and tests to reflect hash-only rebase candidates and the new range/section behaviors.
2026-04-30 00:43:27 +02:00
can1357 47dab9ee97 feat: added atom-mode Lid range edits and hashline shifted-hash recovery
- Added atom edit support for Lid ranges, before-anchor inserts, and no-op Lid=TEXT success handling.
- Expanded hashline recovery to scan shifted hashes, match unique alternates, and emit ±5 anchor-shift hints.
- Added edit failure categorization with per-category counts, percentages, and detailed report lines.
- Added regression coverage for shifted-hash recovery, range continuation, cursor shorthand, and split-file atom ops.
2026-04-29 23:51:22 +02:00
can1357 c6a11079f5 feat: added compact atom-mode parser and execution support
- Added compact Lark grammar processing and applied it to OpenAI custom-format tools before conversion.
- Reworked atom mode into `---PATH` compact commands with new grammar, parser, and rm/mv file operations.
- Updated `hline`/`href`/`hrefr` helper behavior and hashline mismatch guidance using shared anchor state.
- Standardized path formatting with `formatPathRelativeToCwd` across LSP, prompts, and edit/search/write tools.
- Added benchmark run-path handling, including `.gitignore` runs mapping, absolute reports, and safer snapshot output.
- Added tests for compact grammar payloads, atom parsing/execution, renderer streaming, and path-list outputs.
2026-04-29 01:16:33 +02:00
can1357 6409d0b1ff deps: updated lockfile and tooling config for dependency management
- Updated package-manager config to use the hoisted linker and relaxed the @types/bun version range.
- Regenerated bun.lock with updated dependency pins, including postcss, safe-buffer, and string_decoder, plus added jszip nested lock entries.
- Updated the check:tools script to run biome checks without the removed tsgo typecheck step.
2026-04-26 15:12:14 +02:00
can1357 4c1a899679 chore: reformat 2026-04-26 09:05:56 +02:00
can1357 5ea1d55e56 feat: removed chunk-mode modules and read/edit entrypoints from pi-natives
- Removed `pi-natives` chunk language classifier modules and all core chunk subsystems (kind, state, render, edit, resolve).
- Removed chunk-mode CLI/read/edit entrypoints, including `read` command and chunk mode registration/prompt tooling.
- Removed chunk selectors from `read` and `grep` tools, switching behavior to raw/L-range handling.
- Fixed poll wait parsing to keep defaulting to `30s` when the provided value is empty.
2026-04-26 08:19:02 +02:00
can1357 ab1a5e2f3e fix(typescript-edit-benchmark): included timeout transport failures as excluded benchmark runs
- Detected transport failures by checking failed run errors for "Timeout exhausted".
- Counted those failures in benchmark summaries and included them in ghost-like run classification, reducing effective run totals accordingly.
- Updated benchmark reporting to display excluded transport-failure runs when present.
2026-04-26 05:12:09 +02:00
can1357 4f85ebe472 fix(coding-agent): corrected gutter marker alignment in diff rendering
- Added `CodeFrameMarker` and `formatCodeFrameLine()` to centralize code-frame gutter formatting.
- Extended diff rendering to preserve `|` and `│` separators, aligning gutter markers and line numbers.
- Reworked AST, grep, hashline, Vim, and diff renderers to use shared line formatting with computed `lineNumberWidth`.
- Updated atom editing flow and tests, including `resolveAtomEntryPaths` migration and new loc-based/edge-case coverage.
- Adjusted benchmark runner early-stop configuration by passing `buildEarlyStop` through prompt collection.
2026-04-26 03:00:51 +02:00
can1357 cb45cba98e feat: implemented LINEID hashline parsing/output with pre/post atom ops
- Expanded hashline and chunk bigram tables to 647 entries and moved chunk checksums to a 40-item namespace.
- Changed hashline anchors from `LINE#ID`/`:` to concatenated `LINEID`\t forms across parsing and tool outputs.
- Removed line-number padding and routed diff/read/grep/renderer output through raw numbers, tabs, and `toDisplayLine` formatting.
- Renamed atom ops to `pre`/`post`, removed `ins`, and updated schemas, prompts, and tests for new insertion behavior.
2026-04-26 02:18:32 +02:00
can1357 1d244df2c5 feat(typescript-edit-benchmark): added configurable early-stop-on-match behavior
- Added a --no-early-stop-on-match CLI option, passed into benchmark configuration as earlyStopOnMatch.
- Added early-stop support that verified expected files after mutation-tool completion and aborted the prompt loop on a match.
- Captured an earlyStopped flag in task results and emitted an early_stop event when match-based termination occurred.
2026-04-26 01:17:42 +02:00
can1357 27848849fb feat: shorter paths in bench runner 2026-04-26 00:37:01 +02:00
can1357 480d40f2a8 fix: corrected optional edit path fallback across chunk hashline modes
- Added raw read output propagation so read/archive commands bypass anchors, line numbers, and chunk formatting.
- Fixed atom/hashline editing by tightening hashline prefixes and applying grouped anchor edits in stable order.
- Hardened chunk parsing and path handling by using `bigram_end` checks and resolving edit paths via shared `args.path` fallback.
- Updated path-related behavior for chunk, replace, patch, and hashline previews to honor optional edit paths.
- Removed mode-level param validators and replaced them with runtime handling, then removed obsolete validation tests.
2026-04-26 00:29:18 +02:00
can1357 60b2fbf1f9 chore(benchmark): cleaned benchmark edit parsing and token accounting
- Relaxed `--edit-variant` parsing in `packages/typescript-edit-benchmark/src/index.ts` to accept any string.
- Adjusted token delta calculation in `runner.ts` to subtract estimated system-prompt overhead per assistant turn.
- Captured initial system-prompt tokens in `runSingleTask` and added rough `estimateTokens` helper for corrected accounting.
- Reframed `scripts/rate-edit-tool.py` prompts to target edit-tool behavior and constraints for each fixture type.
2026-04-26 00:29:17 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 212d56bc11 feat: added strict-mode fallback for OpenAI tool calls with all_strict
- Added `toolStrictMode` support with `all_strict`/`none`/`mixed` options to OpenAI compatibility.
- Fixed OpenAI-completion strict-mode flows by capturing failed HTTP responses and retrying once as non-strict.
- Fixed completion error reporting by surfacing captured status, headers, and JSON `type`/`param`/`code` details.
- Improved strict-schema enforcement with WeakMap memoization and circular-schema detection in sanitization.
- Fixed OpenRouter provider lookup by resolving fallback model IDs for suffix and date variants in registry resolution.
- Refactored benchmark tooling and added async RPC error-window tracking for scheduled run execution.
2026-04-13 15:46:06 +02:00
can1357 c62ab2d53e chore(benchmarks-misc-fixes): cleaned benchmark task/pid validation
- Added `vim` as an edit variant in benchmark CLI/config and rating script coverage.
- Expanded benchmark execution so `vim` is treated as a mutation tool for retries, stats, and edit intent checks.
- Adjusted `TaskTool` output schema precedence so explicit params override agent frontmatter.
- Fixed `TaskTool` success counting by excluding aborted tasks from success totals.
- Improved validation guidance in `SubmitResultTool`/`TodoWriteTool` for clearer recovery when payloads are missing or invalid.
- Added background command PID regression coverage in `executeBash` to confirm a real, terminateable PID is returned.
2026-04-13 12:27:10 +02:00
can1357 debb76723d chore: catalog imports 2026-04-13 01:06:01 +02:00
can1357 d60ee622a4 refactor: restructured hash API and cache tracking across TypeScript and Rust
- Migrated hash API calls from Bun.hash.xxHash64() to Bun.hash() across TypeScript packages for simplified hash generation.
- Consolidated mermaid cache failure tracking by replacing separate failed Set with null values in cache Map.
- Refactored schema-based child extraction in Rust by introducing promotion_fields tracking and schema_wrapper_child() helper function.
- Updated TypeScript configuration files with reformatted arrays and added compiler options for consistency.
2026-04-10 14:26:06 +02:00
can1357 f49098eb76 config: configured PI_STRICT_EDIT_MODE for model-specific edit behavior
- Added PI_STRICT_EDIT_MODE environment variable to control model-specific edit mode defaults.
- Wrapped model-specific edit mode logic behind PI_STRICT_EDIT_MODE condition for conditional behavior.
- Replaced Bun.env direct access with $env utility for consistent environment variable handling.
- Updated rate-edit-tool and typescript-edit-benchmark to set PI_STRICT_EDIT_MODE in test environments.
2026-04-08 21:30:52 +02:00
can1357 52719d1a7c refactor: restructured monorepo TypeScript config and build tasks for unified setup
- Migrated all package tsconfig files to extend tsconfig.workspace.json for unified TypeScript configuration across monorepo.
- Consolidated build and check scripts across 10+ packages to use biome for linting/formatting with separate type checking via tsgo.
- Renamed build scripts from build:native and build:binary to build for simplified command naming across packages/natives and packages/coding-agent.
- Refactored CI workflow to invoke bun tasks instead of inline shell scripts, reducing workflow complexity by 40+ lines.
- Removed sync-exports.ts and repro-stuck.ts scripts; deleted path aliases from tsconfig.base.json in favor of workspace-based configuration.
- Updated turbo.json with new task definitions (check:types, lint, fmt, fix) and removed build:native/embed:native tasks.
2026-04-08 17:05:20 +02:00
can1357 53c78766f3 feat: introduced code-editing tool evaluation framework with multi-model benchmarking
- Added comprehensive code-editing tool evaluation framework with `rate-edit-tool.py` supporting multi-model benchmarking across TypeScript, Rust, Python, and Markdown.
- Enhanced chunk edit error messages to display fresh chunk context with resolved selectors and anchors for improved debugging.
- Added `--no-lsp` flag to benchmark RPC arguments for TypeScript edit task evaluation.
- Improved chunk body boundary calculation to correctly include closing line indentation in epilogue.
- Added comprehensive benchmark results dataset (`all_models_results.json`) with performance metrics for 6 AI models.
- Enhanced chunk-edit documentation with clarified `@body` selector behavior and append/prepend examples.
2026-04-08 07:11:53 +02:00
can1357 a21a542afd refactor(prompt-templates): migrated prompt utilities to pi-utils package
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).
2026-04-08 05:47:35 +02:00
can1357 b93d1929db fix(edit-benchmark): corrected benchmark statistics by filtering ghost runs with zero activity
- Excluded ghost runs (failed runs with zero activity) from benchmark statistics to prevent skewed averages.
- Introduced isGhostRun() helper function to identify and filter runs with no tokens or tool calls.
- Updated report generation and task summarization to use nonGhostRuns for all metric calculations.
2026-04-06 18:48:15 +02:00
can1357 51467abc5a style(coding-agent): standardized indentation and added connection timeout to benchmark runner
- Fixed indentation and formatting across multiple files for consistency.
- Reformatted code blocks to use tabs instead of spaces and improved line breaks for readability.
- Added connection timeout configuration to benchmark runner for early abort on no events.
- Implemented two-phase timeout strategy in benchmark runner with connection and activity phases.
2026-04-06 18:45:41 +02:00
can1357 d1561beb99 refactor(edit-benchmark): migrated to typescript with in-process client support
- Migrated react-edit-benchmark package to typescript-edit-benchmark with pi-mono source repository.
- Added InProcessClient implementation to eliminate subprocess spawning overhead in benchmark runs.
- Extended benchmark configuration with chunk edit variant, retry limits, and conversation dump support.
- Refactored runner.ts to support both RPC and in-process client modes with improved error telemetry.
- Cleaned up 129 benchmark report files from react-edit-benchmark/runs directory.
2026-04-06 18:39:04 +02:00