61 Commits

Author SHA1 Message Date
metaphorics 3802d2dd80 chore(ts): enforce noImplicitOverride 2026-08-05 02:49:01 +09:00
can1357 a4773c0baa refactor(typescript-edit-benchmark): updated mutation plans
- Update benchmark fixtures archive file.
- Adjust mutation plan block sizes and counts in generator script.
- Remove postmortem quit test.
2026-07-30 07:57:55 +02:00
can1357 38ebd30337 chore: update stale tests 2026-07-30 07:49:43 +02:00
can1357 9856f904d7 feat(typescript-edit-benchmark): introduced empirical edit mutation planning
- Add new structural, multi-edit, and block-level mutation classes with updated category mappings.
- Introduce hunk extraction, placement, rendering, and solver utilities along with unit tests.
- Implement size-based mutation planning, prompt validation logic, and new prompt markdown templates.
- Update benchmark generation scripts and package configurations to support empirical edit shape statistics.
2026-07-30 07:45:01 +02:00
can1357 a6001c04a3 fix(coding-agent): prevented race condition in concurrent createAgentSession calls
- Export AgentRegistry from the SDK to allow passing a private registry instance.
- Provide a dedicated AgentRegistry per in-process client in the benchmark runner.
2026-07-30 06:03:32 +02:00
can1357 52ad6516ef feat(diff): implemented native UTF-16 diff processing and removed jsdiff
- Implemented native UTF-16 text processing in Rust diff module with support for unpaired surrogates.
- Removed `similar` crate from Rust workspace and `diff` npm package from coding-agent, hashline, and natives.
- Removed jsdiff fallback wrappers and `isWellFormed()` guards from TypeScript diff implementations.
- Added comprehensive test suite for native diff functions covering random inputs and edge cases including surrogates and emoji.
- Renamed model `codex-auto-review` to `gpt-5.3-codex-spark` with updated pricing and context window.
2026-07-23 01:35:31 +02:00
can1357 0002905ec5 chore: untrack node_modules symlinks and harden ignore pattern
- An eval-worktree cherry-pick swept 16 packages/*/node_modules symlinks into the index; 'node_modules/' with a trailing slash only matches directories, so symlinked installs bypassed the ignore. Dropped the slash and removed the tracked links.
2026-07-22 21:43:22 +02:00
can1357 b9072f1991 fix(extensibility): validate Type.Unsafe against the draft-2020-12 upgraded schema
Aligns the shim's runtime safeParse/__validator with the wire/tool-call
path, so legacy draft-07 documents (tuple items) accept the same values
validateToolArguments does. Adds a regression test.
2026-07-22 21:13:21 +02:00
can1357 0856055dfe feat(harbor-manager): unified benchmark normalization and reporting
- Implemented a unified benchmark normalization layer to handle metrics and artifacts across harbor, edit, and snapcompact benchmarks.
- Integrated the TypeScript edit benchmark directly into the manager, migrating logic and deprecating the standalone package.
- Updated the API and database schema to support standardized benchmark configurations, metrics, and trace-based reporting.
- Enhanced the UI to visualize comparative performance metrics, including pass rate, cost, and latency deltas for benchmark runs.
2026-07-13 04:21:54 +02:00
can1357 d435385ab1 feat: introduced max reasoning effort tier across model and rpc systems
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
2026-07-10 13:39:42 +02:00
can1357 18f3386aed feat(typescript-edit-benchmark): tracked reasoning tokens
- Added reasoning token counts to benchmark results and reporting.
- Updated session statistics and task summaries to include reasoning metrics.
- Updated report generation to display reasoning token breakdown.
2026-06-28 07:27:02 +02:00
can1357 5a044dc0da feat: consolidated and automate changelog management
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
2026-06-27 03:08:52 +02:00
can1357 ac904fc70c fix: patched Kimi model edit mode fallback
- Added a fallback from `hashline` to `replace` mode for Kimi-family models to resolve compatibility issues.
- Introduced `PI_STRICT_EDIT_MODE` environment variable to bypass automatic model-specific edit-mode fallbacks.
- Updated `getEditVariantForModel` to perform case-insensitive matching for model variant configurations.
- Added comprehensive unit tests for edit mode resolution and settings configuration.
2026-06-26 13:17:35 +02:00
can1357 5bce7ed6df feat: added advisory transcript formatting and one-shot benchmark metrics
- Introduced advisory note output as `<advisory>` tags with optional severity and guidance.
- Updated session transcript formatting to `### Session update` and inline watched role labels.
- Added shared `escapeXmlText` utility and escaped XML-sensitive text in advisor outputs.
- Added one-shot success run token metrics and one-shot statistics reporting.
2026-06-16 18:34:50 +02:00
can1357 8b7dd10a8a feat: added native tool inventory rendering with TypeScript signatures
- Added `jsonSchemaToTypeScript` and `renderToolInventory` to generate tool blocks with TypeScript signatures.
- Added `examples` and `TSchema` fields to dump-tool metadata and passed them through prompt rendering.
- Changed Harmony invocation rendering to omit `<|constrain|>json` markers in tool call payloads.
- Added compact native tool list-mode inventory rendering with full `# Tool:` output elsewhere.
2026-06-15 08:12:58 +02:00
can1357 fbba331f8a feat(cross-cutting): added multi-syntax in-band tool-call support for runtime tool conversion
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
2026-06-15 07:33:24 +02:00
can1357 08a941a14e feat: added standalone snapcompact package and model-specific frame shaping
- Added a new @oh-my-pi/snapcompact package and redirected compaction call sites to it.
- Added provider-aware snapcompact shape resolution for model-specific mixed-frame behavior.
- Added optional image detail support by extending ImageContent and passing hints through OpenAI providers.
- Added native snapcompact render options, including 5x8/8x8 font loading and palette/geometry controls.
2026-06-10 21:50:03 +02:00
can1357 9d457f73d9 test: migrated test imports to package subpath exports
- Replaced relative `../src` imports with `@oh-my-pi/pi-ai` and `@oh-my-pi/pi-agent-core` subpaths.
2026-06-08 19:03:55 +02:00
can1357 c8b4bf09c7 prompts: undo experiment, update system 2026-06-04 17:41:10 +02:00
can1357 95defa712b docs(compaction): rewrote prompts in terse scratchpad style
- Converted compaction, branch, and handoff prompts to fragment voice.
- Replaced "You MUST" phrasing with bare "MUST" directives.
- Applied same rewrite to autoresearch and turn-aborted prompts.
2026-06-04 14:49:45 +02:00
can1357 570b439b31 chore: enabled parallel Bun test runs across package scripts
- Updated each package's test script to run `bun test` with the `--parallel` flag.
- Updated the tui package test command to preserve its `test/*.test.ts` file filter while enabling parallel execution.
2026-05-31 05:24:51 +02:00
can1357 06dc3976f0 fix(coding-agent): probed all Python runtimes to bypass broken managed env
- Replaced single-candidate resolution with `enumeratePythonRuntimes`, returning venv, managed env, and system interpreter in priority order.
- Availability check now probes each candidate and falls through to the first that executes, so a stale `uv`-managed Python no longer fails the whole session.
- Kernel spawn reuses the probed runtime from the availability result instead of re-resolving independently.
- Expanded test coverage for enumeration, fallback, and env isolation between candidates.
2026-05-31 02:00:18 +02:00
can1357 a1ba50b4da feat(benchmark): added median, p1, and p99 token distribution stats
- Added `percentile` and `summarizeTokenDistribution` helpers to runner.
- Extended `BenchmarkSummary` with `medianTokensPerTask`, `p1TokensPerTask`, and `p99TokensPerTask`.
- Updated live progress output and markdown report table to show distribution columns.
- Added unit tests covering percentile interpolation and summary fields.
2026-05-31 01:14:52 +02:00
can1357 64daece638 fix(pi-natives): handled legacy Alt+letter pairs in mixed enhanced-keyboard mode
- Preserved Alt and Ctrl+Alt letter ESC-prefix matching when kitty_protocol_active is true to support mixed tmux and Kitty keyboard modes.
- Parsed two-byte ESC sequences before legacy sequence lookup so mixed-mode Meta pairs are treated as Alt letter keys instead of legacy aliases.
- Updated native and TUI key tests to verify Alt+letter and Alt+Shift+letter parsing and matching while enhanced mode is active.

Fixes #1511
2026-05-29 18:21:30 +02:00
can1357 7c64576524 feat(hashline): replaced file-hash anchors with opaque snapshot-store tags
- Replaced 4-hex content-derived file hashes with 3-hex opaque tags minted by InMemorySnapshotStore, making tags session-bound pointers rather than content fingerprints.
- Removed lru-cache dependency; replaced LRU-bounded per-path rings with a flat 4096-slot global ring using a scrambled permutation to prevent LLM tag extrapolation.
- Made SnapshotStore required in Patcher (was optional); tag resolution now drives stale-anchor detection instead of recomputing hashes at apply time.
- Changed literal payload sigil from `|` to `+` and accepted `^A` shorthand for `^A-A`; added lenient recovery for bare bodies, lone `-` rows, and overlapping bare/concrete block pairs.
2026-05-28 01:00:23 +02:00
can1357 2ef4eb7931 refactor(typescript-edit-benchmark): restructured line pairing to use diff-based matching
- Replaced line-by-line comparison with diff-based pairing to correctly handle insertions and deletions that shift line indices.
- Added logic to preserve trailing newline semantics from the actual content.
- Added test case verifying that indent-only changes are normalized even when earlier insertions shift line positions.
2026-05-27 13:34:19 +02:00
can1357 30793c1655 refactor: restructured hashline to use file-level hash validation with colon separators
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.
2026-05-26 13:25:22 +02:00
can1357 8b41c87ded fix(typescript-edit-benchmark): resolved ts-edit-benchmark exit behavior
- Handled successful `main()` resolution by calling `process.exit(0)`.
- Preserved existing benchmark failure handling by logging the error and exiting with status 1.
2026-05-26 12:31:49 +02:00
can1357 7282d92bf4 fix: normalized line-ending handling for text and TUI output flows
- Unified line-ending normalization to `replace(/\r\n?/g, "\\n")` in editor, scraper, benchmark, and utils modules.
- Added terminal-aware line sanitization in code-cell rendering to collapse inline carriage returns and avoid overwrite corruption.
- Tightened editor and paste sanitizers to trim control characters consistently after CR normalization.
2026-05-15 23:46:23 +02:00
can1357 f1f6516056 refactor: reorganized exports and removed obsolete helper branches
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
2026-05-14 04:36:19 +02:00
can1357 af07faec11 feat: Bun 1.3.14
- Raised the Bun minimum version to >=1.3.14 across package metadata, install scripts, and changelog notes.
- Removed the Photon native image pipeline and added SIXEL-based `sixel` support in pi-natives.
- Migrated coding-agent image handling and resizing to `Bun.Image`, including updated tests and a JPEG quality bump to 80.
- Added HTTP/2 fetch bootstrap with HTTPS-only fallback and updated Bun build flags for autoload suppression/`--keep-names`.
2026-05-13 13:21:59 +02:00
can1357 975941aba4 chore: remove garbage tests 2026-05-12 04:09:33 +02:00
can1357 8c323666be feat: added ordered systemPrompt arrays and normalized context prompts
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
2026-05-04 15:20:26 +02:00
can1357 6e8daae059 feat(coding-agent): added inline hashline parse with conflict checks
- Added inline hashline parse and apply support for `<` prepend and `+` append operations with prefix+suffix edits.
- Added fail-fast behavior to reject inline modify ops combined with delete or replace on same line.
- Renamed HASHLINE_* and mode symbols to HL_* in prompt tooling, read/search checks, and prompt templates.
- Standardized separators to `PI_HL_SEP`/`HL_EDIT_SEP` and fixed `HL_BODY_SEP='|'`, updating parser formatting behavior.
- Updated benchmark subtype constants and python cleanup test setup to use HL_* values and AgentRegistry mock failure injection.
2026-05-03 08:17:09 +02:00
can1357 ffdec04af2 refactor(packages/coding-agent): reorganized atom LID parsing rules
- Centralized `line-hash.ts` hash regex sources and resolved `atom.lark` via `resolveLarkLidPlaceholders`.
- Expanded `computeLineHash` to emit `>[a-z]` and `[a-z]<` hashes for brace-context anchors.
- Replaced atom/hashline parsers' hard-coded lid regex with shared `HASHLINE_HASH_RE_SRC` and lax counterparts.
- Removed `\\TEXT` continuation handling in atom rewrites and switched multi-line replacements to `+TEXT`.
- Added brace-body insertion warning when `@Lid` on `{`-ending lines inserts at non-body-safe indent.
- Suppressed duplicate auto-rebase warnings and kept unmatched `-`/`+` ranges separate in compact previews.
2026-05-02 04:34:39 +02:00
can1357 fbe051bcbd fix(ai): fixed OpenRouter cache write attribution in usage parsing
- Updated parseChunkUsage to subtract prompt_tokens_details.cache_write_tokens from prompt token input so OpenRouter write tokens are not misclassified as billable input.
- Set cacheWrite and total token counts to include cache-write usage, while preserving cache-read behavior from existing cached_tokens handling.
- Added OpenRouter attribution tests verifying cacheWrite and cacheRead totals for write-heavy and cache-warm prompts.
2026-04-30 02:24:57 +02:00
can1357 863560ffb6 fix(coding-agent/edit): added atom range-repair and multi-section preflight checks
- Allowed bare `LidA..LidB` to recover a missing-range-delete typo and accepted `|` as a legacy range replacement separator while validating ranges and replacement text.
- Enabled indented hashline statements to parse as replacement edits and added a preflight pass that validates all atom sections before any file write occurs.
- Updated hash-mismatch messaging, prompt wording, and tests to reflect hash-only rebase candidates and the new range/section behaviors.
2026-04-30 00:43:27 +02:00
can1357 47dab9ee97 feat: added atom-mode Lid range edits and hashline shifted-hash recovery
- Added atom edit support for Lid ranges, before-anchor inserts, and no-op Lid=TEXT success handling.
- Expanded hashline recovery to scan shifted hashes, match unique alternates, and emit ±5 anchor-shift hints.
- Added edit failure categorization with per-category counts, percentages, and detailed report lines.
- Added regression coverage for shifted-hash recovery, range continuation, cursor shorthand, and split-file atom ops.
2026-04-29 23:51:22 +02:00
can1357 c6a11079f5 feat: added compact atom-mode parser and execution support
- Added compact Lark grammar processing and applied it to OpenAI custom-format tools before conversion.
- Reworked atom mode into `---PATH` compact commands with new grammar, parser, and rm/mv file operations.
- Updated `hline`/`href`/`hrefr` helper behavior and hashline mismatch guidance using shared anchor state.
- Standardized path formatting with `formatPathRelativeToCwd` across LSP, prompts, and edit/search/write tools.
- Added benchmark run-path handling, including `.gitignore` runs mapping, absolute reports, and safer snapshot output.
- Added tests for compact grammar payloads, atom parsing/execution, renderer streaming, and path-list outputs.
2026-04-29 01:16:33 +02:00
can1357 6409d0b1ff deps: updated lockfile and tooling config for dependency management
- Updated package-manager config to use the hoisted linker and relaxed the @types/bun version range.
- Regenerated bun.lock with updated dependency pins, including postcss, safe-buffer, and string_decoder, plus added jszip nested lock entries.
- Updated the check:tools script to run biome checks without the removed tsgo typecheck step.
2026-04-26 15:12:14 +02:00
can1357 4c1a899679 chore: reformat 2026-04-26 09:05:56 +02:00
can1357 5ea1d55e56 feat: removed chunk-mode modules and read/edit entrypoints from pi-natives
- Removed `pi-natives` chunk language classifier modules and all core chunk subsystems (kind, state, render, edit, resolve).
- Removed chunk-mode CLI/read/edit entrypoints, including `read` command and chunk mode registration/prompt tooling.
- Removed chunk selectors from `read` and `grep` tools, switching behavior to raw/L-range handling.
- Fixed poll wait parsing to keep defaulting to `30s` when the provided value is empty.
2026-04-26 08:19:02 +02:00
can1357 ab1a5e2f3e fix(typescript-edit-benchmark): included timeout transport failures as excluded benchmark runs
- Detected transport failures by checking failed run errors for "Timeout exhausted".
- Counted those failures in benchmark summaries and included them in ghost-like run classification, reducing effective run totals accordingly.
- Updated benchmark reporting to display excluded transport-failure runs when present.
2026-04-26 05:12:09 +02:00
can1357 4f85ebe472 fix(coding-agent): corrected gutter marker alignment in diff rendering
- Added `CodeFrameMarker` and `formatCodeFrameLine()` to centralize code-frame gutter formatting.
- Extended diff rendering to preserve `|` and `│` separators, aligning gutter markers and line numbers.
- Reworked AST, grep, hashline, Vim, and diff renderers to use shared line formatting with computed `lineNumberWidth`.
- Updated atom editing flow and tests, including `resolveAtomEntryPaths` migration and new loc-based/edge-case coverage.
- Adjusted benchmark runner early-stop configuration by passing `buildEarlyStop` through prompt collection.
2026-04-26 03:00:51 +02:00
can1357 cb45cba98e feat: implemented LINEID hashline parsing/output with pre/post atom ops
- Expanded hashline and chunk bigram tables to 647 entries and moved chunk checksums to a 40-item namespace.
- Changed hashline anchors from `LINE#ID`/`:` to concatenated `LINEID`\t forms across parsing and tool outputs.
- Removed line-number padding and routed diff/read/grep/renderer output through raw numbers, tabs, and `toDisplayLine` formatting.
- Renamed atom ops to `pre`/`post`, removed `ins`, and updated schemas, prompts, and tests for new insertion behavior.
2026-04-26 02:18:32 +02:00
can1357 1d244df2c5 feat(typescript-edit-benchmark): added configurable early-stop-on-match behavior
- Added a --no-early-stop-on-match CLI option, passed into benchmark configuration as earlyStopOnMatch.
- Added early-stop support that verified expected files after mutation-tool completion and aborted the prompt loop on a match.
- Captured an earlyStopped flag in task results and emitted an early_stop event when match-based termination occurred.
2026-04-26 01:17:42 +02:00
can1357 27848849fb feat: shorter paths in bench runner 2026-04-26 00:37:01 +02:00
can1357 480d40f2a8 fix: corrected optional edit path fallback across chunk hashline modes
- Added raw read output propagation so read/archive commands bypass anchors, line numbers, and chunk formatting.
- Fixed atom/hashline editing by tightening hashline prefixes and applying grouped anchor edits in stable order.
- Hardened chunk parsing and path handling by using `bigram_end` checks and resolving edit paths via shared `args.path` fallback.
- Updated path-related behavior for chunk, replace, patch, and hashline previews to honor optional edit paths.
- Removed mode-level param validators and replaced them with runtime handling, then removed obsolete validation tests.
2026-04-26 00:29:18 +02:00
can1357 60b2fbf1f9 chore(benchmark): cleaned benchmark edit parsing and token accounting
- Relaxed `--edit-variant` parsing in `packages/typescript-edit-benchmark/src/index.ts` to accept any string.
- Adjusted token delta calculation in `runner.ts` to subtract estimated system-prompt overhead per assistant turn.
- Captured initial system-prompt tokens in `runSingleTask` and added rough `estimateTokens` helper for corrected accounting.
- Reframed `scripts/rate-edit-tool.py` prompts to target edit-tool behavior and constraints for each fixture type.
2026-04-26 00:29:17 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00