- Switched HL_EDIT_SEP initialization to read from process.env, trimming the value and defaulting to "~" when missing or invalid.
- Expanded the separator documentation with benchmark data and reliability notes explaining why "~" was selected as the fallback over other candidates.
- Removed todo workflow event handling from `scripts/rate-edit-tool.py` by deleting `Todo*` event plumbing and ops.
- Pruned Go/Markdown fixture coverage from prompts and fixture maps, keeping TypeScript/Rust/Python only.
- Dropped todo-based completion gating in `is_effectively_complete` and removed todo columns from progress output.
- Removed PI_STRICT_EDIT_MODE gating from edit-mode resolution so model fallbacks now always apply.
- Stopped injecting PI_STRICT_EDIT_MODE in edit-benchmark.py and rate-edit-tool.py execution environments.
- Removed PI_STRICT_EDIT_MODE from environment-variable documentation and strict-mode test coverage.
- Drop the modern (x86-64-v3) native build matrix entries for darwin-x64 and
win32-x64. Linux x64 still ships both variants. Loader fallback chain
(modern -> baseline -> default) means hosts on those platforms now load
baseline; npm install via @oh-my-pi/pi-natives and the embedded standalone
binary both keep working unchanged.
- Stop downloading .node files in scripts/install.sh and scripts/install.ps1.
The standalone omp binary embeds the native at compile time, so the
separate downloads were redundant and were going to 404 on platforms with
no modern variant.
- Stop publishing standalone .node files to GitHub Releases. The npm
registry is the only consumer that needs them.
- Drop the release-archives build entirely (omp-*.tar.gz). Nothing
consumed them. Deleted scripts/ci-release-build-archives.ts and the
ci:release:build-archives package script.
- Split the monolithic release job into release-github (download omp-*
binaries, upload to GitHub Release) and release-npm (download natives,
bun publish). Each job's steps now obviously serve its single output.
- Fixed bash interceptor to check both raw and cwd-normalized commands, catching commands hidden behind leading `cd ... &&` wrappers.
- Fixed LSP client shutdown to await graceful shutdown with a 5s timeout before killing the process, and parallelized `shutdownAll` via `Promise.allSettled`.
- Fixed concurrent bash command tracking by replacing a single abort controller with a Set, preventing premature cancellation of parallel commands.
- Removed `./hooks` and `./hooks/*` export entries from the coding-agent package exports map.
- Updated pinned Rust nightly toolchain from `nightly-2026-03-27` to `nightly-2026-04-29` in `rust-toolchain.toml` and CI workflow.
- Replaced custom already-published detection in `ci-release-publish.ts` with `bun publish --tolerate-republish` flag.
Standalone Bun binaries on WSL (and any host where the user moves the
binary away from the build-host's checkout) failed to load
pi_natives.<platform>-<arch>*.node. The loader's isCompiledBinary
detection relied on two signals that are both false in shipped binaries:
process.env.PI_COMPILED (bun --define PI_COMPILED=true substitutes the
bare identifier, not property accesses on process.env) and
__filename.includes("$bunfs") (Bun retains the build-host absolute path
in __filename for required CJS modules — only import.meta.url is
rewritten). Detection therefore returned false, embedded-addon
extraction was skipped, and the only candidates probed were the
build-host nativeDir and execDir.
Make embedded-addon presence the authoritative compiled-mode signal
(it is null in the post-build --reset stub, populated when embed:native
ran for the standalone build), eagerly require the manifest, and
extract candidate-path computation into a pure helper covered by a
host-platform-agnostic unit test. Also fix the build-time --define so
process.env.PI_COMPILED is genuinely set at runtime as a defensive
fallback.
Fixes#823
- Lowered the default session scan limit for edits and tools from 100000 to 1000 and updated usage text to match.
- Reworked grand-total and per-tool table rendering in session-stats to compute aligned column widths and include an aggregated "(N others)" row.
- Tuned the session-stats release profile by replacing full stripping with line-tables-only debug symbols and disabled split debuginfo stripping.
- Updated `formatSessionDumpText` to skip `thinking` entries with empty or whitespace-only content.
- Prevented empty `<thinking>` sections from being emitted in session dump output.
- Required top-level `path` in edit requests, removed per-entry `path` fields, and updated docs/tests to match.
- Updated patch/hashline/atom streaming preview generation to return a single request-level path diff preview.
- Changed edit execution to route patch/replace/atom/hashline through single-path handlers sharing `path`.
- Added `scripts/analyze-edit-formats` Go CLI with reports to audit edit-tool usage from session JSONL logs.
- Removed `pi-natives` chunk language classifier modules and all core chunk subsystems (kind, state, render, edit, resolve).
- Removed chunk-mode CLI/read/edit entrypoints, including `read` command and chunk mode registration/prompt tooling.
- Removed chunk selectors from `read` and `grep` tools, switching behavior to raw/L-range handling.
- Fixed poll wait parsing to keep defaulting to `30s` when the provided value is empty.
- Enabled Biome VCS settings and expanded fix scripts with new `fix:all` and `fix:tools:all` commands.
- Updated `fix:ts` and `fix:tools` to separate changed-only checks from full-run formatting and generation tasks.
- Simplified edit-benchmark variant parsing by removing hardcoded validation logic, and removed deprecated flags from `run-rs-task.ts`.
- Relaxed `--edit-variant` parsing in `packages/typescript-edit-benchmark/src/index.ts` to accept any string.
- Adjusted token delta calculation in `runner.ts` to subtract estimated system-prompt overhead per assistant turn.
- Captured initial system-prompt tokens in `runSingleTask` and added rough `estimateTokens` helper for corrected accounting.
- Reframed `scripts/rate-edit-tool.py` prompts to target edit-tool behavior and constraints for each fixture type.
- Replaced the standalone chunk and vim benchmark scripts with a unified `scripts/edit-benchmark.py` entrypoint.
- Added `--variant` and `PI_EDIT_VARIANT` handling with validation for supported edit modes.
- Generated benchmark specs dynamically per variant with mode-specific prompts and retry instructions.
- Added a Go fixture and registered it as `main.go` in the reference files, fixture list, and descriptions.
- Changed prompt construction and workspace setup to exercise every fixture in a single session instead of one file at a time.
- Updated run identifiers and output filenames, and set run results to report a fixture scope of `all`.
- Added a rerun flow that synthesizes oracle findings from existing `review_*.md` files without re-running fixtures.
- Refactored oracle input handling to source `(model, fixture, path)` tuples, and persisted oracle prompt and synthesis outputs in the results directory.
- Removed `OPENROUTER_API_KEY` bootstrapping and passed environment, and now write oracle errors to `oracle_error.txt` on failure.
Compiled Bun binaries unconditionally load bunfig.toml and .env from
the current working directory at runtime, before any application code
runs. Since omp is a coding agent that runs from arbitrary project
directories, it picks up foreign project configs -- most critically
preload directives that cause immediate crashes, but also a potential
security issue since preloads execute arbitrary code.
Add --no-compile-autoload-bunfig and --no-compile-autoload-dotenv to
the bun build --compile invocations (available since Bun v1.3.3).
Note: source installs via 'bun install -g' are still affected because
'bun run' has no equivalent flag. That remains an open issue.
- Updated root package.json @oh-my-pi entries from workspace:* to pinned 14.1.1 versioned dependencies.
- Modified the release script to automatically rewrite @oh-my-pi/* catalog versions in package.json using the release version.
- Kept the update operation before Rust workspace version bump in the release workflow.
- Removed the standalone vim tool and normalized built-in/requested tooling to edit.
- Updated session and SDK tool activation to dedupe lowercase names and track edit state via the edit key.
- Added vim-mode argument detection and delegated edit rendering/execution into Vim handlers under edit.
- Updated Vim step handling to auto-reorder numeric-positioned commands, including cc/C/S/s/i/I/A cases.
- Renamed prompt/changelog text and test expectations to reflect edit-only tool naming and usage.
- Fixed Vim path normalization to allow colon-prefixed targets without raising ToolError.
- Documented the new Vim path behavior in the package changelog.
- Updated chunk-edit benchmark prompts to read `test.rs` for edit and retry operations.
- Updated vim benchmark flow to use `vim`+`read`, import expected content, and display final expected output.
- Expanded benchmark fixtures to generate Rust `test.rs`, compute unified diffs, and add richer pool logic.
- Changed chunk edit normalization to prefer write operations, then replace, insert, then delete, with empty write values now treated as clear-content writes.
- Allowed space motions in vim input as `<Space>`, handled them as movement and rendered in error output via token display.
- Adjusted benchmark retry flow to reset files on each retry, reduced per-run call limits, and lowered the default per-turn timeout.
- Added rollback handling for pending INSERT-mode changes whenever a non-final kbd sequence leaves insert mode, and updated the resulting VimInputError with guidance for using `insert` and escaping insert transitions.
- Hardened VimTool execution by resetting stale insert state before processing commands and by only applying empty inserts when Vim remains in INSERT mode.
- Adjusted Vim search handling to mimic Vim magic escaping and taught `o`/`O` numeric prefixes to act like `Go`/`GO` line inserts, then updated the expected error message test.
- Added `toolStrictMode` support with `all_strict`/`none`/`mixed` options to OpenAI compatibility.
- Fixed OpenAI-completion strict-mode flows by capturing failed HTTP responses and retrying once as non-strict.
- Fixed completion error reporting by surfacing captured status, headers, and JSON `type`/`param`/`code` details.
- Improved strict-schema enforcement with WeakMap memoization and circular-schema detection in sanitization.
- Fixed OpenRouter provider lookup by resolving fallback model IDs for suffix and date variants in registry resolution.
- Refactored benchmark tooling and added async RPC error-window tracking for scheduled run execution.
- Added `<escape>` and `<return>` as special-key aliases in `parser.ts`, mapping to `Esc` and `CR`.
- Clarified Vim tool behavior so `open` replaces the active buffer and non-paused `kbd` calls auto-save.
- Updated the Vim prompt contract to treat `insert` as raw text entered only after entering INSERT mode.
- Increased `scripts/vim-edit-benchmark.py` `request_timeout` from 60s to 120s for longer edit runs.
- Added `vim` as an edit variant in benchmark CLI/config and rating script coverage.
- Expanded benchmark execution so `vim` is treated as a mutation tool for retries, stats, and edit intent checks.
- Adjusted `TaskTool` output schema precedence so explicit params override agent frontmatter.
- Fixed `TaskTool` success counting by excluding aborted tasks from success totals.
- Improved validation guidance in `SubmitResultTool`/`TodoWriteTool` for clearer recovery when payloads are missing or invalid.
- Added background command PID regression coverage in `executeBash` to confirm a real, terminateable PID is returned.
- Added `VimTool` in `src/tools/vim.ts` with `open`, `kbd`, `insert`, and `pause` actions.
- Implemented `src/vim/{buffer,engine,parser,commands,render,types}.ts` for interactive Vim modes and operators.
- Updated `src/tools/index.ts` to normalize `edit`/`vim` tool selection and skip the inactive edit variant.
- Added `vim` registration in `BUILTIN_TOOLS` and `toolRenderers` for discoverability and output formatting.
- Added streaming renderer snapshots with viewport, caret focus, and diff output for `vim` calls.
- Added `test/tools/vim.test.ts` and `scripts/vim-edit-benchmark.py` coverage for the new editor stack.
- Removed Turborepo cache steps from CI and dropped .turbo from gitignore.
- Migrated root scripts from turbo commands to bun workspace execution for build, test, lint, check, format, and fix.
- Removed turbo dependencies and deleted turbo.json task configuration files from the repository.
- Passed the user context when setting the session name during subprocess execution.
- Updated native build scripts to import detectHostAvx2Support from the correct shared module paths.
- Added `/rename <title>` slash command to set explicit session names and update header/tab titles.
- Added `session_name` status segment with hash-derived accent color for session titles.
- Fixed shell execution failure sanitization to preserve all `execResult` fields after redacting stderr.
- Fixed tool execution completion output to pass original `toolResult` text instead of sanitized `content`.
- Added SHORT_LIVED_GIT_CONFIG constants and withShortLivedGitConfig helper for short-lived overrides.
- Updated runCommand to route git args through short-lived config normalization and avoid duplicate -c entries.
- Added scripts/release git() helper and replaced raw release git invocations with config-safe calls.
- Added git-process-config tests asserting status and stage spawn git with disabled core.fsmonitor/untrackedCache.
- Restructured blank-line cleanup logic in chunk editor to correctly handle artifacts from deletions before and after delimiters.
- Refactored rate-edit-tool.py to support per-fixture model runs instead of single model-wide runs, enabling parallel evaluation across multiple code fixtures.
- Extracted fixture metadata into FIXTURES tuple and introduced build_fixture_prompt() helper to customize prompts per fixture.
- Updated ModelRunRecorder and ProgressPrinter to track run_id (model + fixture) separately from model name, enabling independent progress tracking.
- Modified materialize_workspace() to optionally create single-fixture workspaces and updated run_model_sync() to generate fixture-scoped workspace paths and result files.
- Enhanced chunk-edit tool documentation with clarifications on @decl region behavior and guidance for attribute/decorator modifications.
- Improved chunk edit implementation to handle @body region scanning and markdown blank-line preservation with additional test coverage.
- Added oracle model review synthesis to rate-edit-tool.py for aggregating findings across multiple model reviews.
- Updated test infrastructure to run TypeScript and Rust tests in parallel using cargo nextest with improved logging control.
- Updated TypeScript dependencies including @typescript/native-preview and typescript to latest versions.
- Added PI_STRICT_EDIT_MODE environment variable to control model-specific edit mode defaults.
- Wrapped model-specific edit mode logic behind PI_STRICT_EDIT_MODE condition for conditional behavior.
- Replaced Bun.env direct access with $env utility for consistent environment variable handling.
- Updated rate-edit-tool and typescript-edit-benchmark to set PI_STRICT_EDIT_MODE in test environments.
- Migrated all package tsconfig files to extend tsconfig.workspace.json for unified TypeScript configuration across monorepo.
- Consolidated build and check scripts across 10+ packages to use biome for linting/formatting with separate type checking via tsgo.
- Renamed build scripts from build:native and build:binary to build for simplified command naming across packages/natives and packages/coding-agent.
- Refactored CI workflow to invoke bun tasks instead of inline shell scripts, reducing workflow complexity by 40+ lines.
- Removed sync-exports.ts and repro-stuck.ts scripts; deleted path aliases from tsconfig.base.json in favor of workspace-based configuration.
- Updated turbo.json with new task definitions (check:types, lint, fmt, fix) and removed build:native/embed:native tasks.
- Added bounds clamping to prologue and epilogue byte calculations to prevent out-of-range boundary violations.
- Extended chunk boundaries for indent-based languages when epilogue exceeds calculated range with trailing newline.
- Added comprehensive test coverage for Python chunk editing operations including body/head replacement and indentation preservation.
- Extracted working directory formatting logic into reusable utility function and applied tab sanitization to bash command previews.
- Added Auto QA tool (`report_tool_issue`) for automated tracking of unexpected tool behavior with environment variable and setting support.
- Added Python tool environment warmup on first execution to ensure prelude helpers are available before use.
- Fixed Python prelude introspection to respect execution timeout and signal options, preventing hangs.
- Refactored prelude documentation caching and loading logic into reusable helper functions with test environment awareness.
- Enhanced kernel introspection with optional timeout and signal parameters for better execution control.
- Added system prompt guidance to encourage agents to report tool issues via Auto QA when available.
- Added auto-retry event tracking and message-end event handling to improve agent completion detection.
- Implemented is_effectively_complete() method to detect agent completion based on review sections, todo state, and quiet period.
- Enhanced wait_for_settle() logic to handle auto-retry delays and graceful timeout recovery instead of immediate failure.
- Added token usage tracking from partial message updates to capture intermediate token counts.
- Refactored note_tool_end() to only update activity on error, removing redundant success case.
- Removed last_activity updates from todo reminder and auto-clear handlers to simplify state management.
- Added comprehensive code-editing tool evaluation framework with `rate-edit-tool.py` supporting multi-model benchmarking across TypeScript, Rust, Python, and Markdown.
- Enhanced chunk edit error messages to display fresh chunk context with resolved selectors and anchors for improved debugging.
- Added `--no-lsp` flag to benchmark RPC arguments for TypeScript edit task evaluation.
- Improved chunk body boundary calculation to correctly include closing line indentation in epilogue.
- Added comprehensive benchmark results dataset (`all_models_results.json`) with performance metrics for 6 AI models.
- Enhanced chunk-edit documentation with clarified `@body` selector behavior and append/prepend examples.
* fix(ai): update spoofed Gemini CLI User-Agent to v0.35.3 format
Gemini CLI v0.35+ changed its User-Agent format:
- Dropped 'google-api-nodejs-client/10.5.0' suffix
- Added surface field (third param in parens)
- New format: GeminiCLI/VERSION/MODEL (PLATFORM; ARCH; SURFACE)
The stale v0.34.0 UA was contributing to google-gemini-cli/gemini-3.1-pro-preview
being rejected, alongside Google's new abuse detection for third-party CCA usage
(google-gemini/gemini-cli#22970).
Verified: model responds successfully with updated UA.
* feat(scripts): add spoofed version drift checker
Fetches latest stable release from google-gemini/gemini-cli GitHub
releases and compares against the hardcoded version in the provider.
bun check-spoofed-versions # report drift, exit 1 on mismatch
bun check-spoofed-versions --update # apply version bump in-place
Extensible for additional spoofed tools (Antigravity, Copilot headers).
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.