- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Added Anthropic prompt-cache refresh scheduling and state management to keep prompts warm across idle sessions.
- Updated pricing models and database stats tracking to calculate and store cost-weighted cache savings.
- Integrated cache savings metrics and efficiency displays into the stats CLI, dashboard routes, and UI components.
- Added support for package renaming, manifest pointer tracking, and installation migration during CLI updates.
- Add comprehensive Nix flake definitions, derivations, modules, and CI workflows.
- Update tests and executables to resolve binaries from PATH rather than absolute paths.
- Ensure byte reproducibility and zeroed timestamps in embedded dashboard archives.
- Add handling for Nix-managed installations in CLI update checks.
Four tests spawn the full CLI entry graph (or compile a standalone binary)
and declared no timeout, so they inherited Bun's 5s default. Spawning
`src/cli.ts` costs ~900ms warm on a fast machine and ~3.1s cold, so the
budget is spent almost entirely on transpile. When CI runs the native
bucket with OMP_TEST_CONCURRENCY=4, a cold spawn on a contended runner
crosses 5s and the test fails with `timed out after 5000ms` plus a
trailing `killed 1 dangling process` - the subprocess was still alive when
the timeout fired.
Reproduced locally by oversubscribing the box (48 concurrent runs of the
same chunk): 48/48 failed with the identical signature, while 4-way
concurrency - what CI actually configures - passed every time. Only the
subprocess tests starve; the pure-unit tests in the same files pass.
Timeouts are sized to the work, matching existing subprocess tests
(read-cli-mcp-resource 30_000, acp-stdout-hygiene 60_000): 30s for CLI
spawns, 60s for the `bun build --compile` case. After the change, 64
concurrent runs of all three files produce zero timeouts.
`profile-cli.test.ts` gets the timeout on both spawn tests. Only the first
was observed failing, because it warms the transpile cache for its
sibling - that ordering is incidental and would flake if it changed.
These are drift and wiring assertions, not latency assertions, so a
generous ceiling costs nothing on a healthy run.
`omp ttsr test <file>` inferred the match source from the path extension
against a hardcoded allowlist (`SOURCE_FILE_EXT`); a supplied source file
whose extension was absent silently fell through to the text (prose)
context, where tool-scoped rules can never match. The result was a false
negative indistinguishable from a non-matching regex, contradicting the
documented "a positional that resolves to a file defaults to tool/edit
context" contract.
- Emit an explanatory note when a resolvable file path is supplied and the
source is inferred as `text`, pointing at `--source tool --tool edit`.
Surfaced in both text and `--json` output via `TestReport.inferenceNote`.
- Extend the allowlist with the .NET family and other common source
languages (cs, razor, cshtml, fs, fsx, vb, sh, bash, sql, zig, dart,
scala, ex, exs, proto, tf).
Left the fall-through default itself unchanged (inverting the test is a
behaviour change and a maintainer call).
Fixes#6887
A crashed owner's pid can be recycled by an unrelated long-lived
process, so kill(pid, 0) succeeds and the leftover sandbox was pinned
live forever, unreachable by a non-`--all` clear.
The ownership marker now records a process-instance start-time token
alongside the pid (Linux /proc/<pid>/stat field 22, other Unix via
`ps -o lstart`). A live pid whose current token no longer matches the
recorded one is a recycled pid and counts as dead; platforms that can't
report a token degrade to the prior pid-only check.
Fixes#6761
The setup window between writeIsolationOwner and isoStart left the base
dir holding only the marker file and no `m` mount, so classifyDir
returned null and scanWorktrees classified it as a stray — which a
non-`--all` clear removes, defeating the ownership guard mid-setup.
classifyDir now treats the presence of the ownership marker as a
task-isolation signal (in addition to the mount dir), so an in-progress
sandbox with a live owner is preserved throughout backend setup.
Fixes#6761
`omp worktree clear` (without `--all`) removed every task-isolation dir
under the worktree base, including sandboxes owned by subagents running
right now, and the "no live task owns it" reason was asserted from the
mere presence of the `m` mount dir with no ownership check.
`ensureIsolation` now stamps each sandbox base dir with a pid-bearing
ownership marker before the backend materialises `m`, and the worktree
scanner classifies a sandbox as live while its owning process is alive,
so `clear` reclaims only crashed leftovers.
Fixes#6761
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
- Added `off` and `auto` as valid inputs for the `--thinking` CLI flag.
- Centralized thinking level definitions in `CLI_THINKING_LEVELS` to keep flag options, shell completions, and validation in sync.
- Configured CLI parsing to reject `inherit` as an explicit input to prevent unintended configuration suppression.
- Standardized all test temporary file handling to use the system temp directory (`os.tmpdir`) instead of creating artifacts inside the source tree.
- Added lifecycle management to `ttsr-cli.test.ts` to ensure temporary directories are created once per test run and cleaned up automatically.
- Updated file generation paths in CLI and reproduction tests to prevent local state pollution.
- Added the `scan` action to `ttsr` CLI with support for gitignore-aware file globbing.
- Integrated AST pre-filtering and optimized AST/regex matching to evaluate scan rules.
- Implemented file size limits, binary file detection, and custom `--no-gitignore` and `--max-bytes` flags.
- Introduced comprehensive test suites validating directory mapping, exclusions, and size-limit enforcement.
- Added a top-level `ttsr` CLI command with `list` and `test` actions.
- Added snippet input handling for inline text, `--file` path, and stdin via `--file -`.
- Added test-mode context inference and result output for matched and unmatched rules.
- Added CLI tests for source inference, explicit source overrides, JSON output, and list mode.
- Added `omp completions ` command generating scripts from live command/flag metadata.
- Added hidden `omp __complete` helper for dynamic model and session candidates.
- Completions never drift from the CLI: flags, enums, and subcommands are derived from static descriptors.
- Removed `pi-natives` chunk language classifier modules and all core chunk subsystems (kind, state, render, edit, resolve).
- Removed chunk-mode CLI/read/edit entrypoints, including `read` command and chunk mode registration/prompt tooling.
- Removed chunk selectors from `read` and `grep` tools, switching behavior to raw/L-range handling.
- Fixed poll wait parsing to keep defaulting to `30s` when the provided value is empty.
- Added support for embedded URL selectors with `:raw` and `:L#-L#` line range syntax in read command.
- Implemented `parseReadUrlTarget()` function to parse and validate URL read targets with line range support.
- Updated read CLI to delegate URL inputs through read tool pipeline instead of treating as local file paths.
- Added comprehensive test coverage for URL selector parsing and CLI URL delegation.
- Refactored URL handling in read tool to use structured `ParsedReadUrlTarget` object.