Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.
Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init
All 522 previously-failing Windows tests now pass.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
- Reworked the log_experiment flagging test to create sessions and runs directly through storage APIs.
- Logged a baseline run, completed a second run, and then invoked log.execute using the baseline run ID in flag_runs.
- Verified the baseline run was marked flagged with the expected reason via storage.listLoggedRuns output.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
- Added branch-aware session loading so autoresearch state only rehydrates for current branch.
- Replaced user-specified experiment commands with fixed `bash autoresearch.sh` execution flow.
- Enforced safer setup checks, including missing `autoresearch.sh` and uncommitted-worktree errors.
- Added branch-specific storage helpers, baseline-commit persistence, and expanded tests for dirty-path cases.
- Replaced file-backed autoresearch contracts with sqlite-backed session/run storage in `~/.omp/autoresearch`.
- Added `AutoresearchStorage` and rewired `init_experiment`, `run_experiment`, and `log_experiment` to persist sessions and runs.
- Added `update_notes` tool with `body`/`append_idea` inputs and updated prompts to use active-session context.
- Removed `autoresearch.md` contract parsing and checks flow, including `runChecks`, `force`, and timeout schema options.
- Updated autoresearch state/types to persist `goal`, `notes`, `branch`, and `baselineCommit` plus run justification/flag metadata.
- Added `new_segment` parameter to `init_experiment` to force new segment creation when contract fields match.
- Added `skip_restore` parameter to `log_experiment` to preserve working tree state and pre-existing uncommitted changes.
- Changed `log_experiment` to revert only run-modified files instead of entire working tree, preserving user changes.
- Changed `init_experiment` to detect matching contract fields and skip re-initialization unless `new_segment=true`.
- Added git status parsing utilities to track pre-run dirty paths and distinguish tracked vs untracked modifications.
- Removed secondary metrics validation requirement, making secondary metrics informational only.
- Fixed indentation and formatting across multiple files for consistency.
- Reformatted code blocks to use tabs instead of spaces and improved line breaks for readability.
- Added connection timeout configuration to benchmark runner for early abort on no events.
- Implemented two-phase timeout strategy in benchmark runner with connection and activity phases.
- Fixed PR checkout tool to resolve worktree paths to canonical form using fs.realpath().
- Refactored test mocks to use vi.spyOn for git module functions instead of inline implementations.
- Updated test setup to pass Settings.isolated() with edit.manageImports enabled to EditTool.
- Added afterEach hooks across test suites to restore mocks after each test execution.
- Added contract system for validating benchmark commands, metrics, scope paths, constraints, and off-limits paths.
- Contract validation enforces matching initialization parameters against autoresearch.md before init_experiment.
- Segment fingerprinting detects configuration drift and warns when metrics are not directly comparable.
- Added pending run detection and recovery to resume incomplete experiments from .autoresearch/runs/.
- Run directories organize artifacts with benchmark logs and optional checks logs for traceability.
- Extended experiment state to track run number, command, scope, off-limits, constraints, and fingerprint.