Commit Graph
15 Commits
Author SHA1 Message Date
can1357 c03a78917a test(coding-agent): disabled fsmonitor during test repo initialization
- Added core.fsmonitor configuration to the git initialization process in test setup to prevent potential environment-specific conflicts.
2026-06-25 12:55:01 +02:00
oldschoola 14252e71cb fix: Windows test failures — path handling, EBUSY, SQLite handle leaks
Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.

Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
   with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
   and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
   that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
   from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
   now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
   added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
    cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init

All 522 previously-failing Windows tests now pass.
2026-06-18 21:32:38 -07:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
can1357 9d457f73d9 test: migrated test imports to package subpath exports
- Replaced relative `../src` imports with `@oh-my-pi/pi-ai` and `@oh-my-pi/pi-agent-core` subpaths.
2026-06-08 19:03:55 +02:00
can1357 ea14cee2bc test(coding-agent): refined log_experiment flagging test to use storage-run setup
- Reworked the log_experiment flagging test to create sessions and runs directly through storage APIs.
- Logged a baseline run, completed a second run, and then invoked log.execute using the baseline run ID in flag_runs.
- Verified the baseline run was marked flagged with the expected reason via storage.listLoggedRuns output.
2026-06-08 03:48:42 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 31ae1b96fd feat(coding-agent/autoresearch): added branch-specific state restore
- Added branch-aware session loading so autoresearch state only rehydrates for current branch.
- Replaced user-specified experiment commands with fixed `bash autoresearch.sh` execution flow.
- Enforced safer setup checks, including missing `autoresearch.sh` and uncommitted-worktree errors.
- Added branch-specific storage helpers, baseline-commit persistence, and expanded tests for dirty-path cases.
2026-05-02 03:31:27 +02:00
can1357 3293231a9f feat(coding-agent): implemented sqlite-backed storage for autoresearch
- Replaced file-backed autoresearch contracts with sqlite-backed session/run storage in `~/.omp/autoresearch`.
- Added `AutoresearchStorage` and rewired `init_experiment`, `run_experiment`, and `log_experiment` to persist sessions and runs.
- Added `update_notes` tool with `body`/`append_idea` inputs and updated prompts to use active-session context.
- Removed `autoresearch.md` contract parsing and checks flow, including `runChecks`, `force`, and timeout schema options.
- Updated autoresearch state/types to persist `goal`, `notes`, `branch`, and `baselineCommit` plus run justification/flag metadata.
2026-05-02 01:10:00 +02:00
can1357 87a3b6ca03 refactor(patch): restructured chunk edit schema to use explicit op enum and anchor field
- Refactored chunk edit schema to use explicit `op` enum field (replace, delete, append, prepend, after, before) instead of multiple boolean/string flags.
- Renamed chunk edit parameters from `after`/`before` field names to `anchor` for sibling-relative insert operations.
- Simplified chunk edit validation logic by consolidating nested conditionals into switch statement on operation type.
- Fixed `log_experiment` to correctly identify run-modified files by removing pre-run dirty path filtering.
- Simplified chunk edit and read prompt documentation to emphasize read-first workflow with consolidated examples.
2026-04-07 05:53:40 +02:00
can1357 266aabf29f feat(autoresearch): enabled selective segment creation and working tree preservation in experiments
- Added `new_segment` parameter to `init_experiment` to force new segment creation when contract fields match.
- Added `skip_restore` parameter to `log_experiment` to preserve working tree state and pre-existing uncommitted changes.
- Changed `log_experiment` to revert only run-modified files instead of entire working tree, preserving user changes.
- Changed `init_experiment` to detect matching contract fields and skip re-initialization unless `new_segment=true`.
- Added git status parsing utilities to track pre-run dirty paths and distinguish tracked vs untracked modifications.
- Removed secondary metrics validation requirement, making secondary metrics informational only.
2026-04-06 18:49:52 +02:00
can1357 51467abc5a style(coding-agent): standardized indentation and added connection timeout to benchmark runner
- Fixed indentation and formatting across multiple files for consistency.
- Reformatted code blocks to use tabs instead of spaces and improved line breaks for readability.
- Added connection timeout configuration to benchmark runner for early abort on no events.
- Implemented two-phase timeout strategy in benchmark runner with connection and activity phases.
2026-04-06 18:45:41 +02:00
can1357 3d292b0b49 refactor(autoresearch): overhaul tools, state management, and contract handling 2026-04-06 18:41:04 +02:00
can1357 85d0d366c5 fix(coding-agent): resolved PR checkout worktree paths to canonical form
- Fixed PR checkout tool to resolve worktree paths to canonical form using fs.realpath().
- Refactored test mocks to use vi.spyOn for git module functions instead of inline implementations.
- Updated test setup to pass Settings.isolated() with edit.manageImports enabled to EditTool.
- Added afterEach hooks across test suites to restore mocks after each test execution.
2026-04-04 16:08:48 +02:00
can1357 2c93655796 feat(autoresearch): added auto-resume, path validation, and security guards
- Added auto-resume mechanism with state tracking to automatically resume pending experiment runs and prevent duplicate resumptions.
- Added contract path validation to reject unsafe path specifications with absolute paths and parent directory traversal attempts.
- Added secondary metrics input to autoresearch setup flow for specifying tradeoff metrics alongside primary objectives.
- Enhanced command parsing with shell operator detection to reject piped, redirected, or chained autoresearch.sh commands.
- Added prototype pollution guards in object cloning functions to prevent injection via __proto__, constructor, and prototype keys.
- Fixed boundary duplication warnings in hashline detection to properly report multiple overlapping hashline references.
2026-03-23 02:16:39 +01:00
can1357 003f46f42c feat(coding-agent/autoresearch): added contract validation and run tracking system
- Added contract system for validating benchmark commands, metrics, scope paths, constraints, and off-limits paths.
 - Contract validation enforces matching initialization parameters against autoresearch.md before init_experiment.
 - Segment fingerprinting detects configuration drift and warns when metrics are not directly comparable.
 - Added pending run detection and recovery to resume incomplete experiments from .autoresearch/runs/.
 - Run directories organize artifacts with benchmark logs and optional checks logs for traceability.
 - Extended experiment state to track run number, command, scope, off-limits, constraints, and fingerprint.
2026-03-23 01:49:03 +01:00