Commit Graph

15 Commits

Author SHA1 Message Date
can1357 060f4004e7 feat(coding-agent): refactored eval tool to single-step execution
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
2026-06-23 00:59:58 +02:00
oldschoola 14252e71cb fix: Windows test failures — path handling, EBUSY, SQLite handle leaks
Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.

Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
   with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
   and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
   that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
   from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
   now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
   added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
    cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init

All 522 previously-failing Windows tests now pass.
2026-06-18 21:32:38 -07:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 38977b8450 test: replaced exact sleep assertions with tolerance-aware helper
- Added `expectSleepNear` to allow ±100ms variance on sleep duration checks.
- Updated `--thinking` test value from `"extended"` to `Effort.XHigh` enum constant.
2026-06-02 08:57:29 +02:00
can1357 796c437dc1 feat: overhauled stream timeout and eval session management
- Replaced external watchdog timers with per-request SDK timeouts for first-event budget across OpenAI, Anthropic, and Azure providers.
- Keyed Python shared kernels by (sessionId, cwd) to prevent cross-directory state bleed.
- Deduplicated concurrent cold-start session acquisition for JS and Python executors.
- Moved `isOpenAIResponsesProgressEvent` to shared module and scoped display output routing per run for interleaved async cells.
2026-05-26 16:49:11 +02:00
can1357 90b134ca4c test: replaced real timers and sleeps with deterministic test hooks
- Added `providerRetryWait` and `retryWait` hooks to stream/usage options so tests bypass real scheduler delays.
- Parameterized GitHub Copilot poll intervals and Copilot model retry base delay for fast test execution.
- Replaced `Bun.sleep`/`setTimeout` polling loops with `AbortSignal` event listeners in agent session tests.
- Consolidated auth-gateway E2E helpers into a shared `test/helpers` module, eliminating duplicated `checkGatewayAvailable` implementations.
- Migrated credential-disabled tests from SQLite-backed stores to an in-memory store, removing temp-dir lifecycle overhead.
2026-05-17 04:02:09 +02:00
can1357 5523582735 test(coding-agent): minor fixes 2026-05-17 03:12:36 +02:00
can1357 6e8daae059 feat(coding-agent): added inline hashline parse with conflict checks
- Added inline hashline parse and apply support for `<` prepend and `+` append operations with prefix+suffix edits.
- Added fail-fast behavior to reject inline modify ops combined with delete or replace on same line.
- Renamed HASHLINE_* and mode symbols to HL_* in prompt tooling, read/search checks, and prompt templates.
- Standardized separators to `PI_HL_SEP`/`HL_EDIT_SEP` and fixed `HL_BODY_SEP='|'`, updating parser formatting behavior.
- Updated benchmark subtype constants and python cleanup test setup to use HL_* values and AgentRegistry mock failure injection.
2026-05-03 08:17:09 +02:00
can1357 6dccdf9193 refactor(coding-agent/eval): drop python warmup path
The warmup path no longer produces prelude docs, so the cached-session
warmup it implemented added no value over the create-on-first-execute
path that withKernelSession already covers. Remove warmPythonEnvironment,
the backend warm() hook, the eval-tool warmup loop, the createTools
warmup preflight, and the forcePythonWarmup option. Simplify
ExecutorBackendCallOptions into ExecutorBackendExecOptions since execute
is now the only consumer.
2026-05-01 15:27:40 +02:00
can1357 76fe4b4995 test(coding-agent/eval): drop tests for removed helpers and stale warmup count 2026-04-30 18:44:07 +02:00
can1357 b0a31a5956 fix: outdated tests 2026-04-30 18:40:26 +02:00
can1357 cf60e6df51 feat(coding-agent): implemented eval framework and replaced python tool
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
2026-04-30 18:08:37 +02:00
Carl 2b480f7da1 fix(coding-agent): addressed python cleanup review findings
fixed retained-kernel restart and owner cleanup edge cases during recovery and disposal

tracked async user_python hooks during disposal-sensitive execution paths and hardened startup warmup tracking

strengthened cleanup and kernel lifecycle regressions to remove deadlocks, false positives, and timing flakes
2026-04-12 14:45:19 +02:00
Carl 379064a0c9 fix(coding-agent): scoped python kernel cleanup to sessions
scoped retained-kernel ownership to agent sessions and cleaned it up on session disposal

cleaned up warmed python owners on session startup failure and rejected new direct and tool-based python starts during disposal, including async hook, preflight, and warmup races
2026-04-12 13:29:17 +02:00