Commit Graph

118 Commits

Author SHA1 Message Date
can1357 4068bfc304 fix(coding-agent): key python kernel sessions by resolved interpreter 2026-06-10 09:51:45 +02:00
can1357 7ce58fe95c Merge pull request #2204: feat(eval): add python interpreter setting 2026-06-10 08:26:59 +02:00
can1357 24cbc93913 fix(eval): thread python.interpreter from session settings and expand ~
Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.

Addresses review feedback on #2204.
2026-06-10 08:26:00 +02:00
roboomp 8c3149e5a9 fix(ai): scoped antigravity quota blocks by model family
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.

Fixes #2198
2026-06-10 08:26:00 +02:00
danzaio caff395d07 feat(eval): add python interpreter setting 2026-06-10 08:26:00 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 8583421362 perf(coding-agent): deferred heavy imports until first feature use
- Added lazy async module loaders for @babel/parser, linkedom, puppeteer/browsers, @mozilla/readability, @xterm/headless, and mnemopi to avoid loading them during cold startup.
- Added an interactive startup splash before session construction and skipped it for resume/fork/continue, quiet mode, timing mode, or non-TTY runs.
- Updated JS import-rewrite and memory tests to match async parser loading and preloaded mnemopi modules for sync state helpers.
2026-06-10 03:57:32 +02:00
can1357 dc5c93462f feat: rerouted worker subprocesses through the bundled CLI host entrypoint
- Rerouted sync, tab, js-eval, and tiny workers to re-enter CLI modes via `__omp_*` selectors.
- Adjusted `cli.ts` startup to dispatch worker entrypoints before parsing and exit 1 on uncaught errors.
- Bundled CLI as `dist/cli.js` in prepack, switching `omp` binary and published files.
- Removed explicit Bun `--compile` worker entrypoints from build/release scripts in favor of host-entry dispatch.
- Added `declareWorkerHostEntry()` and `workerHostEntry()` environment helpers and `PI_COMPILED` binary detection.
2026-06-10 03:57:31 +02:00
can1357 f638a5b3e0 refactor(coding-agent): added lazy loading for heavy dependencies
- Introduced memoized dynamic import loaders for Babel parser, mnemopi modules, puppeteer, and HTML-related packages.
- Refactored eval import-rewrite helpers and runtime call sites to use asynchronous wrapping and parsing flows.
- Shifted fetch and web-scraper linkedom usage to on-demand imports so heavy modules load only when needed.
2026-06-10 03:09:09 +02:00
can1357 d7eae06830 fix(coding-agent): fixed eval artifact double-writes and kernel I/O capture
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
2026-06-10 01:27:41 +02:00
can1357 3e990c5bea fix(coding-agent): added guarded dynamic import shim for worker fallback
- Rewrote dynamic `import(...)` rewriting to emit a guarded callee that prefers `__omp_import__` and falls back to native `import` when the helper is unavailable.
- Added a shared shim constant and updated import-rewrite tests to verify routed dynamic imports work both with the injected helper and after serializing into a realm without it.
2026-06-09 22:57:19 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
can1357 cfd72ffb40 fix(coding-agent): fixed eval helpers treating local:// as plain paths
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
2026-06-08 18:28:06 +02:00
can1357 59cb767374 perf(coding-agent/eval): forced eval agent subagents to skip LSP startup
- Changed runEvalAgent to always pass enableLsp: false when launching bridge subagents.
- Added a regression test asserting runSubprocess received enableLsp as false even when LSP is enabled by default.
2026-06-08 12:25:53 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
basedcorp99 1e2bbe71aa fix(eval): use withResolvers for worker close 2026-06-08 11:49:23 +02:00
basedcorp99 1136ba81a1 fix(eval): exit graceful-close JS worker 2026-06-08 11:16:49 +02:00
basedcorp99 a6c8014697 fix(coding-agent): close JS eval workers gracefully 2026-06-08 11:16:49 +02:00
can1357 65bddab719 feat(coding-agent/eval): enforced JS eval helper options as trailing object literals
- Added strict option parsing for JavaScript helpers so read/sort/uniq/counter/tree/llm/agent now reject positional options and non-object option values.
- Introduced `optionsArg` and `isPlainObject` to enforce a single trailing options object and throw clear TypeError messages on invalid calls.
- Updated the eval tool prompt to document that JavaScript helpers require one trailing options object and no extra positional arguments.
2026-06-08 03:33:45 +02:00
can1357 088fb7fb75 fix(eval): resolved JS/Python resets by awaiting in-flight operations
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
2026-06-08 02:04:59 +02:00
can1357 49aa6e5839 fix(eval): corrected eval LLM calls and spawn-aware tool descriptions
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.
2026-06-08 02:04:59 +02:00
can1357 f73942eb28 fix(eval/py): corrected read() positional offset/limit support and tests
- Updated `packages/coding-agent/src/eval/py/prelude.py` to accept positional `offset` and `limit` in `read()`.
- Added regression coverage in `packages/coding-agent/src/eval/py/__tests__/prelude.test.ts` for positional read signatures.
2026-06-08 02:04:59 +02:00
can1357 caa9c69309 feat(coding-agent): added provider-priority model selection and refined fallback ordering
- Added first-party-first provider priority defaults for model ranking.
- Consolidated model resolution to use getModelMatchPreferences from session settings.
- Prioritized providerPriorityRank ahead of usage rank when picking preferred models.
- Added second-pass fallback to default-model or API-key-valid matching order.
2026-06-08 01:32:39 +02:00
Can Bölük e9e2d02223 Merge pull request #1982 from AsafMah/fix/js-eval-worker-init-flake
fix(eval): floor JS worker init timeout to stop terminate-mid-init CI flake
2026-06-08 00:19:33 +02:00
can1357 20487d8cb6 chore: fix stale tests 2026-06-07 08:11:38 +02:00
can1357 0914379f49 fix(coding-agent): rendered shared task context as markdown and froze async block borders
- Updated task call and result rendering to process shared context with the Markdown renderer, so context sections are now displayed with proper Markdown formatting.
- Stopped shimmer animation on pending bash/eval/task blocks once async state is `running`, preventing the committed frame from freezing a transient dark border segment.
- Adjusted rule path display to fall back to a root-relative path when cwd-relative resolution is unavailable.
2026-06-07 08:01:52 +02:00
can1357 76af35dcd6 fix(tui): preserved TUI initial scrollback unless clear was requested
- Updated the initial render intent to carry a `clearScrollback` flag.
- Adjusted first-frame rendering logic to preserve existing scrollback by default and clear only when requested.
- Added `resolver` stubs to coding-agent and web-search test registries for interface compatibility.
2026-06-07 07:54:50 +02:00
can1357 c10eb5e50e feat(coding-agent): added resolver-based auth retries to image and search tools
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
2026-06-07 04:49:19 +02:00
can1357 cba6641299 feat: enabled resolver-based API key retries with refresh and rotation
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
2026-06-07 04:20:48 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 5721034739 fix(ui): forced full replay on tool output expand toggle
- Replaced viewport-only repaint with resetDisplay so committed scrollback reflects new heights.
- Added per-server rust-analyzer workspace-ready timing overrides as a test seam.
- Added clearSuppressedSelectors to reset retry-fallback cooldown state.
- Removed obsolete shared eval executors test.
2026-06-06 21:56:00 +02:00
can1357 3b5b182553 Merge remote-tracking branch 'origin/farm/10b5f6a2/surface-subagent-abort-reason' 2026-06-06 21:34:33 +02:00
can1357 133137c9a6 fix(eval): surfaced subagent abort reasons and disabled runtime cap
- Used `||` so empty stderr falls through to abortReason in agent bridge.
- Preferred assistant errorMessage over "Cancelled by caller" on internal aborts.
- Forced `maxRuntimeMs: 0` for eval subagents via ExecutorOptions override.
2026-06-06 21:33:07 +02:00
can1357 bdbbfa9778 fix(eval): surfaced subagent abort reason
- Used `||` so empty stderr no longer masks the real abort reason.
2026-06-06 20:54:00 +02:00
roboomp cab465cba4 fix(eval): surfaced subagent abort reason through python agent() bridge
Python eval agent() collapsed every subagent runtime-limit abort into a
generic 'RuntimeError: bridge call __agent__ failed' instead of the real
reason. runEvalAgent built its failure message with:

  result.error ?? result.stderr ?? result.abortReason ?? <default>

? is nullish-coalescing, so result.stderr = "" (the executor's value for
a runtime-limit abort) short-circuited the chain and never reached
abortReason. The host bridge then shipped {ok: false, error: ""}, and
prelude.py's '<msg> or <fallback>' picked the named-bridge fallback.

Extracted buildSubagentFailureMessage(): aborted subagents prefer the
trimmed abortReason; otherwise fall through error, stderr (trimmed),
abortReason, and the named-bridge default. Empty/whitespace strings no
longer mask anything. The failure-detection condition also accepts
result.aborted so an abort with exitCode 0 (theoretically) still flows
the abort reason out.

Added a regression test asserting that runtime-limit aborts, whitespace
stderr/error, and totally blank aborts all produce non-empty messages
matching the executor's abortReason text.

Fixes #2006
2026-06-06 18:15:04 +00:00
can1357 c057b0ef33 fix(coding-agent/eval): prevented subagent eval session deadlock
- Stopped sharing parentEvalSessionId with bridge-spawned subagents.
- Sharing it deadlocked since the parent's kernel blocks on the bridge call.
- Each subagent now gets its own eval session with an independent kernel.
2026-06-06 16:55:05 +02:00
can1357 001a6ad564 tests: remove useless assertations 2026-06-06 16:00:31 +02:00
can1357 ecd80120a3 feat(coding-agent): added framed tool rendering with capped streaming previews
- Wrapped tool call/result renderers with framed block markers for full-width display.
- Added preview-line caps via capPreviewLines to bound multiline outputs with truncation hints.
- Added code-cell tail rendering so streaming output shows capped tail slices with markers.
- Added setPaddingX in Box and applied framed-inline padding adjustments for tight layouts.
2026-06-06 15:19:06 +02:00
can1357 76ca917bfd fix(coding-agent-eval): suspended idle timeout during delegated bridge calls
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
2026-06-06 14:02:04 +02:00
Asaf Mahlev ea8d58d3c6 docs(eval): clarify indirect-eval cross-reference in worker-init comment
Address review: shared/indirect-eval.ts documents the vm.runInContext mid-execution terminate-race, not a mid-init one. Reword so a reader doesn't grep that file for an init-specific note.
2026-06-06 12:43:53 +03:00
Asaf Mahlev 236bba96b9 fix(eval): floor JS worker init timeout to stop terminate-mid-init flake
Worker-ready wait reused Bun's 5s default per-test timeout as its floor, so a slow cold-start under --isolate + high CI concurrency was aborted at 5s. The catch then terminate()s a still-initializing Bun worker -- the documented SIGILL/SIGTRAP crash trigger -- crashing the whole test file and intermittently failing unrelated PRs.

Introduce WORKER_INIT_TIMEOUT_MS=15s as a fixed infrastructure floor (independent of, still dominated by, a larger per-cell timeout) and set a 20s file-local setDefaultTimeout in js-executor/js-workflow-helpers tests so cold starts complete instead of being torn down.
2026-06-06 11:56:23 +03:00
roboomp 867ee74a3b fix(eval): probed Win32 console directly instead of inferring from stdio
The TTY-OR heuristic still mis-classifies the all-stdio-redirected case:
`omp -p "..." < in.txt > out.txt 2> err.log` from a real Windows Terminal
session has `stdin.isTTY === stdout.isTTY === stderr.isTTY === false`,
but the host process can still own a console that the kernel could
inherit. The TTY signals reflect handle redirection, not console
attachment — `GetConsoleWindow()` is the authoritative Win32 signal.

`spawn-options.ts` now:

- Calls `kernel32!GetConsoleWindow()` via `bun:ffi` on Windows. A non-NULL
  HWND means the host has a console regardless of how the standard
  streams are wired, so `windowsHide` stays `false` and the kernel
  inherits the console — which is what fixes the #1960
  numpy/pandas `LoadLibraryExW` hang and lets SIGINT recover via
  `GenerateConsoleCtrlEvent`.
- Falls back to the TTY-OR heuristic when the FFI probe is unavailable
  or off-Windows. That keeps the predicate working on non-Bun-FFI
  runtimes and on POSIX, where `windowsHide` is a no-op anyway.
- Caches the probe result; console attachment is stable for the host's
  lifetime in practice, and dlopening kernel32 on every kernel spawn
  would be wasteful.

The pure helpers (`shouldHideKernelWindow`, `consoleAttachedViaTTY`)
stay separately exported so they remain unit-testable. The integration
boundary `hostHasInheritableConsole()` is what the kernel spawn site
calls, and its return is asserted to be a concrete boolean (kernel spawn
must commit to a `windowsHide` value).
2026-06-06 01:04:13 +00:00
roboomp 69ffc80368 style: bun run fix 2026-06-06 00:59:23 +00:00
roboomp 8a89ce7de3 fix(eval): widened Python kernel console detection beyond stdout TTY
Stand-alone `process.stdout.isTTY` mis-classifies any partial stdio
redirection — `omp -p "..." > out.txt` reports `stdout.isTTY === false`
even though the parent still owns a console via stdin/stderr. With the
previous check the kernel would have been spawned with `CREATE_NO_WINDOW`
in that scenario and re-hit the #1960 numpy/pandas `LoadLibraryExW` hang
and broken SIGINT.

The host owns a console it can share with the child whenever ANY of
stdin / stdout / stderr is still a TTY; only a fully detached launch
(true service / daemon, or `< in > out 2> err`) has nothing to inherit.

Renamed the predicate parameter to `hostHasInheritableConsole` so the
contract is unambiguous, and added five regression tests covering the
realistic shell redirection combinations.
2026-06-06 00:59:19 +00:00
roboomp 718c8b299b fix(eval): stopped detaching the Python kernel's console on Windows
`PythonKernel.start()` spawned the runner with `windowsHide: true`, which
in Bun maps to the Win32 `CREATE_NO_WINDOW` flag — that detaches the
long-lived child from any inherited console. Two consequences for OMP's
Python eval on Windows:

1. Native extensions that probe the console at init (e.g. NumPy's
   `_core/_multiarray_umath.pyd` plus its bundled OpenBLAS/SLEEF
   thread-pool init) can deadlock inside `LoadLibraryExW`, so the very
   first `import pandas` / `import numpy` after a cold kernel never
   returned. The reporter's faulthandler stack pinned the hang to
   `_bootstrap_external.create_module` → `numpy/_core/multiarray.py:11`.
2. SIGINT cannot be delivered to a console-less process via
   `GenerateConsoleCtrlEvent`, so the host's `proc.kill("SIGINT")`
   silently no-ops — matching the reporter's "kernel unresponsive to
   interrupt" log line. The 5s escalation then hard-kills the kernel.

Plain `python.exe -u -c "import pandas"` from the same venv inherits the
terminal's console and runs to completion, which is the contract this
change restores. The Python kernel now hides its window only when the
host itself has no console to share (service / piped launch); an
interactive TUI launch lets the kernel inherit the parent's console —
analogous to `python.exe` invoked from `cmd.exe`.

The `shouldHideKernelWindow` predicate lives in its own module so it can
be unit-tested without dragging in the kernel's runtime dependencies.
Short-lived helper subprocesses elsewhere (LSP probes, git, plugin
installs) intentionally keep `windowsHide: true` — they don't load
complex native modules and a brief console flash would be user-visible
noise.

Fixes #1960
2026-06-06 00:53:45 +00:00
can1357 5184b565e1 fix(eval): kept Python kernel alive when parallel() cell interrupted
- Resolved in-flight bridge calls the instant the cell's signal aborts.
- Let the kernel unwind via KeyboardInterrupt instead of being hard-killed.
- Preserved persistent session state across wide subagent fan-out teardown.
2026-06-05 20:42:29 +02:00
can1357 c4e157590e feat(eval): replaced fixed concurrency cap with task.maxConcurrency bridge
- Removed the `concurrency` argument from `parallel()` and `pipeline()` in both JS and Python runtimes.
- Added `__concurrency__` bridge to resolve the pool ceiling live from `task.maxConcurrency` (default 32; 0 = unbounded).
- Eval fan-outs now run as wide as a `task` tool batch instead of being capped at 16.
2026-06-02 06:49:15 +02:00
can1357 dbd9489010 refactor(eval): changed timeout from inactivity to wall-clock budget
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
2026-06-01 17:17:00 +02:00
can1357 e18e4ada71 fix(eval): keep idle watchdog armed during in-flight agent()/llm() calls
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.

Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.

- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
2026-06-01 16:38:32 +02:00