Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.
Addresses review feedback on #2204.
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.
Fixes#2198
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Added lazy async module loaders for @babel/parser, linkedom, puppeteer/browsers, @mozilla/readability, @xterm/headless, and mnemopi to avoid loading them during cold startup.
- Added an interactive startup splash before session construction and skipped it for resume/fork/continue, quiet mode, timing mode, or non-TTY runs.
- Updated JS import-rewrite and memory tests to match async parser loading and preloaded mnemopi modules for sync state helpers.
- Rerouted sync, tab, js-eval, and tiny workers to re-enter CLI modes via `__omp_*` selectors.
- Adjusted `cli.ts` startup to dispatch worker entrypoints before parsing and exit 1 on uncaught errors.
- Bundled CLI as `dist/cli.js` in prepack, switching `omp` binary and published files.
- Removed explicit Bun `--compile` worker entrypoints from build/release scripts in favor of host-entry dispatch.
- Added `declareWorkerHostEntry()` and `workerHostEntry()` environment helpers and `PI_COMPILED` binary detection.
- Introduced memoized dynamic import loaders for Babel parser, mnemopi modules, puppeteer, and HTML-related packages.
- Refactored eval import-rewrite helpers and runtime call sites to use asynchronous wrapping and parsing flows.
- Shifted fetch and web-scraper linkedom usage to on-demand imports so heavy modules load only when needed.
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
- Rewrote dynamic `import(...)` rewriting to emit a guarded callee that prefers `__omp_import__` and falls back to native `import` when the helper is unavailable.
- Added a shared shim constant and updated import-rewrite tests to verify routed dynamic imports work both with the injected helper and after serializing into a realm without it.
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
- Changed runEvalAgent to always pass enableLsp: false when launching bridge subagents.
- Added a regression test asserting runSubprocess received enableLsp as false even when LSP is enabled by default.
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
- Added strict option parsing for JavaScript helpers so read/sort/uniq/counter/tree/llm/agent now reject positional options and non-object option values.
- Introduced `optionsArg` and `isPlainObject` to enforce a single trailing options object and throw clear TypeError messages on invalid calls.
- Updated the eval tool prompt to document that JavaScript helpers require one trailing options object and no extra positional arguments.
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.
- Updated `packages/coding-agent/src/eval/py/prelude.py` to accept positional `offset` and `limit` in `read()`.
- Added regression coverage in `packages/coding-agent/src/eval/py/__tests__/prelude.test.ts` for positional read signatures.
- Added first-party-first provider priority defaults for model ranking.
- Consolidated model resolution to use getModelMatchPreferences from session settings.
- Prioritized providerPriorityRank ahead of usage rank when picking preferred models.
- Added second-pass fallback to default-model or API-key-valid matching order.
- Updated task call and result rendering to process shared context with the Markdown renderer, so context sections are now displayed with proper Markdown formatting.
- Stopped shimmer animation on pending bash/eval/task blocks once async state is `running`, preventing the committed frame from freezing a transient dark border segment.
- Adjusted rule path display to fall back to a root-relative path when cwd-relative resolution is unavailable.
- Updated the initial render intent to carry a `clearScrollback` flag.
- Adjusted first-frame rendering logic to preserve existing scrollback by default and clear only when requested.
- Added `resolver` stubs to coding-agent and web-search test registries for interface compatibility.
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
- Used `||` so empty stderr falls through to abortReason in agent bridge.
- Preferred assistant errorMessage over "Cancelled by caller" on internal aborts.
- Forced `maxRuntimeMs: 0` for eval subagents via ExecutorOptions override.
Python eval agent() collapsed every subagent runtime-limit abort into a
generic 'RuntimeError: bridge call __agent__ failed' instead of the real
reason. runEvalAgent built its failure message with:
result.error ?? result.stderr ?? result.abortReason ?? <default>
? is nullish-coalescing, so result.stderr = "" (the executor's value for
a runtime-limit abort) short-circuited the chain and never reached
abortReason. The host bridge then shipped {ok: false, error: ""}, and
prelude.py's '<msg> or <fallback>' picked the named-bridge fallback.
Extracted buildSubagentFailureMessage(): aborted subagents prefer the
trimmed abortReason; otherwise fall through error, stderr (trimmed),
abortReason, and the named-bridge default. Empty/whitespace strings no
longer mask anything. The failure-detection condition also accepts
result.aborted so an abort with exitCode 0 (theoretically) still flows
the abort reason out.
Added a regression test asserting that runtime-limit aborts, whitespace
stderr/error, and totally blank aborts all produce non-empty messages
matching the executor's abortReason text.
Fixes#2006
- Stopped sharing parentEvalSessionId with bridge-spawned subagents.
- Sharing it deadlocked since the parent's kernel blocks on the bridge call.
- Each subagent now gets its own eval session with an independent kernel.
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
Address review: shared/indirect-eval.ts documents the vm.runInContext mid-execution terminate-race, not a mid-init one. Reword so a reader doesn't grep that file for an init-specific note.
Worker-ready wait reused Bun's 5s default per-test timeout as its floor, so a slow cold-start under --isolate + high CI concurrency was aborted at 5s. The catch then terminate()s a still-initializing Bun worker -- the documented SIGILL/SIGTRAP crash trigger -- crashing the whole test file and intermittently failing unrelated PRs.
Introduce WORKER_INIT_TIMEOUT_MS=15s as a fixed infrastructure floor (independent of, still dominated by, a larger per-cell timeout) and set a 20s file-local setDefaultTimeout in js-executor/js-workflow-helpers tests so cold starts complete instead of being torn down.
The TTY-OR heuristic still mis-classifies the all-stdio-redirected case:
`omp -p "..." < in.txt > out.txt 2> err.log` from a real Windows Terminal
session has `stdin.isTTY === stdout.isTTY === stderr.isTTY === false`,
but the host process can still own a console that the kernel could
inherit. The TTY signals reflect handle redirection, not console
attachment — `GetConsoleWindow()` is the authoritative Win32 signal.
`spawn-options.ts` now:
- Calls `kernel32!GetConsoleWindow()` via `bun:ffi` on Windows. A non-NULL
HWND means the host has a console regardless of how the standard
streams are wired, so `windowsHide` stays `false` and the kernel
inherits the console — which is what fixes the #1960
numpy/pandas `LoadLibraryExW` hang and lets SIGINT recover via
`GenerateConsoleCtrlEvent`.
- Falls back to the TTY-OR heuristic when the FFI probe is unavailable
or off-Windows. That keeps the predicate working on non-Bun-FFI
runtimes and on POSIX, where `windowsHide` is a no-op anyway.
- Caches the probe result; console attachment is stable for the host's
lifetime in practice, and dlopening kernel32 on every kernel spawn
would be wasteful.
The pure helpers (`shouldHideKernelWindow`, `consoleAttachedViaTTY`)
stay separately exported so they remain unit-testable. The integration
boundary `hostHasInheritableConsole()` is what the kernel spawn site
calls, and its return is asserted to be a concrete boolean (kernel spawn
must commit to a `windowsHide` value).
Stand-alone `process.stdout.isTTY` mis-classifies any partial stdio
redirection — `omp -p "..." > out.txt` reports `stdout.isTTY === false`
even though the parent still owns a console via stdin/stderr. With the
previous check the kernel would have been spawned with `CREATE_NO_WINDOW`
in that scenario and re-hit the #1960 numpy/pandas `LoadLibraryExW` hang
and broken SIGINT.
The host owns a console it can share with the child whenever ANY of
stdin / stdout / stderr is still a TTY; only a fully detached launch
(true service / daemon, or `< in > out 2> err`) has nothing to inherit.
Renamed the predicate parameter to `hostHasInheritableConsole` so the
contract is unambiguous, and added five regression tests covering the
realistic shell redirection combinations.
`PythonKernel.start()` spawned the runner with `windowsHide: true`, which
in Bun maps to the Win32 `CREATE_NO_WINDOW` flag — that detaches the
long-lived child from any inherited console. Two consequences for OMP's
Python eval on Windows:
1. Native extensions that probe the console at init (e.g. NumPy's
`_core/_multiarray_umath.pyd` plus its bundled OpenBLAS/SLEEF
thread-pool init) can deadlock inside `LoadLibraryExW`, so the very
first `import pandas` / `import numpy` after a cold kernel never
returned. The reporter's faulthandler stack pinned the hang to
`_bootstrap_external.create_module` → `numpy/_core/multiarray.py:11`.
2. SIGINT cannot be delivered to a console-less process via
`GenerateConsoleCtrlEvent`, so the host's `proc.kill("SIGINT")`
silently no-ops — matching the reporter's "kernel unresponsive to
interrupt" log line. The 5s escalation then hard-kills the kernel.
Plain `python.exe -u -c "import pandas"` from the same venv inherits the
terminal's console and runs to completion, which is the contract this
change restores. The Python kernel now hides its window only when the
host itself has no console to share (service / piped launch); an
interactive TUI launch lets the kernel inherit the parent's console —
analogous to `python.exe` invoked from `cmd.exe`.
The `shouldHideKernelWindow` predicate lives in its own module so it can
be unit-tested without dragging in the kernel's runtime dependencies.
Short-lived helper subprocesses elsewhere (LSP probes, git, plugin
installs) intentionally keep `windowsHide: true` — they don't load
complex native modules and a brief console flash would be user-visible
noise.
Fixes#1960
- Resolved in-flight bridge calls the instant the cell's signal aborts.
- Let the kernel unwind via KeyboardInterrupt instead of being hard-killed.
- Preserved persistent session state across wide subagent fan-out teardown.
- Removed the `concurrency` argument from `parallel()` and `pipeline()` in both JS and Python runtimes.
- Added `__concurrency__` bridge to resolve the pool ceiling live from `task.maxConcurrency` (default 32; 0 = unbounded).
- Eval fan-outs now run as wide as a `task` tool batch instead of being capped at 16.
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.
Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.
- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render