Commit Graph

35 Commits

Author SHA1 Message Date
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
can1357 cfd72ffb40 fix(coding-agent): fixed eval helpers treating local:// as plain paths
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
2026-06-08 18:28:06 +02:00
can1357 59cb767374 perf(coding-agent/eval): forced eval agent subagents to skip LSP startup
- Changed runEvalAgent to always pass enableLsp: false when launching bridge subagents.
- Added a regression test asserting runSubprocess received enableLsp as false even when LSP is enabled by default.
2026-06-08 12:25:53 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
basedcorp99 1136ba81a1 fix(eval): exit graceful-close JS worker 2026-06-08 11:16:49 +02:00
basedcorp99 a6c8014697 fix(coding-agent): close JS eval workers gracefully 2026-06-08 11:16:49 +02:00
can1357 49aa6e5839 fix(eval): corrected eval LLM calls and spawn-aware tool descriptions
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.
2026-06-08 02:04:59 +02:00
can1357 20487d8cb6 chore: fix stale tests 2026-06-07 08:11:38 +02:00
can1357 0914379f49 fix(coding-agent): rendered shared task context as markdown and froze async block borders
- Updated task call and result rendering to process shared context with the Markdown renderer, so context sections are now displayed with proper Markdown formatting.
- Stopped shimmer animation on pending bash/eval/task blocks once async state is `running`, preventing the committed frame from freezing a transient dark border segment.
- Adjusted rule path display to fall back to a root-relative path when cwd-relative resolution is unavailable.
2026-06-07 08:01:52 +02:00
can1357 76af35dcd6 fix(tui): preserved TUI initial scrollback unless clear was requested
- Updated the initial render intent to carry a `clearScrollback` flag.
- Adjusted first-frame rendering logic to preserve existing scrollback by default and clear only when requested.
- Added `resolver` stubs to coding-agent and web-search test registries for interface compatibility.
2026-06-07 07:54:50 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 5721034739 fix(ui): forced full replay on tool output expand toggle
- Replaced viewport-only repaint with resetDisplay so committed scrollback reflects new heights.
- Added per-server rust-analyzer workspace-ready timing overrides as a test seam.
- Added clearSuppressedSelectors to reset retry-fallback cooldown state.
- Removed obsolete shared eval executors test.
2026-06-06 21:56:00 +02:00
roboomp cab465cba4 fix(eval): surfaced subagent abort reason through python agent() bridge
Python eval agent() collapsed every subagent runtime-limit abort into a
generic 'RuntimeError: bridge call __agent__ failed' instead of the real
reason. runEvalAgent built its failure message with:

  result.error ?? result.stderr ?? result.abortReason ?? <default>

? is nullish-coalescing, so result.stderr = "" (the executor's value for
a runtime-limit abort) short-circuited the chain and never reached
abortReason. The host bridge then shipped {ok: false, error: ""}, and
prelude.py's '<msg> or <fallback>' picked the named-bridge fallback.

Extracted buildSubagentFailureMessage(): aborted subagents prefer the
trimmed abortReason; otherwise fall through error, stderr (trimmed),
abortReason, and the named-bridge default. Empty/whitespace strings no
longer mask anything. The failure-detection condition also accepts
result.aborted so an abort with exitCode 0 (theoretically) still flows
the abort reason out.

Added a regression test asserting that runtime-limit aborts, whitespace
stderr/error, and totally blank aborts all produce non-empty messages
matching the executor's abortReason text.

Fixes #2006
2026-06-06 18:15:04 +00:00
can1357 001a6ad564 tests: remove useless assertations 2026-06-06 16:00:31 +02:00
can1357 ecd80120a3 feat(coding-agent): added framed tool rendering with capped streaming previews
- Wrapped tool call/result renderers with framed block markers for full-width display.
- Added preview-line caps via capPreviewLines to bound multiline outputs with truncation hints.
- Added code-cell tail rendering so streaming output shows capped tail slices with markers.
- Added setPaddingX in Box and applied framed-inline padding adjustments for tight layouts.
2026-06-06 15:19:06 +02:00
can1357 76ca917bfd fix(coding-agent-eval): suspended idle timeout during delegated bridge calls
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
2026-06-06 14:02:04 +02:00
roboomp 867ee74a3b fix(eval): probed Win32 console directly instead of inferring from stdio
The TTY-OR heuristic still mis-classifies the all-stdio-redirected case:
`omp -p "..." < in.txt > out.txt 2> err.log` from a real Windows Terminal
session has `stdin.isTTY === stdout.isTTY === stderr.isTTY === false`,
but the host process can still own a console that the kernel could
inherit. The TTY signals reflect handle redirection, not console
attachment — `GetConsoleWindow()` is the authoritative Win32 signal.

`spawn-options.ts` now:

- Calls `kernel32!GetConsoleWindow()` via `bun:ffi` on Windows. A non-NULL
  HWND means the host has a console regardless of how the standard
  streams are wired, so `windowsHide` stays `false` and the kernel
  inherits the console — which is what fixes the #1960
  numpy/pandas `LoadLibraryExW` hang and lets SIGINT recover via
  `GenerateConsoleCtrlEvent`.
- Falls back to the TTY-OR heuristic when the FFI probe is unavailable
  or off-Windows. That keeps the predicate working on non-Bun-FFI
  runtimes and on POSIX, where `windowsHide` is a no-op anyway.
- Caches the probe result; console attachment is stable for the host's
  lifetime in practice, and dlopening kernel32 on every kernel spawn
  would be wasteful.

The pure helpers (`shouldHideKernelWindow`, `consoleAttachedViaTTY`)
stay separately exported so they remain unit-testable. The integration
boundary `hostHasInheritableConsole()` is what the kernel spawn site
calls, and its return is asserted to be a concrete boolean (kernel spawn
must commit to a `windowsHide` value).
2026-06-06 01:04:13 +00:00
roboomp 8a89ce7de3 fix(eval): widened Python kernel console detection beyond stdout TTY
Stand-alone `process.stdout.isTTY` mis-classifies any partial stdio
redirection — `omp -p "..." > out.txt` reports `stdout.isTTY === false`
even though the parent still owns a console via stdin/stderr. With the
previous check the kernel would have been spawned with `CREATE_NO_WINDOW`
in that scenario and re-hit the #1960 numpy/pandas `LoadLibraryExW` hang
and broken SIGINT.

The host owns a console it can share with the child whenever ANY of
stdin / stdout / stderr is still a TTY; only a fully detached launch
(true service / daemon, or `< in > out 2> err`) has nothing to inherit.

Renamed the predicate parameter to `hostHasInheritableConsole` so the
contract is unambiguous, and added five regression tests covering the
realistic shell redirection combinations.
2026-06-06 00:59:19 +00:00
roboomp 718c8b299b fix(eval): stopped detaching the Python kernel's console on Windows
`PythonKernel.start()` spawned the runner with `windowsHide: true`, which
in Bun maps to the Win32 `CREATE_NO_WINDOW` flag — that detaches the
long-lived child from any inherited console. Two consequences for OMP's
Python eval on Windows:

1. Native extensions that probe the console at init (e.g. NumPy's
   `_core/_multiarray_umath.pyd` plus its bundled OpenBLAS/SLEEF
   thread-pool init) can deadlock inside `LoadLibraryExW`, so the very
   first `import pandas` / `import numpy` after a cold kernel never
   returned. The reporter's faulthandler stack pinned the hang to
   `_bootstrap_external.create_module` → `numpy/_core/multiarray.py:11`.
2. SIGINT cannot be delivered to a console-less process via
   `GenerateConsoleCtrlEvent`, so the host's `proc.kill("SIGINT")`
   silently no-ops — matching the reporter's "kernel unresponsive to
   interrupt" log line. The 5s escalation then hard-kills the kernel.

Plain `python.exe -u -c "import pandas"` from the same venv inherits the
terminal's console and runs to completion, which is the contract this
change restores. The Python kernel now hides its window only when the
host itself has no console to share (service / piped launch); an
interactive TUI launch lets the kernel inherit the parent's console —
analogous to `python.exe` invoked from `cmd.exe`.

The `shouldHideKernelWindow` predicate lives in its own module so it can
be unit-tested without dragging in the kernel's runtime dependencies.
Short-lived helper subprocesses elsewhere (LSP probes, git, plugin
installs) intentionally keep `windowsHide: true` — they don't load
complex native modules and a brief console flash would be user-visible
noise.

Fixes #1960
2026-06-06 00:53:45 +00:00
can1357 5184b565e1 fix(eval): kept Python kernel alive when parallel() cell interrupted
- Resolved in-flight bridge calls the instant the cell's signal aborts.
- Let the kernel unwind via KeyboardInterrupt instead of being hard-killed.
- Preserved persistent session state across wide subagent fan-out teardown.
2026-06-05 20:42:29 +02:00
can1357 c4e157590e feat(eval): replaced fixed concurrency cap with task.maxConcurrency bridge
- Removed the `concurrency` argument from `parallel()` and `pipeline()` in both JS and Python runtimes.
- Added `__concurrency__` bridge to resolve the pool ceiling live from `task.maxConcurrency` (default 32; 0 = unbounded).
- Eval fan-outs now run as wide as a `task` tool batch instead of being capped at 16.
2026-06-02 06:49:15 +02:00
can1357 dbd9489010 refactor(eval): changed timeout from inactivity to wall-clock budget
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
2026-06-01 17:17:00 +02:00
can1357 e18e4ada71 fix(eval): keep idle watchdog armed during in-flight agent()/llm() calls
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.

Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.

- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
2026-06-01 16:38:32 +02:00
can1357 2003d7382e feat(eval): added per-cell inactivity timeout budgets in eval executors
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
2026-05-31 10:11:59 +02:00
can1357 9fabd5e4d5 feat(coding-agent): added onStatus in eval backends for status streams
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
2026-05-31 08:45:12 +02:00
can1357 2ddc9c5bc9 feat(coding-agent): added turn-budget parsing, multipliers and hard caps
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
2026-05-31 08:03:45 +02:00
can1357 7613c2a913 feat(coding-agent): expanded eval execution and bridge paths with workflow helpers
Expand eval execution/bridge paths with args/log/phase/budget support, add workflow helpers (parallel/pipeline), expose usage statistics, and add eval integration tests and docs.
2026-05-31 07:40:18 +02:00
can1357 cf621d0abf feat(coding-agent-eval): added runEvalAgent bridge for agent plan checks
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
2026-05-31 06:57:02 +02:00
can1357 3f2e4b2fe7 fix(eval): link cyclic local module graphs in one pass
The JS eval kernel's LocalModuleLoader linked and evaluated every local
module individually inside the recursive vm.SourceTextModule linker
callback. On any import cycle that re-enters Bun's node:vm linker
mid-instantiation and segfaults JSC (getImportedModule on a null record,
SIGTRAP at 0xFFFFFFFFFFFFFFF8) — e.g. `await import(".../edit/streaming.ts")`,
whose relative-import subtree is cyclic.

Construct the entire local module graph first, then drive a single
link() + evaluate() from the graph root so cyclic graphs instantiate in
one pass. The shared static linker only constructs dependencies;
external (node_modules) modules stay eagerly loaded since they carry no
imports and cannot form a cycle. Failed loads invalidate any
not-fully-evaluated modules so a retry reconstructs them.

Upstream Bun bug: https://github.com/oven-sh/bun/issues/31623
2026-05-31 03:17:47 +02:00
can1357 91513cdbf3 fix(coding-agent/eval): fixed TypeScript type-only import rewriting
- Updated local module loading to force TS syntax stripping for .ts/.tsx/.mts modules and use the matching Bun transpiler loader.
- Extended TypeScript stripping to detect `import type`/`export type` syntax and rewired wrapping to strip TS syntax after final-expression extraction with TypeScript-aware parsing.
- Added tests verifying type-only imports are handled correctly in evaluator modules and rewritten code no longer contains type-only import declarations.
2026-05-30 18:35:01 +02:00
can1357 8715ed207c feat(coding-agent-eval): added oneshot llm helper and __llm__ bridge
- Added one-shot `llm(prompt, opts)` helpers in JS and Python eval runtimes.
- Added `__llm__` eval bridge wiring for synthetic LLM tool dispatch and status/event output.
- Added `runEvalLlm` with tier-to-model resolution, effort handling, and oneshot completion execution.
- Added structured schema output handling via `respond` tool and JSON fallback parsing.
- Documented new llm behavior in eval docs/changelog and added tests for tier mapping and error cases.
2026-05-30 00:27:04 +02:00
can1357 796c437dc1 feat: overhauled stream timeout and eval session management
- Replaced external watchdog timers with per-request SDK timeouts for first-event budget across OpenAI, Anthropic, and Azure providers.
- Keyed Python shared kernels by (sessionId, cwd) to prevent cross-directory state bleed.
- Deduplicated concurrent cold-start session acquisition for JS and Python executors.
- Moved `isOpenAIResponsesProgressEvent` to shared module and scoped display output routing per run for interleaved async cells.
2026-05-26 16:49:11 +02:00
can1357 9a2cc3bda0 feat(eval/py): added runtime environment and working directory support to Python kernel execution
- Added `cwd` and `env` optional parameters to kernel execution API for runtime working directory and environment variable control.
- Implemented runtime environment setup in Python runner with `_apply_request_runtime()` to apply cwd and env from request before code execution.
- Enhanced SIGINT handler management with `active_executions` counter and `_begin_exec_sigint()` / `_end_exec_sigint()` functions to prevent state mutation during concurrent execution.
- Changed `SearchRenderArgs.paths` parameter type from `string[]` to `string | string[]` to accept single string paths.
- Added comprehensive test coverage for kernel cwd updates, timeout interruption safety, and SystemExit handling in shared executor sessions.
2026-05-26 15:26:29 +02:00
can1357 8a5b3e9552 feat(eval): added shared executor inheritance for subagents with concurrent async cells
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
2026-05-26 14:37:56 +02:00