Commit Graph
55 Commits
Author SHA1 Message Date
can1357 088fb7fb75 fix(eval): resolved JS/Python resets by awaiting in-flight operations
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
2026-06-08 02:04:59 +02:00
Can BölükandGitHub e9e2d02223 Merge pull request #1982 from AsafMah/fix/js-eval-worker-init-flake
fix(eval): floor JS worker init timeout to stop terminate-mid-init CI flake
2026-06-08 00:19:33 +02:00
can1357 76ca917bfd fix(coding-agent-eval): suspended idle timeout during delegated bridge calls
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
2026-06-06 14:02:04 +02:00
Asaf Mahlev ea8d58d3c6 docs(eval): clarify indirect-eval cross-reference in worker-init comment
Address review: shared/indirect-eval.ts documents the vm.runInContext mid-execution terminate-race, not a mid-init one. Reword so a reader doesn't grep that file for an init-specific note.
2026-06-06 12:43:53 +03:00
Asaf Mahlev 236bba96b9 fix(eval): floor JS worker init timeout to stop terminate-mid-init flake
Worker-ready wait reused Bun's 5s default per-test timeout as its floor, so a slow cold-start under --isolate + high CI concurrency was aborted at 5s. The catch then terminate()s a still-initializing Bun worker -- the documented SIGILL/SIGTRAP crash trigger -- crashing the whole test file and intermittently failing unrelated PRs.

Introduce WORKER_INIT_TIMEOUT_MS=15s as a fixed infrastructure floor (independent of, still dominated by, a larger per-cell timeout) and set a 20s file-local setDefaultTimeout in js-executor/js-workflow-helpers tests so cold starts complete instead of being torn down.
2026-06-06 11:56:23 +03:00
can1357 c4e157590e feat(eval): replaced fixed concurrency cap with task.maxConcurrency bridge
- Removed the `concurrency` argument from `parallel()` and `pipeline()` in both JS and Python runtimes.
- Added `__concurrency__` bridge to resolve the pool ceiling live from `task.maxConcurrency` (default 32; 0 = unbounded).
- Eval fan-outs now run as wide as a `task` tool batch instead of being capped at 16.
2026-06-02 06:49:15 +02:00
can1357 dbd9489010 refactor(eval): changed timeout from inactivity to wall-clock budget
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
2026-06-01 17:17:00 +02:00
can1357 e18e4ada71 fix(eval): keep idle watchdog armed during in-flight agent()/llm() calls
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.

Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.

- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
2026-06-01 16:38:32 +02:00
can1357 2003d7382e feat(eval): added per-cell inactivity timeout budgets in eval executors
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
2026-05-31 10:11:59 +02:00
can1357 9fabd5e4d5 feat(coding-agent): added onStatus in eval backends for status streams
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
2026-05-31 08:45:12 +02:00
can1357 2ddc9c5bc9 feat(coding-agent): added turn-budget parsing, multipliers and hard caps
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
2026-05-31 08:03:45 +02:00
can1357 4ab40764d1 refactor(coding-agent): removed args global from eval runtimes
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
2026-05-31 07:54:52 +02:00
can1357 7613c2a913 feat(coding-agent): expanded eval execution and bridge paths with workflow helpers
Expand eval execution/bridge paths with args/log/phase/budget support, add workflow helpers (parallel/pipeline), expose usage statistics, and add eval integration tests and docs.
2026-05-31 07:40:18 +02:00
can1357 cf621d0abf feat(coding-agent-eval): added runEvalAgent bridge for agent plan checks
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
2026-05-31 06:57:02 +02:00
can1357 3f2e4b2fe7 fix(eval): link cyclic local module graphs in one pass
The JS eval kernel's LocalModuleLoader linked and evaluated every local
module individually inside the recursive vm.SourceTextModule linker
callback. On any import cycle that re-enters Bun's node:vm linker
mid-instantiation and segfaults JSC (getImportedModule on a null record,
SIGTRAP at 0xFFFFFFFFFFFFFFF8) — e.g. `await import(".../edit/streaming.ts")`,
whose relative-import subtree is cyclic.

Construct the entire local module graph first, then drive a single
link() + evaluate() from the graph root so cyclic graphs instantiate in
one pass. The shared static linker only constructs dependencies;
external (node_modules) modules stay eagerly loaded since they carry no
imports and cannot form a cycle. Failed loads invalidate any
not-fully-evaluated modules so a retry reconstructs them.

Upstream Bun bug: https://github.com/oven-sh/bun/issues/31623
2026-05-31 03:17:47 +02:00
can1357 91513cdbf3 fix(coding-agent/eval): fixed TypeScript type-only import rewriting
- Updated local module loading to force TS syntax stripping for .ts/.tsx/.mts modules and use the matching Bun transpiler loader.
- Extended TypeScript stripping to detect `import type`/`export type` syntax and rewired wrapping to strip TS syntax after final-expression extraction with TypeScript-aware parsing.
- Added tests verifying type-only imports are handled correctly in evaluator modules and rewritten code no longer contains type-only import declarations.
2026-05-30 18:35:01 +02:00
can1357 e831c2c758 chore: reformat 2026-05-30 18:08:51 +02:00
can1357 8715ed207c feat(coding-agent-eval): added oneshot llm helper and __llm__ bridge
- Added one-shot `llm(prompt, opts)` helpers in JS and Python eval runtimes.
- Added `__llm__` eval bridge wiring for synthetic LLM tool dispatch and status/event output.
- Added `runEvalLlm` with tier-to-model resolution, effort handling, and oneshot completion execution.
- Added structured schema output handling via `respond` tool and JSON fallback parsing.
- Documented new llm behavior in eval docs/changelog and added tests for tier mapping and error cases.
2026-05-30 00:27:04 +02:00
can1357 796c437dc1 feat: overhauled stream timeout and eval session management
- Replaced external watchdog timers with per-request SDK timeouts for first-event budget across OpenAI, Anthropic, and Azure providers.
- Keyed Python shared kernels by (sessionId, cwd) to prevent cross-directory state bleed.
- Deduplicated concurrent cold-start session acquisition for JS and Python executors.
- Moved `isOpenAIResponsesProgressEvent` to shared module and scoped display output routing per run for interleaved async cells.
2026-05-26 16:49:11 +02:00
can1357 8a5b3e9552 feat(eval): added shared executor inheritance for subagents with concurrent async cells
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
2026-05-26 14:37:56 +02:00
can1357 7f0208ac83 feat(coding-agent/eval): added console.table bridge to runtime text output
- Added a `console.table` helper in the JS prelude that forwards calls to the runtime `__omp_table__` hook.
- Implemented `__omp_table__` in the runtime to render tables through `node:console.Console` and emit text via `onText`.
- Added tests verifying `console.table` produced formatted table output and respected the optional columns filter.
2026-05-26 13:16:59 +02:00
can1357 9456e4fe2c fix(coding-agent): coalesce user timeout with JS worker startup window
READY_TIMEOUT_MS was a hard 5s ceiling, ignoring the per-cell timeout the
caller supplied. Cold Bun starts on slow machines routinely exceed 5s
and the worker init failed before user code ran. The window now takes
the larger of the default and the caller timeout.
2026-05-19 19:16:08 +09:00
can1357 7aef1f1cd5 fix(coding-agent): rewrite trailing return statement into final expression
A top-level `return value;` in a JS eval cell was previously swallowed:
returnFinalExpression only handled ExpressionStatement, so ReturnStatement
flipped the IIFE wrapper which discarded the value. Rewrite the trailing
return into __omp_set_final_expr__((expr)) so the existing
final-expression channel surfaces the value just like a trailing
expression.
2026-05-19 19:16:08 +09:00
can1357 226fe87345 fix(coding-agent/eval): hardened image display value coercion to strict base64
- Implemented strict base64 validation and normalization for image `displayValue` payloads in the JS runtime, supporting strict strings, `Uint8Array`, `Buffer`, `ArrayBuffer`, typed-array views, and JSON `Buffer` objects.
- Dropped unrecognized image payloads while emitting a warning and fallback text instead of forwarding malformed data.
- Added tests covering successful coercions and invalid image data rejection paths.
2026-05-19 04:36:30 +02:00
can1357 7901cecf80 feat(coding-agent): added module cache busting for local imports in JS runtime
- Appended a unique nonce query param to local file imports so Bun treats each reload as a fresh module record.
- Restricted cache busting to relative/absolute path specifiers; bare packages and built-ins are left unchanged.
2026-05-16 20:27:45 +02:00
can1357 110a6a3244 fix(coding-agent): persisted bindings from async-wrapped cells via globalThis publish
- Detected async wrapper need via AST traversal instead of regex, enabling pre-demote analysis.
- Published demoted var bindings back to `this` (worker global) when inside the async wrapper, preventing function-scope from hiding them across cells.
- Added `collectBindingNames`/`getLexicalBindingNames` to extract names from destructured patterns.
2026-05-15 18:31:12 +02:00
can1357 f1f6516056 refactor: reorganized exports and removed obsolete helper branches
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
2026-05-14 04:36:19 +02:00
can1357 28b9ce7a0c feat(coding-agent): added middle-elision caps to OutputSink truncation
- Added `tools.artifactHeadBytes` and `tools.outputMaxColumns` settings with defaults in `SETTINGS_SCHEMA`.
- Expanded `OutputSink` with `headBytes`/`maxColumns` and middle truncate logic with elision markers and tracking.
- Updated output-meta to resolve sink settings, emit truncation metrics, and use `truncateMiddle` for spills.
- Integrated head and column limits into JS/Python/Bash/SSH/read output flows, with `:raw` skipping read truncation.
- Documented new output middle-elision and column-cap behavior in `CHANGELOG.md`.
- Added truncation tests for `OutputSink`, `truncateMiddle`, and read-tool line handling.
2026-05-13 11:19:11 +02:00
can1357 db1a3fd7b1 fix(packages/coding-agent): corrected js import rewriting empty AST body
- Added guard for ASTs with no body and trimmed trailing EmptyStatement nodes before final-expression capture.
- Exposed wrapCode via context-manager export for external JS import-rewrite callers.
- Added regression test asserting final-expression wrapping when trailing semicolons follow await.
2026-05-13 05:24:33 +02:00
can1357 218fe8b892 fix(packages/coding-agent): resolved JS display fallback for noncloneable values
- Handled structured-clone failures in JsRuntime.display by falling back to text output.
- Added regression coverage for non-structured-cloneable JS display output.
2026-05-13 05:24:33 +02:00
can1357 1ccc56aca6 fix(coding-agent/eval): routed JS import calls through session-aware import helper
- Updated JS import rewriting to route top-level `import` declarations through `__omp_import__` with support for import attributes.
- Added AST traversal to replace `import(...)` call callee nodes with `__omp_import__` so dynamic imports resolve via session-aware helper.
- Updated runtime `__omp_import__` to accept an optional options object and pass it through to `import(target, options)`.
2026-05-13 04:47:49 +02:00
can1357 a6f80d6a86 Merge PR #1024: feat(coding-agent): return JS eval final expressions
Rewrites the final top-level expression statement to a runtime hook so
its value surfaces without an explicit return/display, while preserving
top-level binding lifetime across cells. Promise-valued and thenable
finals are awaited exactly once; import-only cells stay silent.

Closes #1024
2026-05-13 02:25:04 +02:00
jiwangyihaoandcan1357 56428618c2 fix(js-eval): 不捕获改写后的 import 表达式 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 370fda14ff fix(js-eval): 使用私有最终表达式槽 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 7855028814 fix(js-eval): 仅在重写后读取最终表达式标记 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 51234ad929 fix(js-eval): 忽略继承的最终表达式标记 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 1b0ed01d32 fix(coding-agent): 解析所有最终 Promise 表达式 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 efbf6b3833 fix(coding-agent): 等待最终 Promise 表达式 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 3af9369536 fix(coding-agent): 保留 JS eval 顶层绑定 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 c270e966ba feat(coding-agent): 返回 JS eval 最终表达式 2026-05-13 02:24:10 +02:00
jiwangyihaoandcan1357 0f73414d3b fix(coding-agent): 识别顶层工具错误标记 2026-05-13 02:18:56 +02:00
jiwangyihaoandcan1357 e043125e11 fix(coding-agent): 标记 JS eval 工具错误结果 2026-05-13 02:18:56 +02:00
can1357 1cb451a071 fix(workers): replaced file-URL import pattern with compiled-binary-aware spawn
- Replaced `with { type: "file" }` worker imports with `isCompiledBinary()` hybrid: literal string for `--compile` static analysis, `new URL(import.meta.url)` for dev portability.
- Added worker entrypoints as explicit `--compile` args in `build-binary.ts` so Bun emits them into bunfs.
- Added `smokeTestSyncWorker` and `omp --smoke-test` to catch silent worker-load failures in compiled binaries (fixes #1011, #1027).
- Added `isCompiledBinary()` utility to `@oh-my-pi/pi-utils` detecting bunfs path markers.
2026-05-12 16:40:15 +02:00
can1357 dcf3f1c8a3 feat(coding-agent): added tool.<name>(args) execution in Python prelude
- Added `tool.<name>(args)` execution in Python prelude and docs, enabling direct tool call syntax.
- Added per-execution Python tool-bridge registration, env wiring, and executor lifecycle management.
- Implemented authenticated loopback `/v1/tool` handling with session/name validation and structured responses.
- Updated runtime and lint handling to permit controlled global eval usage in indirect eval paths.
- Added tests for success, failures, invalid bodies, and `emitStatus` propagation in Python tool bridges.
2026-05-12 08:43:36 +02:00
can1357 2f7b2aaccd feat(coding-agent/eval): stripped TypeScript syntax before import rewrites
- Added a stripTypeScript pass that heuristically detects TypeScript-like syntax and transpiles it with Bun's ts loader before rewriting.
- Integrated the new pass into wrapCode so import and lexical rewrites operate on transpiled input.
- Added an executeJs test that verifies interfaces, type annotations, and type assertions execute to the expected numeric result.
2026-05-12 08:42:52 +02:00
can1357 2d4d70e47e feat(coding-agent): added shared JS runtime for worker execution
- Added shared JS runtime infrastructure with RuntimeHooks, RuntimeOptions, JsRuntime, and indirectEval-based execution.
- Reworked worker execution to use runtime.run, track active runs, and clean pending tool calls on close.
- Added tool-call and tool-reply message variants and wired supervisor/worker to route run-scoped calls.
- Added shared helper globals (read, writeFile, diff, tree, env) for browser-tab JS execution.
2026-05-12 08:42:30 +02:00
can1357 efbaa59a05 fix(coding-agent): routed rewritten static imports through cwd-aware helper
- Rewrote static import lowering to emit `__omp_import__` calls instead of direct `import(...)` so rewritten code routes through worker context helpers.
- Added a `__omp_import__` runtime helper in `WorkerCore` that resolves specifiers against the session cwd via `Bun.resolveSync` while passing through URL-like modules unchanged.
- Updated static import rewrite tests to expect the new helper-based `__omp_import__` form, including import attribute cases.
2026-05-12 08:22:55 +02:00
can1357 f56cbbf4c6 feat(stats): added frustration metrics to stats parser and charting
- Added new frustration and revised behavior metrics across stats types, parser mappings, and UI charts.
- Reworked session sync to fan out parsing across a capped worker pool with SyncOptions progress callbacks.
- Added worker-based parse messaging and throttled TTY progress rendering in the stats sync command.
- Updated behavior scoring to strip structured content, ignore noisy prompts, and fix profanity boundary regressions.
- Expanded user message schema and migrations, then repaired assistant model/provider links during sync backfill.
- Updated worker-import guidance in AGENTS.md and recorded sync-progress/metrics updates in package changelogs.
2026-05-12 07:55:23 +02:00
can1357 a7209c2b4f feat(coding-agent-eval): added worker-backed JS runtime with queued runs
- Added worker-backed JS runtime via Transport and WorkerCore, with queued runs, status/error output, and clean teardown.
- Refactored executeInVmContext to reuse worker sessions, route pending runs by id, and fallback inline on spawn failure.
- Added rewrite helpers to convert top-level imports and demote top-level const/let/class declarations for persistence.
- Handled aborts by canceling in-flight tool calls, hard-killing worker sessions, and dropping JS timeout enforcement.
2026-05-12 07:20:43 +02:00
can1357 7d233725b6 refactor(coding-agent/eval): removed eval shell run helpers from JS and Python preludes
- Removed the JS and Python eval prelude `run` helpers, including their shell execution and timeout/cwd option handling.
- Updated the JS VM helper set to expose `Bun` and removed the deleted `run` entry from the prelude.
- Revised eval docs to drop `run` from the helper surface and note the new JS `Bun` global.
2026-05-12 05:43:37 +02:00