Commit Graph

77 Commits

Author SHA1 Message Date
roboomp 03489d1ebe fix(eval): coalesced python kernel replacement
Tracked retained Python kernel generations and shared one replacement promise per dead generation. Reset and disposal now invalidate and drain replacement work before allowing a new session to take ownership.

Added deterministic fake-kernel coverage for concurrent callers, cancellation, reset, owner/global disposal, and independent cwd keys.

Fixes #6367
2026-07-23 19:00:08 +00:00
can1357 b6e68fd243 fix(coding-agent): updated mentions of explore -> scout 2026-07-23 00:12:07 +02:00
vmcall d944879f21 feat(task): unified structured subagent execution
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.

Fixes #5279
2026-07-17 17:36:59 +02:00
can1357 55f5ebec49 chore: reformat 2026-07-15 00:08:50 +02:00
can1357 a9c038818d feat(tools): removed separate selector args from read and grep APIs
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
2026-07-15 00:04:02 +02:00
can1357 37b1263a26 Merge PR #5494: fix(eval): delegate Python URI reads to host resolver (@roboomp) 2026-07-14 23:11:12 +02:00
roboomp 2bfdf0ab07 fix(eval): passed uri selectors separately
- Kept opaque MCP resource paths unchanged during pagination.
- Sent line ranges through the read tool selector field.
- Covered paged artifact and MCP reads in the Python prelude test.

Fixes #5353
2026-07-14 20:20:42 +00:00
roboomp 5d6fcff2f1 fix(eval): delegated python uri reads to host resolver
- Routed non-local URI reads through the session read tool.
- Preserved offset and limit as host line selectors.
- Covered artifact delegation with the shipped Python prelude.

Fixes #5353
2026-07-14 19:12:06 +00:00
can1357 5d28a319fd Merge PR #5015: fix(eval): shield Python agent bridge aborts (@roboomp) 2026-07-14 18:45:03 +02:00
Colin Mason c40ccdc684 isolate eval runtimes from terminal 2026-07-13 09:57:19 -04:00
roboomp a86c1ec46d fix(eval): shield agent bridge aborts
- Deferred eval cancellation while Python bridge calls are paused so already-started agent() subagents can finish and persist output.
- Rejected new Python bridge calls after an external abort is pending to prevent post-abort fan-out waves.
- Added regression coverage for the bridge shield and Python parallel agent() interruption path.

Fixes #5005
2026-07-10 01:21:39 +00:00
can1357 db8c79cc93 style: apply formatter and fix devin test enum
biome + cargo fmt over merge-sweep eval-fix commits; correct StopReason.END_TURN (nonexistent) to StopReason.FUNCTION_CALL in devin streaming test
2026-07-01 04:45:28 +02:00
can1357 5888bccba0 fix(eval): truncate oversized Python shell output 2026-07-01 04:40:55 +02:00
roboomp e348e8d9b5 fix(eval): bound python shell helper output
Stream Python eval shell helper stdout in fixed-size chunks instead of buffering through subprocess.run or newline-bound text iteration.

Added regression coverage for !cmd and newline-free %%bash streaming.

Fixes #3950
2026-07-01 01:15:38 +00:00
can1357 1dd78b207e feat(coding-agent): removed unused eval helper functions
- Removed deprecated eval prelude helpers `append`, `tree`, `diff`, `sort`, `uniq`, and `counter` from all supported runtimes.
- Cleaned up runtime implementations, protocol definitions, and UI rendering logic associated with the removed helpers.
- Updated project documentation, prompts, and test suites to reflect the reduced helper API surface.
- Recorded functional changes in the package changelog.
2026-06-23 01:39:24 +02:00
can1357 9e6eb98db5 refactor(coding-agent/eval): removed unused sh cell magic
- Removed the `_magic_cell_sh` function as `sh` cell magic was no longer required.
2026-06-23 00:27:45 +02:00
can1357 24d58033bb refactor(eval): rename agent() params agent_type→agent, return_handle→handle
The eval agent() helper used `agent_type`/`return_handle` (snake_case) in
Python/Ruby/Julia and `agentType`/`returnHandle` (camelCase) in JS, forcing
the prelude docs to repeat every option twice ("JS same but camelcased").
Both are now single lowercase words identical across all four runtimes, and
`agent` matches the `task` tool's existing agent-selection parameter.

- Renamed across py/js/rb/jl preludes (signatures, forwarding, docstrings).
- Renamed the `__agent__` bridge wire protocol + `EvalAgentArgs` (`agentType`
  → `agent`, `returnHandle` → `handle`) so no prelude-side remap is needed.
- Updated prompt docs (workflow-notice.md, tools/eval.md), repo docs
  (docs/tools/eval.md, docs/python-repl.md), and all bridge/prelude tests.
- CHANGELOG: Breaking Changes entry under [Unreleased].
2026-06-22 22:12:24 +02:00
can1357 2f2acaf928 Merge pull request #3205: feat(eval): isolated/apply/merge options for agent() helper
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.

Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
  forwards isolated/apply/merge (as booleans) plus returnHandle, and the
  return_handle node carries isolated/patch_path/branch_name/
  nested_patches/changes_applied/isolation_summary.

Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
  task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
  the stale "defaults track task.isolation.mode" wording to the final
  strict opt-in behavior, and note all four runtimes.

Fixes #3196
2026-06-22 21:56:45 +02:00
roboomp fbb7255e8e fix(eval): flipped isolated default to strict opt-in
Per maintainer ruling on #3196, eval agent() now defaults to non-isolated regardless of task.isolation.mode, mirroring the task tool. isolated=true is the only way to turn it on; isolated=true while task.isolation.mode === "none" still throws the same clear error.

Updated tests, workflow-notice.md, and Python agent() docstring to reflect the strict opt-in contract. Existing isolation tests now pass isolated:true explicitly; the inherit-from-settings assertion is replaced with a default-off + isolated=true opt-in regression.

Fixes #3196
2026-06-22 18:29:28 +00:00
can1357 a94cadf723 refactor(coding-agent): standardized evaluation kernel and executor architecture
- Extracted common kernel and executor logic into `BaseKernel` and `executor-base` to eliminate duplicated implementations for Julia, Python, and Ruby.
- Migrated shared operational workflows--including session namespacing, environment filtering, and result mapping--to centralized backend helpers.
- Consolidated runtime discovery and resolution logic into a unified `runtime-env` utility module.
- Simplified language-specific modules by delegating subprocess lifecycle, IPC, and configuration management to the newly established base classes.
2026-06-22 07:22:10 +02:00
roboomp 3677e71e9d fix(eval): exposed nested apply=false patches
Branch-mode isolation can capture nested repository changes without creating a root branch. Eval agent() with apply=false previously treated that shape as no captured changes and returned no recoverable nested patch payload after the isolation worktree was removed.

Expose captured nested patches in EvalAgentResult details and copy them onto JS/Python returnHandle nodes (nestedPatches / nested_patches). Document the return_handle escape hatch and add regression coverage for branch-mode nested-only apply=false runs.

Fixes #3196
2026-06-22 04:08:56 +00:00
roboomp a6ee95df08 fix(eval): preserved returnHandle artifacts and nested branch patches
Eval preludes now forward returnHandle to the bridge so no-session eval runs can preserve the temp artifacts backing returned agent:// handles. The bridge keeps those temporary artifact directories whenever returnHandle is requested, including non-isolated runs and successful isolated applies.

Branch-mode isolation now treats nested-only changes as merge-eligible even when no root branch was produced, letting callers apply nested patches instead of dropping them when the root repo had no diff.

Added regression coverage for returnHandle artifact preservation and nested-only branch isolation.

Fixes #3196
2026-06-22 03:59:17 +00:00
roboomp 54f4652573 fix(eval): exposed apply=false artifacts on the agent() return_handle node
When agent() ran with schema and apply=false, the bridge correctly returned the captured patch/branch in details, but the preludes only forwarded id/agent/handle/data on the returnHandle node. Structured workflows had no way to recover the artifact for a manual apply.

Both runtimes now copy isolated, patchPath/branchName, changesApplied, and isolationSummary onto the returnHandle node (snake_case in Python, camelCase in JS), keeping null changesApplied so apply=false stays distinguishable from a successful apply. Updated the workflow notice and the Python agent() docstring to point callers at return_handle as the artifact escape hatch for isolated+apply=false runs. Added prelude tests locking the new node shape in both runtimes.

Fixes #3196
2026-06-21 17:35:26 +00:00
roboomp eeed4b93b6 feat(eval): added isolated/apply/merge options to agent() helper
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.

Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.

Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.

Fixes #3196
2026-06-21 17:20:07 +00:00
can1357 7d28c60c86 refactor: renamed intent field from _i to i
- Renamed the global `INTENT_FIELD` constant from `_i` to `i`.
- Updated documentation strings, type annotations, and test expectations across packages to reflect the new field name.
- Ensured consistent usage of the constant in tool schema construction and intent serialization.
2026-06-19 16:42:33 +02:00
can1357 4f03180ae6 refactor(deps): moved intent field constant to pi-wire
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
2026-06-19 04:58:29 +02:00
can1357 dc915dfef1 feat(coding-agent/eval): tracked and forwarded session id and working directory
- Added `sessionId` and `cwd` tracking to JS evaluator session instances.
- Updated `acquireSession` in the JS context manager to update existing session identity and directory metadata on reuse.
- Forwarded the current working directory from `executePerCall` and `executeOnSession` callers to the underlying Python kernel tasks.
2026-06-18 00:59:56 +02:00
can1357 e4a87fa3ce feat(coding-agent): added matplotlib rendering and image persistence across session reload
- Added Matplotlib figure PNG rendering and display tracking in Python runner to emit PNG output immediately when figures are displayed via display(fig).
- Extended session persistence to externalize oversized image payloads in both content and details.images, enabling tool result images to survive session reload.
- Enhanced session loader to resolve image data payloads and blob references across content and details.images during session reconstruction.
- Added image cache invalidation in TUI image component when image protocol, cell dimensions, or Kitty Unicode placeholder mode changes.
- Added comprehensive test coverage for Matplotlib display, image persistence across reload, and TUI image rendering with protocol and dimension changes.
2026-06-17 13:19:28 +02:00
metaphorics 5d95423515 feat(coding-agent/eval): added return_handle to agent() for DAG handle piping
- Added a return_handle (Python) / returnHandle (JS) option to the eval agent() helper that returns a DAG node dict { text, output, handle, id, agent } instead of bare text, where handle is the spawned agent's recoverable agent:// URI.
- Enabled downstream pipeline/parallel stages to reference a large transcript by handle/output instead of re-inlining it; the default path stays backward compatible (bare text, or the parsed object under schema).
- Documented return_handle and the acyclic DAG-wiring pattern in the eval tool description and added a VM-level prelude regression test for the node shape and the no-details fallback.
2026-06-15 21:54:39 +09:00
can1357 a92d2ce989 feat(coding-agent): removed context argument from eval agent() spawn
Shared background now flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each prompt instead of a context string forwarded into the subagent's system prompt. The JS and Python preludes drop the context kwarg from agent(), the subagent system prompt drops the {{#if context}} block and the conversation-context file pointer, and runEvalAgent no longer writes a per-call conversation context file. AgentSession sheds the now-unused formatCompactContext() helper that supplied the file's body, and ToolSession.getCompactContext is removed alongside it.
2026-06-10 17:49:01 +02:00
can1357 4068bfc304 fix(coding-agent): key python kernel sessions by resolved interpreter 2026-06-10 09:51:45 +02:00
can1357 24cbc93913 fix(eval): thread python.interpreter from session settings and expand ~
Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.

Addresses review feedback on #2204.
2026-06-10 08:26:00 +02:00
danzaio caff395d07 feat(eval): add python interpreter setting 2026-06-10 08:26:00 +02:00
can1357 d7eae06830 fix(coding-agent): fixed eval artifact double-writes and kernel I/O capture
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
2026-06-10 01:27:41 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
can1357 cfd72ffb40 fix(coding-agent): fixed eval helpers treating local:// as plain paths
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
2026-06-08 18:28:06 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
can1357 088fb7fb75 fix(eval): resolved JS/Python resets by awaiting in-flight operations
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
2026-06-08 02:04:59 +02:00
can1357 f73942eb28 fix(eval/py): corrected read() positional offset/limit support and tests
- Updated `packages/coding-agent/src/eval/py/prelude.py` to accept positional `offset` and `limit` in `read()`.
- Added regression coverage in `packages/coding-agent/src/eval/py/__tests__/prelude.test.ts` for positional read signatures.
2026-06-08 02:04:59 +02:00
can1357 76ca917bfd fix(coding-agent-eval): suspended idle timeout during delegated bridge calls
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
2026-06-06 14:02:04 +02:00
roboomp 867ee74a3b fix(eval): probed Win32 console directly instead of inferring from stdio
The TTY-OR heuristic still mis-classifies the all-stdio-redirected case:
`omp -p "..." < in.txt > out.txt 2> err.log` from a real Windows Terminal
session has `stdin.isTTY === stdout.isTTY === stderr.isTTY === false`,
but the host process can still own a console that the kernel could
inherit. The TTY signals reflect handle redirection, not console
attachment — `GetConsoleWindow()` is the authoritative Win32 signal.

`spawn-options.ts` now:

- Calls `kernel32!GetConsoleWindow()` via `bun:ffi` on Windows. A non-NULL
  HWND means the host has a console regardless of how the standard
  streams are wired, so `windowsHide` stays `false` and the kernel
  inherits the console — which is what fixes the #1960
  numpy/pandas `LoadLibraryExW` hang and lets SIGINT recover via
  `GenerateConsoleCtrlEvent`.
- Falls back to the TTY-OR heuristic when the FFI probe is unavailable
  or off-Windows. That keeps the predicate working on non-Bun-FFI
  runtimes and on POSIX, where `windowsHide` is a no-op anyway.
- Caches the probe result; console attachment is stable for the host's
  lifetime in practice, and dlopening kernel32 on every kernel spawn
  would be wasteful.

The pure helpers (`shouldHideKernelWindow`, `consoleAttachedViaTTY`)
stay separately exported so they remain unit-testable. The integration
boundary `hostHasInheritableConsole()` is what the kernel spawn site
calls, and its return is asserted to be a concrete boolean (kernel spawn
must commit to a `windowsHide` value).
2026-06-06 01:04:13 +00:00
roboomp 69ffc80368 style: bun run fix 2026-06-06 00:59:23 +00:00
roboomp 8a89ce7de3 fix(eval): widened Python kernel console detection beyond stdout TTY
Stand-alone `process.stdout.isTTY` mis-classifies any partial stdio
redirection — `omp -p "..." > out.txt` reports `stdout.isTTY === false`
even though the parent still owns a console via stdin/stderr. With the
previous check the kernel would have been spawned with `CREATE_NO_WINDOW`
in that scenario and re-hit the #1960 numpy/pandas `LoadLibraryExW` hang
and broken SIGINT.

The host owns a console it can share with the child whenever ANY of
stdin / stdout / stderr is still a TTY; only a fully detached launch
(true service / daemon, or `< in > out 2> err`) has nothing to inherit.

Renamed the predicate parameter to `hostHasInheritableConsole` so the
contract is unambiguous, and added five regression tests covering the
realistic shell redirection combinations.
2026-06-06 00:59:19 +00:00
roboomp 718c8b299b fix(eval): stopped detaching the Python kernel's console on Windows
`PythonKernel.start()` spawned the runner with `windowsHide: true`, which
in Bun maps to the Win32 `CREATE_NO_WINDOW` flag — that detaches the
long-lived child from any inherited console. Two consequences for OMP's
Python eval on Windows:

1. Native extensions that probe the console at init (e.g. NumPy's
   `_core/_multiarray_umath.pyd` plus its bundled OpenBLAS/SLEEF
   thread-pool init) can deadlock inside `LoadLibraryExW`, so the very
   first `import pandas` / `import numpy` after a cold kernel never
   returned. The reporter's faulthandler stack pinned the hang to
   `_bootstrap_external.create_module` → `numpy/_core/multiarray.py:11`.
2. SIGINT cannot be delivered to a console-less process via
   `GenerateConsoleCtrlEvent`, so the host's `proc.kill("SIGINT")`
   silently no-ops — matching the reporter's "kernel unresponsive to
   interrupt" log line. The 5s escalation then hard-kills the kernel.

Plain `python.exe -u -c "import pandas"` from the same venv inherits the
terminal's console and runs to completion, which is the contract this
change restores. The Python kernel now hides its window only when the
host itself has no console to share (service / piped launch); an
interactive TUI launch lets the kernel inherit the parent's console —
analogous to `python.exe` invoked from `cmd.exe`.

The `shouldHideKernelWindow` predicate lives in its own module so it can
be unit-tested without dragging in the kernel's runtime dependencies.
Short-lived helper subprocesses elsewhere (LSP probes, git, plugin
installs) intentionally keep `windowsHide: true` — they don't load
complex native modules and a brief console flash would be user-visible
noise.

Fixes #1960
2026-06-06 00:53:45 +00:00
can1357 5184b565e1 fix(eval): kept Python kernel alive when parallel() cell interrupted
- Resolved in-flight bridge calls the instant the cell's signal aborts.
- Let the kernel unwind via KeyboardInterrupt instead of being hard-killed.
- Preserved persistent session state across wide subagent fan-out teardown.
2026-06-05 20:42:29 +02:00
can1357 c4e157590e feat(eval): replaced fixed concurrency cap with task.maxConcurrency bridge
- Removed the `concurrency` argument from `parallel()` and `pipeline()` in both JS and Python runtimes.
- Added `__concurrency__` bridge to resolve the pool ceiling live from `task.maxConcurrency` (default 32; 0 = unbounded).
- Eval fan-outs now run as wide as a `task` tool batch instead of being capped at 16.
2026-06-02 06:49:15 +02:00
can1357 dbd9489010 refactor(eval): changed timeout from inactivity to wall-clock budget
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
2026-06-01 17:17:00 +02:00
can1357 e18e4ada71 fix(eval): keep idle watchdog armed during in-flight agent()/llm() calls
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.

Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.

- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
2026-06-01 16:38:32 +02:00
can1357 2003d7382e feat(eval): added per-cell inactivity timeout budgets in eval executors
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
2026-05-31 10:11:59 +02:00
can1357 9fabd5e4d5 feat(coding-agent): added onStatus in eval backends for status streams
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
2026-05-31 08:45:12 +02:00