Commit Graph
79 Commits
Author SHA1 Message Date
can1357 4f03180ae6 refactor(deps): moved intent field constant to pi-wire
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
2026-06-19 04:58:29 +02:00
can1357 dc915dfef1 feat(coding-agent/eval): tracked and forwarded session id and working directory
- Added `sessionId` and `cwd` tracking to JS evaluator session instances.
- Updated `acquireSession` in the JS context manager to update existing session identity and directory metadata on reuse.
- Forwarded the current working directory from `executePerCall` and `executeOnSession` callers to the underlying Python kernel tasks.
2026-06-18 00:59:56 +02:00
can1357 a050474af7 feat: migrated validation schemas and tool definitions from Zod to ArkType
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
2026-06-18 00:59:53 +02:00
can1357 db58a163b2 fix(coding-agent): isolated JS eval runtime globals across runtime lifetimes
- Made `JsRuntime` track the realm globals it installs (`#globalOwner`, `#ownedGlobalKeys`, `PRELUDE_GLOBAL_KEYS`) and reclaim them in `dispose()`, re-binding ownership through `#activateGlobals()` on cwd/run-scope/run; disposing an older inline/direct runtime no longer deletes a newer runtime's helper globals, and overlapping same-realm runs are serialized via `enterGlobalRun`.
- Added `WorkerCore.dispose()` (rejects pending tool calls with `ToolError` and disposes the runtime) and invoked it from `spawnInlineWorker` teardown in `context-manager.ts`.
- Updated the `js-static-import-rewrite` test to snapshot and restore any pre-existing `__omp_import__` global instead of unconditionally deleting it.
- Added `runtime-global-dispose.test.ts` covering newer-runtime global survival, older-runtime reactivation, and cross-runtime mutation rejection.
- Clarified the `runWorkerEntrypoint` comment in `cli.ts` about `parentPort` sync-prefix buffering for tab/eval workers.
2026-06-17 01:21:33 +02:00
metaphorics 5d95423515 feat(coding-agent/eval): added return_handle to agent() for DAG handle piping
- Added a return_handle (Python) / returnHandle (JS) option to the eval agent() helper that returns a DAG node dict { text, output, handle, id, agent } instead of bare text, where handle is the spawned agent's recoverable agent:// URI.
- Enabled downstream pipeline/parallel stages to reference a large transcript by handle/output instead of re-inlining it; the default path stays backward compatible (bare text, or the parsed object under schema).
- Documented return_handle and the acyclic DAG-wiring pattern in the eval tool description and added a VM-level prelude regression test for the node shape and the no-details fallback.
2026-06-15 21:54:39 +09:00
can1357 3ab675a83b fix: added buffered worker inboxing and standardized worker selectors
- Added `WorkerInbox` and `installWorkerInbox(port)` to queue worker messages before bind.
- Added `consumeWorkerInbox()` to replay buffered messages and clear one active inbox.
- Added buffered inbox consumption in JS and tab worker transports before direct message handlers.
- Normalized worker selector arguments to the `__omp_worker_*` naming across workers and tests.
2026-06-15 11:59:44 +02:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
can1357 eaffa00177 fix(coding-agent/eval): routed JS worker ready signaling through init response
- Updated worker core so it emitted `ready` only after receiving an `init` message instead of on construction.
- Changed worker startup to send `init` before waiting for readiness, enabling startup errors to fail fast and trigger fallback.
- Expanded JS worker tests with startup-error simulation and verified execution falls back to the inline worker when spawning fails.
2026-06-15 08:49:00 +02:00
can1357 ebe05c3fc2 fix(coding-agent/eval): fixed JS eval worker initialization with inline fallback on failure
- Initialized JS workers through a new helper that attaches message and error listeners before initialization handshake and enforces the existing ready timeout.
- Handled startup failures by rejecting on init or error events, cleaning up listeners and replacing dead workers with a retry path.
- Retried session creation with an inline worker after non-inline init failure so requests no longer stall on worker startup timeouts.
2026-06-15 08:45:15 +02:00
can1357 7b319f0f02 fix(coding-agent/eval): routed process stdio writes to active runtime text sink
- Added runtime hook resolvers so each `JsRuntime` instance can expose hooks for its active run.
- Patched `process.stdout` and `process.stderr` writes once per stream to route output chunks through active run text hooks and preserve existing worker logging when no run is active.
- Added chunk-to-string conversion for write payloads and encoding-aware forwarding while keeping callback semantics intact.
2026-06-15 03:04:10 +02:00
can1357 74bfcabc35 fix(coding-agent): fixed JS eval helper positional args and URI read slicing
- Updated JS eval helper option parsing to accept positional optional args or a trailing plain-object options argument, while rejecting mixed or invalid forms.
- Updated `read` to route non-`local://` URI paths through the `read` tool with line selectors so offset/limit slicing works for artifact-style resources.
- Expanded JS executor tests for positional reads, nullable slot skipping, and delegated URI read slicing.
2026-06-14 07:27:16 +02:00
can1357 a92d2ce989 feat(coding-agent): removed context argument from eval agent() spawn
Shared background now flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each prompt instead of a context string forwarded into the subagent's system prompt. The JS and Python preludes drop the context kwarg from agent(), the subagent system prompt drops the {{#if context}} block and the conversation-context file pointer, and runEvalAgent no longer writes a per-call conversation context file. AgentSession sheds the now-unused formatCompactContext() helper that supplied the file's body, and ToolSession.getCompactContext is removed alongside it.
2026-06-10 17:49:01 +02:00
can1357 8583421362 perf(coding-agent): deferred heavy imports until first feature use
- Added lazy async module loaders for @babel/parser, linkedom, puppeteer/browsers, @mozilla/readability, @xterm/headless, and mnemopi to avoid loading them during cold startup.
- Added an interactive startup splash before session construction and skipped it for resume/fork/continue, quiet mode, timing mode, or non-TTY runs.
- Updated JS import-rewrite and memory tests to match async parser loading and preloaded mnemopi modules for sync state helpers.
2026-06-10 03:57:32 +02:00
can1357 dc5c93462f feat: rerouted worker subprocesses through the bundled CLI host entrypoint
- Rerouted sync, tab, js-eval, and tiny workers to re-enter CLI modes via `__omp_*` selectors.
- Adjusted `cli.ts` startup to dispatch worker entrypoints before parsing and exit 1 on uncaught errors.
- Bundled CLI as `dist/cli.js` in prepack, switching `omp` binary and published files.
- Removed explicit Bun `--compile` worker entrypoints from build/release scripts in favor of host-entry dispatch.
- Added `declareWorkerHostEntry()` and `workerHostEntry()` environment helpers and `PI_COMPILED` binary detection.
2026-06-10 03:57:31 +02:00
can1357 f638a5b3e0 refactor(coding-agent): added lazy loading for heavy dependencies
- Introduced memoized dynamic import loaders for Babel parser, mnemopi modules, puppeteer, and HTML-related packages.
- Refactored eval import-rewrite helpers and runtime call sites to use asynchronous wrapping and parsing flows.
- Shifted fetch and web-scraper linkedom usage to on-demand imports so heavy modules load only when needed.
2026-06-10 03:09:09 +02:00
can1357 d7eae06830 fix(coding-agent): fixed eval artifact double-writes and kernel I/O capture
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
2026-06-10 01:27:41 +02:00
can1357 3e990c5bea fix(coding-agent): added guarded dynamic import shim for worker fallback
- Rewrote dynamic `import(...)` rewriting to emit a guarded callee that prefers `__omp_import__` and falls back to native `import` when the helper is unavailable.
- Added a shared shim constant and updated import-rewrite tests to verify routed dynamic imports work both with the injected helper and after serializing into a realm without it.
2026-06-09 22:57:19 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
can1357 cfd72ffb40 fix(coding-agent): fixed eval helpers treating local:// as plain paths
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
2026-06-08 18:28:06 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
basedcorp99 1e2bbe71aa fix(eval): use withResolvers for worker close 2026-06-08 11:49:23 +02:00
basedcorp99 1136ba81a1 fix(eval): exit graceful-close JS worker 2026-06-08 11:16:49 +02:00
basedcorp99 a6c8014697 fix(coding-agent): close JS eval workers gracefully 2026-06-08 11:16:49 +02:00
can1357 65bddab719 feat(coding-agent/eval): enforced JS eval helper options as trailing object literals
- Added strict option parsing for JavaScript helpers so read/sort/uniq/counter/tree/llm/agent now reject positional options and non-object option values.
- Introduced `optionsArg` and `isPlainObject` to enforce a single trailing options object and throw clear TypeError messages on invalid calls.
- Updated the eval tool prompt to document that JavaScript helpers require one trailing options object and no extra positional arguments.
2026-06-08 03:33:45 +02:00
can1357 088fb7fb75 fix(eval): resolved JS/Python resets by awaiting in-flight operations
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
2026-06-08 02:04:59 +02:00
Can BölükandGitHub e9e2d02223 Merge pull request #1982 from AsafMah/fix/js-eval-worker-init-flake
fix(eval): floor JS worker init timeout to stop terminate-mid-init CI flake
2026-06-08 00:19:33 +02:00
can1357 76ca917bfd fix(coding-agent-eval): suspended idle timeout during delegated bridge calls
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
2026-06-06 14:02:04 +02:00
Asaf Mahlev ea8d58d3c6 docs(eval): clarify indirect-eval cross-reference in worker-init comment
Address review: shared/indirect-eval.ts documents the vm.runInContext mid-execution terminate-race, not a mid-init one. Reword so a reader doesn't grep that file for an init-specific note.
2026-06-06 12:43:53 +03:00
Asaf Mahlev 236bba96b9 fix(eval): floor JS worker init timeout to stop terminate-mid-init flake
Worker-ready wait reused Bun's 5s default per-test timeout as its floor, so a slow cold-start under --isolate + high CI concurrency was aborted at 5s. The catch then terminate()s a still-initializing Bun worker -- the documented SIGILL/SIGTRAP crash trigger -- crashing the whole test file and intermittently failing unrelated PRs.

Introduce WORKER_INIT_TIMEOUT_MS=15s as a fixed infrastructure floor (independent of, still dominated by, a larger per-cell timeout) and set a 20s file-local setDefaultTimeout in js-executor/js-workflow-helpers tests so cold starts complete instead of being torn down.
2026-06-06 11:56:23 +03:00
can1357 c4e157590e feat(eval): replaced fixed concurrency cap with task.maxConcurrency bridge
- Removed the `concurrency` argument from `parallel()` and `pipeline()` in both JS and Python runtimes.
- Added `__concurrency__` bridge to resolve the pool ceiling live from `task.maxConcurrency` (default 32; 0 = unbounded).
- Eval fan-outs now run as wide as a `task` tool batch instead of being capped at 16.
2026-06-02 06:49:15 +02:00
can1357 dbd9489010 refactor(eval): changed timeout from inactivity to wall-clock budget
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
2026-06-01 17:17:00 +02:00
can1357 e18e4ada71 fix(eval): keep idle watchdog armed during in-flight agent()/llm() calls
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.

Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.

- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
2026-06-01 16:38:32 +02:00
can1357 2003d7382e feat(eval): added per-cell inactivity timeout budgets in eval executors
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
2026-05-31 10:11:59 +02:00
can1357 9fabd5e4d5 feat(coding-agent): added onStatus in eval backends for status streams
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
2026-05-31 08:45:12 +02:00
can1357 2ddc9c5bc9 feat(coding-agent): added turn-budget parsing, multipliers and hard caps
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
2026-05-31 08:03:45 +02:00
can1357 4ab40764d1 refactor(coding-agent): removed args global from eval runtimes
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
2026-05-31 07:54:52 +02:00
can1357 7613c2a913 feat(coding-agent): expanded eval execution and bridge paths with workflow helpers
Expand eval execution/bridge paths with args/log/phase/budget support, add workflow helpers (parallel/pipeline), expose usage statistics, and add eval integration tests and docs.
2026-05-31 07:40:18 +02:00
can1357 cf621d0abf feat(coding-agent-eval): added runEvalAgent bridge for agent plan checks
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
2026-05-31 06:57:02 +02:00
can1357 3f2e4b2fe7 fix(eval): link cyclic local module graphs in one pass
The JS eval kernel's LocalModuleLoader linked and evaluated every local
module individually inside the recursive vm.SourceTextModule linker
callback. On any import cycle that re-enters Bun's node:vm linker
mid-instantiation and segfaults JSC (getImportedModule on a null record,
SIGTRAP at 0xFFFFFFFFFFFFFFF8) — e.g. `await import(".../edit/streaming.ts")`,
whose relative-import subtree is cyclic.

Construct the entire local module graph first, then drive a single
link() + evaluate() from the graph root so cyclic graphs instantiate in
one pass. The shared static linker only constructs dependencies;
external (node_modules) modules stay eagerly loaded since they carry no
imports and cannot form a cycle. Failed loads invalidate any
not-fully-evaluated modules so a retry reconstructs them.

Upstream Bun bug: https://github.com/oven-sh/bun/issues/31623
2026-05-31 03:17:47 +02:00
can1357 91513cdbf3 fix(coding-agent/eval): fixed TypeScript type-only import rewriting
- Updated local module loading to force TS syntax stripping for .ts/.tsx/.mts modules and use the matching Bun transpiler loader.
- Extended TypeScript stripping to detect `import type`/`export type` syntax and rewired wrapping to strip TS syntax after final-expression extraction with TypeScript-aware parsing.
- Added tests verifying type-only imports are handled correctly in evaluator modules and rewritten code no longer contains type-only import declarations.
2026-05-30 18:35:01 +02:00
can1357 e831c2c758 chore: reformat 2026-05-30 18:08:51 +02:00
can1357 8715ed207c feat(coding-agent-eval): added oneshot llm helper and __llm__ bridge
- Added one-shot `llm(prompt, opts)` helpers in JS and Python eval runtimes.
- Added `__llm__` eval bridge wiring for synthetic LLM tool dispatch and status/event output.
- Added `runEvalLlm` with tier-to-model resolution, effort handling, and oneshot completion execution.
- Added structured schema output handling via `respond` tool and JSON fallback parsing.
- Documented new llm behavior in eval docs/changelog and added tests for tier mapping and error cases.
2026-05-30 00:27:04 +02:00
can1357 796c437dc1 feat: overhauled stream timeout and eval session management
- Replaced external watchdog timers with per-request SDK timeouts for first-event budget across OpenAI, Anthropic, and Azure providers.
- Keyed Python shared kernels by (sessionId, cwd) to prevent cross-directory state bleed.
- Deduplicated concurrent cold-start session acquisition for JS and Python executors.
- Moved `isOpenAIResponsesProgressEvent` to shared module and scoped display output routing per run for interleaved async cells.
2026-05-26 16:49:11 +02:00
can1357 8a5b3e9552 feat(eval): added shared executor inheritance for subagents with concurrent async cells
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
2026-05-26 14:37:56 +02:00
can1357 7f0208ac83 feat(coding-agent/eval): added console.table bridge to runtime text output
- Added a `console.table` helper in the JS prelude that forwards calls to the runtime `__omp_table__` hook.
- Implemented `__omp_table__` in the runtime to render tables through `node:console.Console` and emit text via `onText`.
- Added tests verifying `console.table` produced formatted table output and respected the optional columns filter.
2026-05-26 13:16:59 +02:00
can1357 9456e4fe2c fix(coding-agent): coalesce user timeout with JS worker startup window
READY_TIMEOUT_MS was a hard 5s ceiling, ignoring the per-cell timeout the
caller supplied. Cold Bun starts on slow machines routinely exceed 5s
and the worker init failed before user code ran. The window now takes
the larger of the default and the caller timeout.
2026-05-19 19:16:08 +09:00
can1357 7aef1f1cd5 fix(coding-agent): rewrite trailing return statement into final expression
A top-level `return value;` in a JS eval cell was previously swallowed:
returnFinalExpression only handled ExpressionStatement, so ReturnStatement
flipped the IIFE wrapper which discarded the value. Rewrite the trailing
return into __omp_set_final_expr__((expr)) so the existing
final-expression channel surfaces the value just like a trailing
expression.
2026-05-19 19:16:08 +09:00
can1357 226fe87345 fix(coding-agent/eval): hardened image display value coercion to strict base64
- Implemented strict base64 validation and normalization for image `displayValue` payloads in the JS runtime, supporting strict strings, `Uint8Array`, `Buffer`, `ArrayBuffer`, typed-array views, and JSON `Buffer` objects.
- Dropped unrecognized image payloads while emitting a warning and fallback text instead of forwarding malformed data.
- Added tests covering successful coercions and invalid image data rejection paths.
2026-05-19 04:36:30 +02:00
can1357 7901cecf80 feat(coding-agent): added module cache busting for local imports in JS runtime
- Appended a unique nonce query param to local file imports so Bun treats each reload as a fresh module record.
- Restricted cache busting to relative/absolute path specifiers; bare packages and built-ins are left unchanged.
2026-05-16 20:27:45 +02:00
can1357 110a6a3244 fix(coding-agent): persisted bindings from async-wrapped cells via globalThis publish
- Detected async wrapper need via AST traversal instead of regex, enabling pre-demote analysis.
- Published demoted var bindings back to `this` (worker global) when inside the async wrapper, preventing function-scope from hiding them across cells.
- Added `collectBindingNames`/`getLexicalBindingNames` to extract names from destructured patterns.
2026-05-15 18:31:12 +02:00