Commit Graph
198 Commits
Author SHA1 Message Date
can1357 70754dfa01 fix(coding-agent): hardened same-realm runtime guards from PR review
- setCwd now updates the saved __omp_session__ stack entry so a deferred
  cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
  a first init during another runtime's live run fails via init-failed
  instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
  setup throw under an already-aborted signal cannot become an unhandled
  rejection (review P2)
- credited #4907 in the changelog entry
2026-07-10 12:33:53 +02:00
ben 71c9cc0324 fix(coding-agent): keep same-realm setCwd from killing TUI sessions
When the JS eval worker falls back to the in-process inline path, concurrent
JsRuntime instances share one realm. setCwd used to throw on exclusive-owner
conflicts, and the microtask delivery path turned that into a fatal
unhandledRejection that postmortem exited on. Stamp local cwd without
stealing the active realm, report init failures over the worker protocol,
and cover process survival with in-process and child-process regressions.
2026-07-09 16:02:23 +08:00
can1357 f03f310495 merge PR #4839: fix(coding-agent): handle JS eval worker cwd conflicts via worker protocol 2026-07-08 15:21:30 +02:00
can1357 38486e56db feat(coding-agent): handled unawaited promise rejections in eval cells
- Introduced a rejection interception mechanism to capture unhandled promise rejections from eval cell code.
- Attributed floating rejections to specific runs to fail the owning cell instead of crashing the process or worker.
- Downgraded rejections occurring after a cell finished to warn logs to prevent silent failures.
2026-07-08 14:51:25 +02:00
ben 9cb051a4b1 Fix JS worker cwd conflict handling 2026-07-08 18:32:05 +08:00
can1357 64eab4e1c6 Merge PR #4273: fix(eval): surface unexpected JS eval worker exit via close listener (@metaphorics) 2026-07-05 13:39:09 +02:00
can1357 af146b1118 Merge PR #4439: fix(task): preserve isolated commits on transfer failure (@roboomp) 2026-07-05 13:10:25 +02:00
roboomp a52ed682c7 fix(ai): separated codex orchestration usage
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.

Fixes #4469
2026-07-03 16:44:12 +00:00
roboomp 437a871304 fix(eval): surfaced preserved patch path in agent() failure
runEvalAgent short-circuits on subagent error before mergeIsolatedChanges runs, so branch-mode transfer failures previously threw with just result.error and the caller never learned where the preserved patch lived. Enrich the pre-merge failure and reuse a shared recovery-hint builder for the apply/nested-patch failure paths so the captured patch path, branch name, and nested-patch artifacts are always surfaced.

Fixes #4437
2026-07-03 14:51:52 +00:00
metaphorics 6c94081888 fix(eval): surface unexpected JS eval worker exit via close listener
In `wrapBunWorker` (`packages/coding-agent/src/eval/js/context-manager.ts`), `error` and `messageerror` listeners were registered during normal operation, but the `close` listener was only added inside `close()`. If user code called `process.exit(0)` or the Bun worker otherwise exited cleanly, the `close` event fired with no handler, so `runOnce` never rejected and callers hung until the cell timeout.

Add a normal-operation `close` listener in `wrapBunWorker.onError` that forwards `new Error("JS eval worker exited")` through the existing error handler path, and remove it in the returned unsubscribe callback.

Verified with `bun run check:types` and `bun test test/tools/eval-*.test.ts test/core/eval-workflow-helpers.integration.test.ts` (32 pass, 0 fail).

Closes #4244
2026-07-02 17:29:42 +09:00
can1357 237b672849 refactor(coding-agent): conditionalized output schema properties
- Conditionalized output schema properties using object spreading to avoid passing explicit undefined values.
2026-07-02 02:18:54 +02:00
can1357 1cb8608a58 Merge PR #3896: fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths (@roboomp)
# Conflicts:
#	packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts
#	packages/coding-agent/src/eval/agent-bridge.ts
2026-07-01 22:07:28 +02:00
can1357 88e3e77f3e Merge PR #3927: fix(agent): reject stale yield labels for override schemas (@roboomp) 2026-07-01 21:50:54 +02:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 b8a996ac56 Merge branch 'farm/af5c9fdb/fix-eval-spawn-default' 2026-07-01 04:46:23 +02:00
can1357 db8c79cc93 style: apply formatter and fix devin test enum
biome + cargo fmt over merge-sweep eval-fix commits; correct StopReason.END_TURN (nonexistent) to StopReason.FUNCTION_CALL in devin streaming test
2026-07-01 04:45:28 +02:00
roboomp 8614b4c086 fix(task): respected restricted spawn defaults
Resolved eval agent() and task tool defaults from the active spawn policy so restricted agents advertise and execute an allowed default.

Fixes #3973
2026-07-01 02:44:08 +00:00
can1357 5888bccba0 fix(eval): truncate oversized Python shell output 2026-07-01 04:40:55 +02:00
roboomp e348e8d9b5 fix(eval): bound python shell helper output
Stream Python eval shell helper stdout in fixed-size chunks instead of buffering through subprocess.run or newline-bound text iteration.

Added regression coverage for !cmd and newline-free %%bash streaming.

Fixes #3950
2026-07-01 01:15:38 +00:00
roboomp 00ef58f843 fix(agent): steered override-schema subagents
- Marked eval agent schema calls as caller overrides so subagent prompts can revoke native output/yield instructions.\n- Added override-schema prompt guidance telling agents to ignore conflicting native output labels and terminal-yield the caller schema object.\n- Added prompt coverage for the override notice.\n\nRefs #3926
2026-06-30 22:30:15 +00:00
roboomp 1b9c6be129 fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths
Three independent paths bypassed the user's subagent caps:

1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
   task.maxConcurrency only on first use and never re-read the setting,
   so lowering the cap mid-session left every later spawn running
   against the old ceiling. Resize the live semaphore against the
   current setting on each acquire.

2. The task tool prompt threaded MAX_CONCURRENCY through to the
   template but never rendered it. A model with task.maxConcurrency=1
   could still emit oversized tasks[] batches that registered
   immediately and piled up behind the semaphore. Render a 'Concurrency
   cap' directive in task.md whenever the setting is bounded.

3. The eval agent() bridge's assertDepthAllowed gated only against
   the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
   task.maxRecursionDepth, so a user-tightened recursion limit
   (0='None', 1='Single') still let cell-spawned subagents recurse
   to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
   by the hard ceiling.

Fixes #3895
2026-06-30 11:37:58 +00:00
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
can1357 037c528e54 Merge PR #3350: increase Julia prelude test timeout to fix CI flake (@oldschoola)
# Conflicts:
#	packages/coding-agent/src/eval/__tests__/julia-prelude.test.ts
2026-06-28 18:41:29 +02:00
can1357 2900730aa6 fix(eval): serialize vm link passes across concurrent local imports
Concurrent graph-root loads in the JS eval kernel segfaulted Bun
(SIGSEGV at 0xFFFFFFFFFFFFFFF8, getImportedModule on a null record in
JSC::AbstractModuleRecord::innerModuleLinking). Two roots loaded at once
over an overlapping local-import graph — e.g.
Promise.all([import("./a.ts"), import("./b.ts")]) sharing a dependency —
each launched its own module.link() over the same shared
vm.SourceTextModule instances. The async link resolver yields
mid-instantiation, letting the two link passes interleave and re-enter
Bun's node:vm linker, which crashes instead of throwing.

The earlier single-pass fix (15.7.2) only serialized linking within one
graph root; it did not guard two concurrent roots over shared modules.

LocalModuleLoader now routes every module.link() through a promise-chain
mutex (#serializeLink), so the linker is never re-entered
mid-instantiation. The lock is held only across link(), not evaluate(),
so a dynamic import during evaluation re-acquires it without deadlock,
and the chain swallows rejections so a failed link cannot wedge later
imports. Reproduced and verified against the user crash trace: the
concurrent-diamond repro went from 5/6 crashes to 0/N with correct
namespaces.
2026-06-28 07:27:01 +02:00
can1357 577d2a8eb8 style: biome format/organize-imports across integrated PRs 2026-06-27 02:06:38 +02:00
can1357 29269d9547 Fix eval agent isolation artifact recovery 2026-06-27 01:39:35 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 84ef014faa Merge PR #3407: fix(eval): dispose one-shot eval subagents after run (@korri123) 2026-06-26 23:27:40 +02:00
can1357 305db55cd7 fix(coding-agent/eval): surfaced the exception type and message in the
- Surface the exception type and message in the error display by prepending the formatted error string to the traceback array.
- Prevent the host from hiding the actual error by ensuring the traceback is not empty, consistent with other language runners.
2026-06-26 10:06:48 +02:00
can1357 938489f3fd feat(coding-agent): added configurable service tier settings for subagents and advisor
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
2026-06-25 22:35:15 +02:00
Kormákur 0a7ec389a0 fix(coding-agent): verify eval subagent cleanup 2026-06-25 09:27:38 +00:00
Kormákur 9118e2cb95 fix(coding-agent): dispose eval subagents after run 2026-06-24 22:15:03 +00:00
oldschoola 332a761625 test(eval): increase Julia prelude test and afterEach timeout to fix CI flake 2026-06-23 13:37:43 -07:00
can1357 1dd78b207e feat(coding-agent): removed unused eval helper functions
- Removed deprecated eval prelude helpers `append`, `tree`, `diff`, `sort`, `uniq`, and `counter` from all supported runtimes.
- Cleaned up runtime implementations, protocol definitions, and UI rendering logic associated with the removed helpers.
- Updated project documentation, prompts, and test suites to reflect the reduced helper API surface.
- Recorded functional changes in the package changelog.
2026-06-23 01:39:24 +02:00
can1357 9e6eb98db5 refactor(coding-agent/eval): removed unused sh cell magic
- Removed the `_magic_cell_sh` function as `sh` cell magic was no longer required.
2026-06-23 00:27:45 +02:00
can1357 dc8dfa48c9 feat(coding-agent/eval): expanded rich media support for cross-language evaluation
- Added headless plot configuration for Julia to prevent GUI popup windows during execution.
- Implemented robust mime-bundle serialization in Julia using `invokelatest` to handle runtime-loaded library methods.
- Added IRuby-protocol and magic-byte image sniffing support to the Ruby runner to enable inline rendering for graphics gems like Gruff, ChunkyPNG, and RMagick.
2026-06-23 00:15:28 +02:00
can1357 24d58033bb refactor(eval): rename agent() params agent_type→agent, return_handle→handle
The eval agent() helper used `agent_type`/`return_handle` (snake_case) in
Python/Ruby/Julia and `agentType`/`returnHandle` (camelCase) in JS, forcing
the prelude docs to repeat every option twice ("JS same but camelcased").
Both are now single lowercase words identical across all four runtimes, and
`agent` matches the `task` tool's existing agent-selection parameter.

- Renamed across py/js/rb/jl preludes (signatures, forwarding, docstrings).
- Renamed the `__agent__` bridge wire protocol + `EvalAgentArgs` (`agentType`
  → `agent`, `returnHandle` → `handle`) so no prelude-side remap is needed.
- Updated prompt docs (workflow-notice.md, tools/eval.md), repo docs
  (docs/tools/eval.md, docs/python-repl.md), and all bridge/prelude tests.
- CHANGELOG: Breaking Changes entry under [Unreleased].
2026-06-22 22:12:24 +02:00
roboompandcan1357 5b6e9f904d fix(eval): paused timeout over baseline capture and surfaced nested stash-restore failures
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:

1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.

2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.

Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.

Fixes #3196
2026-06-22 21:57:35 +02:00
can1357 2f2acaf928 Merge pull request #3205: feat(eval): isolated/apply/merge options for agent() helper
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.

Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
  forwards isolated/apply/merge (as booleans) plus returnHandle, and the
  return_handle node carries isolated/patch_path/branch_name/
  nested_patches/changes_applied/isolation_summary.

Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
  task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
  the stale "defaults track task.isolation.mode" wording to the final
  strict opt-in behavior, and note all four runtimes.

Fixes #3196
2026-06-22 21:56:45 +02:00
roboomp 28137f46d9 refactor(task): deduped nested patch apply + commit-message factory
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.

Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.

Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.

Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.

Fixes #3196
2026-06-22 19:22:24 +00:00
roboomp 4c3d785048 test(eval): raised Julia prelude per-test timeout to 30s
The two julia-prelude tests pay a ~11-12s Julia kernel cold-start
(JIT + package precompile) when run with reset: true, exceeding
Bun's default 5000ms per-test timeout. CI exposed this since 33e2594f0
(Julia eval support) ran on PR runners (ubuntu-22.04, Julia
preinstalled) rather than the omp-kata self-hosted runners that lack
Julia and skip the suite.

Bumped both tests to 30_000ms, matching the precedent in
agent-bridge.test.ts:540 for similar persistent-kernel tests.

Fixes #3274
2026-06-22 18:41:45 +00:00
roboomp 15ae84f0de fix(eval): paused timeout through isolation merge/apply
Previously withBridgeTimeoutPause only wrapped the subagent subprocess; mergeIsolatedChanges, applyNestedPatches, nested commit-message generation, and artifact cleanup ran with the eval watchdog re-armed. A cherry-pick or large patch apply could trip the cell timeout and abort successful post-processing.

Moved the entire bridge work (subprocess + merge + nested apply + cleanup + usage recording) inside one withBridgeTimeoutPause block. The pause helper still resumes via its finally on success and on throw, so existing failure paths are unchanged.

Added a regression that captures the emitted op order and asserts merge fires after timeout-pause and before timeout-resume.

Fixes #3196
2026-06-22 18:41:44 +00:00
roboomp 7438efc5b3 style: bun run fix 2026-06-22 18:29:34 +00:00
roboomp fbb7255e8e fix(eval): flipped isolated default to strict opt-in
Per maintainer ruling on #3196, eval agent() now defaults to non-isolated regardless of task.isolation.mode, mirroring the task tool. isolated=true is the only way to turn it on; isolated=true while task.isolation.mode === "none" still throws the same clear error.

Updated tests, workflow-notice.md, and Python agent() docstring to reflect the strict opt-in contract. Existing isolation tests now pass isolated:true explicitly; the inherit-from-settings assertion is replaced with a default-off + isolated=true opt-in regression.

Fixes #3196
2026-06-22 18:29:28 +00:00
can1357 a94cadf723 refactor(coding-agent): standardized evaluation kernel and executor architecture
- Extracted common kernel and executor logic into `BaseKernel` and `executor-base` to eliminate duplicated implementations for Julia, Python, and Ruby.
- Migrated shared operational workflows--including session namespacing, environment filtering, and result mapping--to centralized backend helpers.
- Consolidated runtime discovery and resolution logic into a unified `runtime-env` utility module.
- Simplified language-specific modules by delegating subprocess lifecycle, IPC, and configuration management to the newly established base classes.
2026-06-22 07:22:10 +02:00
roboomp a48f1ea813 fix(eval): persisted nested patches before throwing apply failures
When an isolated apply fails the bridge throws a ToolError and never returns details, so the nested-patch payload that previously lived in details.nestedPatches was unrecoverable after the isolation worktree was torn down.

The bridge now writes each captured nested patch to a file under the per-call artifacts dir (e.g. <agentId>.nested-<index>-<slug>.patch) before throwing and includes the resolved paths in the error message so the caller can apply them manually.

Added a regression test verifying the persisted file exists with the original patch contents and that the path is surfaced in the thrown error.

Fixes #3196
2026-06-22 04:21:11 +00:00
roboomp c1b01f1cb9 fix(eval): surfaced isolated apply failures as ToolError
A failed isolated apply (changesApplied === false) previously only set details.isolationSummary and returned the subagent text. Schema-backed agent() calls then parsed the JSON and returned the object, so workflows saw a successful structured result while none of the edits had landed.

Throw a ToolError when mergeIsolatedChanges reports a failed apply, with the merge summary plus a recovery hint pointing at the preserved patch/branch/nested artifacts so the caller can apply manually.

Added regression tests for the schema and non-schema apply-failure paths.

Fixes #3196
2026-06-22 04:13:43 +00:00
can1357 33e2594f03 feat(coding-agent): added support for Julia and display language icons in code cells
- Added Julia language support to the theme symbol maps.
- Enabled language icons in code cell headers for the eval tool renderer.
2026-06-22 06:13:01 +02:00
can1357 1f3f3cf5d1 feat: added ruby and julia language support to coding-agent
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
2026-06-22 06:13:01 +02:00
roboomp 3677e71e9d fix(eval): exposed nested apply=false patches
Branch-mode isolation can capture nested repository changes without creating a root branch. Eval agent() with apply=false previously treated that shape as no captured changes and returned no recoverable nested patch payload after the isolation worktree was removed.

Expose captured nested patches in EvalAgentResult details and copy them onto JS/Python returnHandle nodes (nestedPatches / nested_patches). Document the return_handle escape hatch and add regression coverage for branch-mode nested-only apply=false runs.

Fixes #3196
2026-06-22 04:08:56 +00:00