Commit Graph

225 Commits

Author SHA1 Message Date
vmcall d53cf023b0 fix(task): reconciled structured subagents with upstream
- Preserved the plan-mode capability clamp after upstream removed report_finding.
- Updated persisted-revival coverage for mounted xdev tool activation.
- Applied current formatter output to conflicted runtime files.
2026-07-17 17:38:12 +02:00
vmcall b4952e27c4 refactor(eval): removed dead isolation recovery formatter 2026-07-17 17:38:12 +02:00
vmcall 414ef80c41 fix(task): restored goal activation outside restricted sessions
- Preserved goal-mode tool injection for ordinary explicit tool lists.
- Kept plan-mode LSP and IRC unavailable under the host capability clamp.
- Added regressions for both capability boundaries.
2026-07-17 17:36:59 +02:00
vmcall d944879f21 feat(task): unified structured subagent execution
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.

Fixes #5279
2026-07-17 17:36:59 +02:00
can1357 55f5ebec49 chore: reformat 2026-07-15 00:08:50 +02:00
can1357 a9c038818d feat(tools): removed separate selector args from read and grep APIs
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
2026-07-15 00:04:02 +02:00
can1357 37b1263a26 Merge PR #5494: fix(eval): delegate Python URI reads to host resolver (@roboomp) 2026-07-14 23:11:12 +02:00
can1357 6b381eb765 Merge PR #5451: fix(eval): honor timeout zero and classify session deadlines (@roboomp) 2026-07-14 22:58:49 +02:00
roboomp eef79705a6 fix(eval): passed js uri selectors separately
- Kept opaque MCP resource paths unchanged during JS pagination.
- Sent JS line ranges through the read tool selector field.
- Covered the shipped JS prelude and updated the changelog.

Fixes #5353
2026-07-14 20:31:58 +00:00
roboomp 2bfdf0ab07 fix(eval): passed uri selectors separately
- Kept opaque MCP resource paths unchanged during pagination.
- Sent line ranges through the read tool selector field.
- Covered paged artifact and MCP reads in the Python prelude test.

Fixes #5353
2026-07-14 20:20:42 +00:00
roboomp 5d6fcff2f1 fix(eval): delegated python uri reads to host resolver
- Routed non-local URI reads through the session read tool.
- Preserved offset and limit as host line selectors.
- Covered artifact delegation with the shipped Python prelude.

Fixes #5353
2026-07-14 19:12:06 +00:00
roboomp 8e6d26b1e8 fix(eval): honored unlimited cell timeouts
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.

Fixes #5250
2026-07-14 17:33:56 +00:00
can1357 5e74444c27 fix(test): repair auth-storage-rotation merge resolution and normalize changelogs 2026-07-14 18:50:47 +02:00
can1357 7a2b027e47 fix(eval): keep completion aborts interruptible 2026-07-14 18:45:03 +02:00
can1357 5d28a319fd Merge PR #5015: fix(eval): shield Python agent bridge aborts (@roboomp) 2026-07-14 18:45:03 +02:00
can1357 bfdca36f44 test(coding-agent): use portable eval isolation probe 2026-07-14 18:41:16 +02:00
can1357 79d7b74e24 Merge PR #5327: fix(coding-agent): isolate eval runtimes from terminal (@masonc15) 2026-07-14 18:41:16 +02:00
can1357 f9f6ed9e8d feat(coding-agent): replaced legacy pi/ role alias prefix with
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
2026-07-13 23:26:33 +02:00
Colin Mason ffa879ba2c keep process cwd while other cells are live 2026-07-13 09:57:20 -04:00
Colin Mason 640d7c83b9 propagate ipc clone errors to eval cells 2026-07-13 09:57:20 -04:00
Colin Mason 95620ccd6c preserve async js fallback ladder 2026-07-13 09:57:20 -04:00
Colin Mason 55d9fcfd14 fall back to bun worker on eval spawn failure 2026-07-13 09:57:20 -04:00
Colin Mason 1a527d9a2f mirror session cwd in js eval subprocess 2026-07-13 09:57:19 -04:00
Colin Mason 9c9da2387d harden js eval subprocess 2026-07-13 09:57:19 -04:00
Colin Mason c40ccdc684 isolate eval runtimes from terminal 2026-07-13 09:57:19 -04:00
can1357 b6f83021c9 fix(coding-agent): ensured top-level declarations persist in async cells
- Updated import rewriting to identify and publish `var` and `function` declarations to the global scope when a cell contains top-level `await`.
- Prevented these declarations from being trapped within the async wrapper's function scope, allowing them to remain accessible to subsequent evaluation cells.
2026-07-12 13:21:50 +02:00
can1357 70754dfa01 fix(coding-agent): hardened same-realm runtime guards from PR review
- setCwd now updates the saved __omp_session__ stack entry so a deferred
  cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
  a first init during another runtime's live run fails via init-failed
  instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
  setup throw under an already-aborted signal cannot become an unhandled
  rejection (review P2)
- credited #4907 in the changelog entry
2026-07-10 12:33:53 +02:00
roboomp a86c1ec46d fix(eval): shield agent bridge aborts
- Deferred eval cancellation while Python bridge calls are paused so already-started agent() subagents can finish and persist output.
- Rejected new Python bridge calls after an external abort is pending to prevent post-abort fan-out waves.
- Added regression coverage for the bridge shield and Python parallel agent() interruption path.

Fixes #5005
2026-07-10 01:21:39 +00:00
ben 71c9cc0324 fix(coding-agent): keep same-realm setCwd from killing TUI sessions
When the JS eval worker falls back to the in-process inline path, concurrent
JsRuntime instances share one realm. setCwd used to throw on exclusive-owner
conflicts, and the microtask delivery path turned that into a fatal
unhandledRejection that postmortem exited on. Stamp local cwd without
stealing the active realm, report init failures over the worker protocol,
and cover process survival with in-process and child-process regressions.
2026-07-09 16:02:23 +08:00
can1357 f03f310495 merge PR #4839: fix(coding-agent): handle JS eval worker cwd conflicts via worker protocol 2026-07-08 15:21:30 +02:00
can1357 38486e56db feat(coding-agent): handled unawaited promise rejections in eval cells
- Introduced a rejection interception mechanism to capture unhandled promise rejections from eval cell code.
- Attributed floating rejections to specific runs to fail the owning cell instead of crashing the process or worker.
- Downgraded rejections occurring after a cell finished to warn logs to prevent silent failures.
2026-07-08 14:51:25 +02:00
ben 9cb051a4b1 Fix JS worker cwd conflict handling 2026-07-08 18:32:05 +08:00
can1357 64eab4e1c6 Merge PR #4273: fix(eval): surface unexpected JS eval worker exit via close listener (@metaphorics) 2026-07-05 13:39:09 +02:00
can1357 af146b1118 Merge PR #4439: fix(task): preserve isolated commits on transfer failure (@roboomp) 2026-07-05 13:10:25 +02:00
roboomp a52ed682c7 fix(ai): separated codex orchestration usage
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.

Fixes #4469
2026-07-03 16:44:12 +00:00
roboomp 437a871304 fix(eval): surfaced preserved patch path in agent() failure
runEvalAgent short-circuits on subagent error before mergeIsolatedChanges runs, so branch-mode transfer failures previously threw with just result.error and the caller never learned where the preserved patch lived. Enrich the pre-merge failure and reuse a shared recovery-hint builder for the apply/nested-patch failure paths so the captured patch path, branch name, and nested-patch artifacts are always surfaced.

Fixes #4437
2026-07-03 14:51:52 +00:00
metaphorics 6c94081888 fix(eval): surface unexpected JS eval worker exit via close listener
In `wrapBunWorker` (`packages/coding-agent/src/eval/js/context-manager.ts`), `error` and `messageerror` listeners were registered during normal operation, but the `close` listener was only added inside `close()`. If user code called `process.exit(0)` or the Bun worker otherwise exited cleanly, the `close` event fired with no handler, so `runOnce` never rejected and callers hung until the cell timeout.

Add a normal-operation `close` listener in `wrapBunWorker.onError` that forwards `new Error("JS eval worker exited")` through the existing error handler path, and remove it in the returned unsubscribe callback.

Verified with `bun run check:types` and `bun test test/tools/eval-*.test.ts test/core/eval-workflow-helpers.integration.test.ts` (32 pass, 0 fail).

Closes #4244
2026-07-02 17:29:42 +09:00
can1357 237b672849 refactor(coding-agent): conditionalized output schema properties
- Conditionalized output schema properties using object spreading to avoid passing explicit undefined values.
2026-07-02 02:18:54 +02:00
can1357 1cb8608a58 Merge PR #3896: fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths (@roboomp)
# Conflicts:
#	packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts
#	packages/coding-agent/src/eval/agent-bridge.ts
2026-07-01 22:07:28 +02:00
can1357 88e3e77f3e Merge PR #3927: fix(agent): reject stale yield labels for override schemas (@roboomp) 2026-07-01 21:50:54 +02:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 b8a996ac56 Merge branch 'farm/af5c9fdb/fix-eval-spawn-default' 2026-07-01 04:46:23 +02:00
can1357 db8c79cc93 style: apply formatter and fix devin test enum
biome + cargo fmt over merge-sweep eval-fix commits; correct StopReason.END_TURN (nonexistent) to StopReason.FUNCTION_CALL in devin streaming test
2026-07-01 04:45:28 +02:00
roboomp 8614b4c086 fix(task): respected restricted spawn defaults
Resolved eval agent() and task tool defaults from the active spawn policy so restricted agents advertise and execute an allowed default.

Fixes #3973
2026-07-01 02:44:08 +00:00
can1357 5888bccba0 fix(eval): truncate oversized Python shell output 2026-07-01 04:40:55 +02:00
roboomp e348e8d9b5 fix(eval): bound python shell helper output
Stream Python eval shell helper stdout in fixed-size chunks instead of buffering through subprocess.run or newline-bound text iteration.

Added regression coverage for !cmd and newline-free %%bash streaming.

Fixes #3950
2026-07-01 01:15:38 +00:00
roboomp 00ef58f843 fix(agent): steered override-schema subagents
- Marked eval agent schema calls as caller overrides so subagent prompts can revoke native output/yield instructions.\n- Added override-schema prompt guidance telling agents to ignore conflicting native output labels and terminal-yield the caller schema object.\n- Added prompt coverage for the override notice.\n\nRefs #3926
2026-06-30 22:30:15 +00:00
roboomp 1b9c6be129 fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths
Three independent paths bypassed the user's subagent caps:

1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
   task.maxConcurrency only on first use and never re-read the setting,
   so lowering the cap mid-session left every later spawn running
   against the old ceiling. Resize the live semaphore against the
   current setting on each acquire.

2. The task tool prompt threaded MAX_CONCURRENCY through to the
   template but never rendered it. A model with task.maxConcurrency=1
   could still emit oversized tasks[] batches that registered
   immediately and piled up behind the semaphore. Render a 'Concurrency
   cap' directive in task.md whenever the setting is bounded.

3. The eval agent() bridge's assertDepthAllowed gated only against
   the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
   task.maxRecursionDepth, so a user-tightened recursion limit
   (0='None', 1='Single') still let cell-spawned subagents recurse
   to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
   by the hard ceiling.

Fixes #3895
2026-06-30 11:37:58 +00:00
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
can1357 037c528e54 Merge PR #3350: increase Julia prelude test timeout to fix CI flake (@oldschoola)
# Conflicts:
#	packages/coding-agent/src/eval/__tests__/julia-prelude.test.ts
2026-06-28 18:41:29 +02:00