- Preserved goal-mode tool injection for ordinary explicit tool lists.
- Kept plan-mode LSP and IRC unavailable under the host capability clamp.
- Added regressions for both capability boundaries.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
- Kept opaque MCP resource paths unchanged during JS pagination.
- Sent JS line ranges through the read tool selector field.
- Covered the shipped JS prelude and updated the changelog.
Fixes#5353
- Kept opaque MCP resource paths unchanged during pagination.
- Sent line ranges through the read tool selector field.
- Covered paged artifact and MCP reads in the Python prelude test.
Fixes#5353
- Routed non-local URI reads through the session read tool.
- Preserved offset and limit as host line selectors.
- Covered artifact delegation with the shipped Python prelude.
Fixes#5353
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.
Fixes#5250
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Updated import rewriting to identify and publish `var` and `function` declarations to the global scope when a cell contains top-level `await`.
- Prevented these declarations from being trapped within the async wrapper's function scope, allowing them to remain accessible to subsequent evaluation cells.
- setCwd now updates the saved __omp_session__ stack entry so a deferred
cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
a first init during another runtime's live run fails via init-failed
instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
setup throw under an already-aborted signal cannot become an unhandled
rejection (review P2)
- credited #4907 in the changelog entry
- Deferred eval cancellation while Python bridge calls are paused so already-started agent() subagents can finish and persist output.
- Rejected new Python bridge calls after an external abort is pending to prevent post-abort fan-out waves.
- Added regression coverage for the bridge shield and Python parallel agent() interruption path.
Fixes#5005
When the JS eval worker falls back to the in-process inline path, concurrent
JsRuntime instances share one realm. setCwd used to throw on exclusive-owner
conflicts, and the microtask delivery path turned that into a fatal
unhandledRejection that postmortem exited on. Stamp local cwd without
stealing the active realm, report init failures over the worker protocol,
and cover process survival with in-process and child-process regressions.
- Introduced a rejection interception mechanism to capture unhandled promise rejections from eval cell code.
- Attributed floating rejections to specific runs to fail the owning cell instead of crashing the process or worker.
- Downgraded rejections occurring after a cell finished to warn logs to prevent silent failures.
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.
Fixes#4469
runEvalAgent short-circuits on subagent error before mergeIsolatedChanges runs, so branch-mode transfer failures previously threw with just result.error and the caller never learned where the preserved patch lived. Enrich the pre-merge failure and reuse a shared recovery-hint builder for the apply/nested-patch failure paths so the captured patch path, branch name, and nested-patch artifacts are always surfaced.
Fixes#4437
In `wrapBunWorker` (`packages/coding-agent/src/eval/js/context-manager.ts`), `error` and `messageerror` listeners were registered during normal operation, but the `close` listener was only added inside `close()`. If user code called `process.exit(0)` or the Bun worker otherwise exited cleanly, the `close` event fired with no handler, so `runOnce` never rejected and callers hung until the cell timeout.
Add a normal-operation `close` listener in `wrapBunWorker.onError` that forwards `new Error("JS eval worker exited")` through the existing error handler path, and remove it in the returned unsubscribe callback.
Verified with `bun run check:types` and `bun test test/tools/eval-*.test.ts test/core/eval-workflow-helpers.integration.test.ts` (32 pass, 0 fail).
Closes#4244
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Stream Python eval shell helper stdout in fixed-size chunks instead of buffering through subprocess.run or newline-bound text iteration.
Added regression coverage for !cmd and newline-free %%bash streaming.
Fixes#3950
Three independent paths bypassed the user's subagent caps:
1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
task.maxConcurrency only on first use and never re-read the setting,
so lowering the cap mid-session left every later spawn running
against the old ceiling. Resize the live semaphore against the
current setting on each acquire.
2. The task tool prompt threaded MAX_CONCURRENCY through to the
template but never rendered it. A model with task.maxConcurrency=1
could still emit oversized tasks[] batches that registered
immediately and piled up behind the semaphore. Render a 'Concurrency
cap' directive in task.md whenever the setting is bounded.
3. The eval agent() bridge's assertDepthAllowed gated only against
the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
task.maxRecursionDepth, so a user-tightened recursion limit
(0='None', 1='Single') still let cell-spawned subagents recurse
to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
by the hard ceiling.
Fixes#3895
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.