- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
Dirty isolated baselines can be accidentally committed by subagents that run git add -A. Fetching the raw isolation HEAD then cherry-picking the range would replay that baseline WIP into parent history.
Add a dirty-baseline replay path that rewrites each agent commit against the captured baseline tree, preserving the agent commit message and author while excluding staged, unstaged, and untracked changes that existed before isolation started. Clean baselines still use the raw git fetch path, and nested-only changes keep returning patches without creating an empty root branch.
Add a regression for baseline staged + untracked WIP committed by the agent, asserting the task branch contains only the agent file and parent WIP remains staged/untracked after merge.
Fixes#3842
When an isolated task agent commits its own changes before yielding, the
harness used to collapse the captured delta into one AI-summarized commit
and discard the agent's commit messages and authorship entirely. This
violated commit discipline for agentic swarms — multiple logical commits
("fix bug" + "add test") became a single opaque commit, and the
agent's commit object (which lived in isolation/.git/objects under
overlayfs/rcopy) was lost when cleanupIsolation tore down the overlay.
commitToBranch now detects when isolation HEAD moved past baseline.root
.headCommit. When it has, the function git-fetches the agent's HEAD into
the parent repo as omp/task/${taskId} so the commit objects survive
cleanupIsolation, and stamps the captured baselineSha onto the returned
CommitToBranchResult. mergeTaskBranches cherry-picks the inclusive range
baseSha..branchName when baseSha is provided, replaying each agent
commit verbatim with its original message and author. Any uncommitted
leftover (staged, unstaged, untracked) on top of the agent's last commit
becomes one trailing AI-summarized commit on the same branch.
Falls back to the legacy single-commit path when the agent never moved
HEAD (purely dirty working tree); existing patch-mode flow is untouched.
Fixes#3842
Applied isolated branch patches with three-way fallback when unrelated parent dirt appears in patch context.
Surfaced branch preparation failures instead of reporting no changes.
Fixes#3841
Used compact hashed isolation directory segments and the short m mount dir so long task ids are not copied into subagent working paths.
Kept worktree cleanup compatible with legacy merged task-isolation directories.
Fixes#3756
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.
Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).
Fixes#3749
- Refactored payload assembly to distinguish between incremental sections and terminal results.
- Prevented terminal markers from being incorrectly treated as section labels to avoid payload nesting issues.
- Updated section processing to ignore non-incremental terminal items, resolving incorrectly missing data in output-schema validation.
- Improved terminal item resolution to correctly fallback to the last assistant text when no explicit data is provided.
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
- Extracted taskToolRenderer to a dedicated renderer file to resolve circular dependencies.
- Updated all references to the renderer to point to the new location.
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
- Updated task rendering to use `previewLine` for truncating descriptions consistently.
- Ensured progress and result descriptions are truncated before formatting.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
In #registerSpawnJob the markRunning()/reportProgress() calls sat between
semaphore.acquire() and the try whose finally releases the slot. If progress
reporting threw there, the acquired task.maxConcurrency slot leaked and
permanently shrank subagent concurrency. Move those statements inside the try
so finally always releases. The abort-before-execution branch is unchanged
(releases once and throws before the try is entered — not a double release).
Refs #3464
The per-provider subagent limiter (providers.ollama-cloud.maxConcurrency)
created a fresh Semaphore whenever the configured limit changed, orphaning
in-flight slots on the old instance so a runtime or mixed limit value could
exceed the cap. getProviderSemaphore now always hands out one shared limiter
(Infinity when unlimited, so every run is still counted) and resizes it in
place. Semaphore.release() decrements before admitting, and the new
Semaphore.resize() raises the ceiling by admitting queued waiters while
lowering it drains in-flight holders without admitting past the new cap.
Refs #3464
Semaphore.acquire now accepts an AbortSignal so a queued waiter that is cancelled (parent task abort, wall-clock budget elapsing) removes itself from the wait queue instead of being resolved by the next release. The provider semaphore in runSubprocess passes the run's abortSignal through, preventing aborted ollama-cloud subagents from permanently draining the provider concurrency budget.
Fixes#3464
Normalize the configured semaphore max before deciding whether the spawn
limit is bounded. Fractional values between 0 and 1 now truncate to 0 and
fall through to the unbounded path instead of storing 0 and deadlocking
the first acquire.
Fixes#3305
The session-scoped spawn Semaphore clamped its max via Math.max(1, max), so
task.maxConcurrency: 0 — labeled 'Unlimited' in the settings UI — serialized
subagent spawns one at a time instead of releasing every eligible seat.
The constructor now treats max <= 0 (and any non-finite input) as unbounded
via Number.POSITIVE_INFINITY, so the existing 'current < max' check naturally
permits every acquire. Mirrors the eval parallel()/pipeline() worker-pool
semantics (runEvalConcurrency in eval/concurrency-bridge.ts), which already
treats 0 as 'run every item at once'.
Fixes#3305
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:
1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.
2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.
Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.
Fixes#3196
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.
Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
forwards isolated/apply/merge (as booleans) plus returnHandle, and the
return_handle node carries isolated/patch_path/branch_name/
nested_patches/changes_applied/isolation_summary.
Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
the stale "defaults track task.isolation.mode" wording to the final
strict opt-in behavior, and note all four runtimes.
Fixes#3196
git stash pop without --index restores stashed staged changes as unstaged. When a nested repo had staged WIP before the isolated agent ran, the pop in applyNestedPatches() brought the content back but lost the user's index state.
Pass { index: true } so pop uses --index, matching the root merge path that already does the same thing.
Added a regression test that stages a pre-existing edit in the nested repo, runs applyNestedPatches, and asserts the file is still in the index (porcelain "M " with the trailing space) and the cached diff still shows the staged WIP.
Fixes#3196
applyNestedPatches() applied the captured patch then ran git.stage.files(nestedDir), which stages every working-tree change in the nested repo. A nested repo that was already dirty before the agent ran ended up with the user's unrelated work-in-progress committed alongside the agent delta.
Stash any pre-existing dirty state (tracked + untracked) before applying the patch and pop it back in the finally block after the commit, so the agent commit contains only the captured patch and the user's in-flight work is restored on top of it. A failing stash pop logs a warning and leaves the stash entry intact for manual recovery; the broader nested-apply failure path is already non-fatal.
Added a worktree integration test that confirms a pre-existing untracked file in the nested repo is not staged into the agent commit and is still present in the working tree afterwards.
Fixes#3196
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.
Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.
Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.
Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.
Fixes#3196
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
Eval preludes now forward returnHandle to the bridge so no-session eval runs can preserve the temp artifacts backing returned agent:// handles. The bridge keeps those temporary artifact directories whenever returnHandle is requested, including non-isolated runs and successful isolated applies.
Branch-mode isolation now treats nested-only changes as merge-eligible even when no root branch was produced, letting callers apply nested patches instead of dropping them when the root repo had no diff.
Added regression coverage for returnHandle artifact preservation and nested-only branch isolation.
Fixes#3196
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.
Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.
Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.
Fixes#3196
- Add `onFirstChatDispatch` hook to `CreateAgentSessionOptions` to track the boundary between session creation and the initial model request.
- Update `runSubprocess` and `TaskTool` to measure and log detailed latency metrics across the subagent lifecycle, including semaphore queue wait, setup time, and dispatch latency.
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
- Hardened the `hashline` parameters parsing pipeline to guarantee presence of the `input` field.
- Enforced input length limits on task roles to secure against oversized payloads.
- Configured Arktype schemas to reject or delete extra, undeclared fields in task and inspect-image payloads.
- Updated mock test parameters to align with corrected success exit codes.
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
- Added explicit ArkType schema descriptions across all coding agent tool definitions.
- Updated schema definitions in autoresearch and commit tools with descriptive wrappers.
- Documented tool schema enhancements in the packages/coding-agent CHANGELOG.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.