The per-provider subagent limiter (providers.ollama-cloud.maxConcurrency)
created a fresh Semaphore whenever the configured limit changed, orphaning
in-flight slots on the old instance so a runtime or mixed limit value could
exceed the cap. getProviderSemaphore now always hands out one shared limiter
(Infinity when unlimited, so every run is still counted) and resizes it in
place. Semaphore.release() decrements before admitting, and the new
Semaphore.resize() raises the ceiling by admitting queued waiters while
lowering it drains in-flight holders without admitting past the new cap.
Refs #3464
Semaphore.acquire now accepts an AbortSignal so a queued waiter that is cancelled (parent task abort, wall-clock budget elapsing) removes itself from the wait queue instead of being resolved by the next release. The provider semaphore in runSubprocess passes the run's abortSignal through, preventing aborted ollama-cloud subagents from permanently draining the provider concurrency budget.
Fixes#3464
Normalize the configured semaphore max before deciding whether the spawn
limit is bounded. Fractional values between 0 and 1 now truncate to 0 and
fall through to the unbounded path instead of storing 0 and deadlocking
the first acquire.
Fixes#3305
The session-scoped spawn Semaphore clamped its max via Math.max(1, max), so
task.maxConcurrency: 0 — labeled 'Unlimited' in the settings UI — serialized
subagent spawns one at a time instead of releasing every eligible seat.
The constructor now treats max <= 0 (and any non-finite input) as unbounded
via Number.POSITIVE_INFINITY, so the existing 'current < max' check naturally
permits every acquire. Mirrors the eval parallel()/pipeline() worker-pool
semantics (runEvalConcurrency in eval/concurrency-bridge.ts), which already
treats 0 as 'run every item at once'.
Fixes#3305
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:
1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.
2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.
Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.
Fixes#3196
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.
Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
forwards isolated/apply/merge (as booleans) plus returnHandle, and the
return_handle node carries isolated/patch_path/branch_name/
nested_patches/changes_applied/isolation_summary.
Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
the stale "defaults track task.isolation.mode" wording to the final
strict opt-in behavior, and note all four runtimes.
Fixes#3196
git stash pop without --index restores stashed staged changes as unstaged. When a nested repo had staged WIP before the isolated agent ran, the pop in applyNestedPatches() brought the content back but lost the user's index state.
Pass { index: true } so pop uses --index, matching the root merge path that already does the same thing.
Added a regression test that stages a pre-existing edit in the nested repo, runs applyNestedPatches, and asserts the file is still in the index (porcelain "M " with the trailing space) and the cached diff still shows the staged WIP.
Fixes#3196
applyNestedPatches() applied the captured patch then ran git.stage.files(nestedDir), which stages every working-tree change in the nested repo. A nested repo that was already dirty before the agent ran ended up with the user's unrelated work-in-progress committed alongside the agent delta.
Stash any pre-existing dirty state (tracked + untracked) before applying the patch and pop it back in the finally block after the commit, so the agent commit contains only the captured patch and the user's in-flight work is restored on top of it. A failing stash pop logs a warning and leaves the stash entry intact for manual recovery; the broader nested-apply failure path is already non-fatal.
Added a worktree integration test that confirms a pre-existing untracked file in the nested repo is not staged into the agent commit and is still present in the working tree afterwards.
Fixes#3196
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.
Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.
Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.
Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.
Fixes#3196
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
Eval preludes now forward returnHandle to the bridge so no-session eval runs can preserve the temp artifacts backing returned agent:// handles. The bridge keeps those temporary artifact directories whenever returnHandle is requested, including non-isolated runs and successful isolated applies.
Branch-mode isolation now treats nested-only changes as merge-eligible even when no root branch was produced, letting callers apply nested patches instead of dropping them when the root repo had no diff.
Added regression coverage for returnHandle artifact preservation and nested-only branch isolation.
Fixes#3196
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.
Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.
Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.
Fixes#3196
- Add `onFirstChatDispatch` hook to `CreateAgentSessionOptions` to track the boundary between session creation and the initial model request.
- Update `runSubprocess` and `TaskTool` to measure and log detailed latency metrics across the subagent lifecycle, including semaphore queue wait, setup time, and dispatch latency.
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
- Hardened the `hashline` parameters parsing pipeline to guarantee presence of the `input` field.
- Enforced input length limits on task roles to secure against oversized payloads.
- Configured Arktype schemas to reject or delete extra, undeclared fields in task and inspect-image payloads.
- Updated mock test parameters to align with corrected success exit codes.
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
- Added explicit ArkType schema descriptions across all coding agent tool definitions.
- Updated schema definitions in autoresearch and commit tools with descriptive wrappers.
- Documented tool schema enhancements in the packages/coding-agent CHANGELOG.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
- Fixed cold revival flow so parked subagents are restored from persisted sessions at startup.
- Fixed session-init persistence to include spawns and readSummarize fields for replay accuracy.
- Fixed latest-session lookup by adding peekSessionInit for lock-free persisted contract access.
- Added lifecycle and session tests for cold-revive success, decline, and retry paths.
- Forwarded `parentAgentId` through task and eval launch paths when spawning subagents.
- Mapped `parentAgentId` to `parentId` in `createAgentSession`.
- Passed each caller's session `getAgentId` (or `MAIN_AGENT_ID`) as the parent for spawned agents.
- Removed render_mermaid from tool discovery, task definitions, and registries.
- Removed renderMermaid setting and prompt/docs references tied to the deleted tool.
- Added maxWidth and theme color options to Mermaid ASCII resolution in markdown flow.
- Re-rendered Mermaid ASCII in both directions and clipped output to available width.
- Added a `suppressBreadcrumb` option to `SessionManager.open()` so headless opens skip writing the per-TTY `--continue` breadcrumb, and passed it from the subagent opens in `task/executor.ts` and the HTML export open in `export/html/index.ts`, which run in the parent's terminal and were clobbering the breadcrumb with their own artifact-dir session file.
- Added `resolveBreadcrumbToInteractiveRoot()` and applied it in `continueRecent()` so already-poisoned breadcrumbs pointing inside a parent's artifacts dir (`<parent>/<agentId>.jsonl`) resolve back up to the top-level interactive session.
- Added `subagent-breadcrumb.test.ts` covering both that a subagent open keeps `--continue` on the parent and that a stale subagent-pointing breadcrumb is recovered.
Kept OpenRouter and Vercel upstream routing suffixes in subagent retry fallback selectors so same-base routed candidates stay distinct.
Resolved retry fallback candidates from the raw selector before model switching so routed fallback models keep their requested upstream route.
Placed subagent-scoped fallback chains ahead of inherited retry chains so overlapping primary selectors prefer the explicit subagent model order.
Expanded the regression to pin fallback chain insertion order when a global chain shares the same primary selector.
Installed subagent-scoped retry fallback chains from ordered task model candidates so provider failures can advance to the next configured worker model.
Updated task result tracking to surface the fallback-applied final model and added focused regression coverage for the executor wiring.
Fixes#2750
- Passed USER_INTERRUPT_LABEL through abort paths in collab, ACP, RPC, runtime, and SDK flows.
- Added userInitiated to synthetic continue inputs and session prompt calls.
- Suppressed advisor auto-resume during user aborts and preserved queued concerns.
- Cleared suppression on user prompts and reclaimed parked advisor cards on abort settle.
Persisted isolated subagents created their fresh JSONL session through
SessionManager.open(), which fell back to getProjectDir() when the file had no
header. createAgentSession still received the isolated worktree cwd, but built-in
tools resolve paths through sessionManager.getCwd(), so file tools could target
the parent repository while patch capture saw no isolated delta.
Allow SessionManager.open() to take an initial cwd for empty/missing session
files and pass the isolated worktree cwd from task execution. Non-empty resumes
still use the persisted header cwd. Add regression coverage asserting persisted
isolated subagent sessions expose the worktree cwd through their session
manager.
- Added cycle and depth guards for nested task progress rendering so async fan-out snapshots cannot recurse until the TUI crashes.
- Shortened long Windows '/data/workspaces/can1357__oh-my-pi__2551/.omp-session/2026-06-14T07-09-37-753Z_019ec4f6-ee59-7000-8226-e1b7ed0680e9/local' roots into temp-backed session roots before plan/handoff writes hit MAX_PATH.
Fixes#2551
- Removed the `task` embedded agent's explicit `thinkingLevel` assignment.
- Updated the `quick_task` agent to use `Effort.Medium` instead of `Effort.Minimal`.
Two controller bugs from review:
- The post-stop nudge still queued a passive `nextTurn` message during goal
mode (goal mode only disabled `autoContinue`). That message rides the goal
continuation and can divert the goal loop into capture. Return early when
goal mode is active.
- `#suppressNext` was latched before the fire-and-forget `sendCustomMessage`.
It must arm synchronously (the synthetic turn's `agent_end` fires inside
`sendCustomMessage` before it resolves), but a rejected or *deferred* dispatch
(ACP clients downgrade `triggerTurn` to a queue) then produces no `agent_end`,
so the latch swallowed the next real stop. `sendCustomMessage` now returns
whether it actually started a turn; the controller disarms the latch when no
turn ran (rejection or deferral).
`task/executor.ts` widens its pending-message array to `Promise<unknown>[]` to
absorb the new return type (the resolved values are discarded).
Addresses review threads on PR #2542 (threads 5, 9, 12).
When one task call spawns two or more live siblings with spawn capacity
and IRC enabled, TaskTool.execute appends a coordinate-via-irc
suggestion, composed onto the specialization advisory through the same
seam. Tighten the subagent COOP section and irc tool prompt so guidance
spans discovery (list who/what), coordination (message before
overlapping edits), and follow-up (replyTo/await) instead of only
assuming agents resolve collisions on their own.
Refs #2471
Op: extend
Add a display-only `activity` field to AgentRef plus `setActivity`, fed
from the subagent progress chokepoint with a short gist of the agent's
latest intent (or current tool). Render it in the `irc list` output, the
subagent peer roster, and the TUI peer card, beside the role-derived
display name. setActivity emits no event — the roster reads on demand —
so the per-tool-call rate stays off the registry listener path. Peers
with no activity render without a dangling clause.
Refs #2470
Op: extend
When a spawner with remaining depth capacity spawns generic role-less
workers (a task/quick_task spawn without a `role`, or the same agent
cloned >=2x all without roles), TaskTool.execute appends a non-blocking
advisory steering it toward tailored specialists. Gated on DepthCapacity
so a leaf at max recursion is never nudged; the task-tool depth gate is
extracted into a shared `canSpawnAtDepth` helper reused by both the tool
gate and the advisory.
Refs #2469
Op: extend
Add an optional `role` field to the task spawn contract, threaded end to
end through resolveSpawnItems/spawnParamsFor into the executor. A role
injects a specialization preamble into the subagent system prompt and
becomes the subagent's display name and telemetry identity (label
normalized, length capped), so delegated trees stop being clones of one
generic worker. Empty/absent roles fall back to the agent type name.
Refs #2467
Op: extend
Always-on LoopWatchdog (armed in TUI.start/stop) logs ui.loop-blocked with blockedMs and the current loop phase on the rising edge of a late probe tick. New pushLoopPhase/popLoopPhase/currentLoopPhase stack in pi-utils feeds it; breadcrumbs at in-process subagent dispatch (subagent:<id>) and the SelectList fuzzy filter (ui.select-filter) attribute residual main-thread stalls.
- Added detached spawn metadata to lifecycle, progress, and executor session payloads so task and evaluator runs can mark background jobs.
- Updated subagent session tracking and HUD rendering to only show active detached spawns.
- Extended HUD tests to verify non-detached sync and eval spawns are excluded while detached flags propagate through events.
- Changed collapsed progress rendering to keep the most recent live agents visible, adding a summary line for folded-away rows.
- Updated collapsed result rendering to preserve failed and aborted agents in the visible set while trimming other completions.
- Refreshed job polling text to document waiting on all running jobs when `poll` is omitted and added tests for both collapsed progress and result display behavior.