- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
- Extracted taskToolRenderer to a dedicated renderer file to resolve circular dependencies.
- Updated all references to the renderer to point to the new location.
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
Normalize the configured semaphore max before deciding whether the spawn
limit is bounded. Fractional values between 0 and 1 now truncate to 0 and
fall through to the unbounded path instead of storing 0 and deadlocking
the first acquire.
Fixes#3305
The session-scoped spawn Semaphore clamped its max via Math.max(1, max), so
task.maxConcurrency: 0 — labeled 'Unlimited' in the settings UI — serialized
subagent spawns one at a time instead of releasing every eligible seat.
The constructor now treats max <= 0 (and any non-finite input) as unbounded
via Number.POSITIVE_INFINITY, so the existing 'current < max' check naturally
permits every acquire. Mirrors the eval parallel()/pipeline() worker-pool
semantics (runEvalConcurrency in eval/concurrency-bridge.ts), which already
treats 0 as 'run every item at once'.
Fixes#3305
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:
1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.
2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.
Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.
Fixes#3196
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.
Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
forwards isolated/apply/merge (as booleans) plus returnHandle, and the
return_handle node carries isolated/patch_path/branch_name/
nested_patches/changes_applied/isolation_summary.
Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
the stale "defaults track task.isolation.mode" wording to the final
strict opt-in behavior, and note all four runtimes.
Fixes#3196
git stash pop without --index restores stashed staged changes as unstaged. When a nested repo had staged WIP before the isolated agent ran, the pop in applyNestedPatches() brought the content back but lost the user's index state.
Pass { index: true } so pop uses --index, matching the root merge path that already does the same thing.
Added a regression test that stages a pre-existing edit in the nested repo, runs applyNestedPatches, and asserts the file is still in the index (porcelain "M " with the trailing space) and the cached diff still shows the staged WIP.
Fixes#3196
applyNestedPatches() applied the captured patch then ran git.stage.files(nestedDir), which stages every working-tree change in the nested repo. A nested repo that was already dirty before the agent ran ended up with the user's unrelated work-in-progress committed alongside the agent delta.
Stash any pre-existing dirty state (tracked + untracked) before applying the patch and pop it back in the finally block after the commit, so the agent commit contains only the captured patch and the user's in-flight work is restored on top of it. A failing stash pop logs a warning and leaves the stash entry intact for manual recovery; the broader nested-apply failure path is already non-fatal.
Added a worktree integration test that confirms a pre-existing untracked file in the nested repo is not staged into the agent commit and is still present in the working tree afterwards.
Fixes#3196
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.
Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.
Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.
Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.
Fixes#3196
Eval preludes now forward returnHandle to the bridge so no-session eval runs can preserve the temp artifacts backing returned agent:// handles. The bridge keeps those temporary artifact directories whenever returnHandle is requested, including non-isolated runs and successful isolated applies.
Branch-mode isolation now treats nested-only changes as merge-eligible even when no root branch was produced, letting callers apply nested patches instead of dropping them when the root repo had no diff.
Added regression coverage for returnHandle artifact preservation and nested-only branch isolation.
Fixes#3196
- Moved authentication logic to `perplexity-auth.ts` to share logic between search providers and CLI commands.
- Updated authentication priority to prefer browser cookies over OAuth tokens during search operations.
- Modified the `token` CLI command to display active OAuth tokens when both an OAuth token and an API key are configured.
- Added comprehensive unit tests in `perplexity.test.ts` to verify authentication priority and precedence.
- Replaced usage of `ReturnType<typeof setTimeout>` and `ReturnType<typeof setInterval>` with the explicit `Timer` type across the codebase.
- Updated several type definitions and function signatures to use concrete types instead of inferred return types for improved clarity and maintainability.
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
- Switch all UI components and tests from sharp box corners (`boxSharp`) to rounded ones (`boxRound`).
- Update `Theme` to re-export sharp junction symbols (tees and cross) under `boxRound` to ensure consistent divider rendering in rounded boxes.
- Remove outdated architectural notes regarding forced tool-choice queues in documentation.
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
- Hardened the `hashline` parameters parsing pipeline to guarantee presence of the `input` field.
- Enforced input length limits on task roles to secure against oversized payloads.
- Configured Arktype schemas to reject or delete extra, undeclared fields in task and inspect-image payloads.
- Updated mock test parameters to align with corrected success exit codes.
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
- Forwarded `parentAgentId` through task and eval launch paths when spawning subagents.
- Mapped `parentAgentId` to `parentId` in `createAgentSession`.
- Passed each caller's session `getAgentId` (or `MAIN_AGENT_ID`) as the parent for spawned agents.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
Persisted isolated subagents created their fresh JSONL session through
SessionManager.open(), which fell back to getProjectDir() when the file had no
header. createAgentSession still received the isolated worktree cwd, but built-in
tools resolve paths through sessionManager.getCwd(), so file tools could target
the parent repository while patch capture saw no isolated delta.
Allow SessionManager.open() to take an initial cwd for empty/missing session
files and pass the isolated worktree cwd from task execution. Non-empty resumes
still use the persisted header cwd. Add regression coverage asserting persisted
isolated subagent sessions expose the worktree cwd through their session
manager.
- Added a setup-system-deps action with preloaded-runner guards and apt fallbacks.
- Updated CI workflows to download Linux x64 native artifacts and gate on native job success.
- Renamed coding-agent fast mode to singleton in scripts and test partitioning logic.
- Added settings test-state begin/restore helpers with recursive cleanup in affected tests.
- Added cycle and depth guards for nested task progress rendering so async fan-out snapshots cannot recurse until the TUI crashes.
- Shortened long Windows '/data/workspaces/can1357__oh-my-pi__2551/.omp-session/2026-06-14T07-09-37-753Z_019ec4f6-ee59-7000-8226-e1b7ed0680e9/local' roots into temp-backed session roots before plan/handoff writes hit MAX_PATH.
Fixes#2551
When one task call spawns two or more live siblings with spawn capacity
and IRC enabled, TaskTool.execute appends a coordinate-via-irc
suggestion, composed onto the specialization advisory through the same
seam. Tighten the subagent COOP section and irc tool prompt so guidance
spans discovery (list who/what), coordination (message before
overlapping edits), and follow-up (replyTo/await) instead of only
assuming agents resolve collisions on their own.
Refs #2471
Op: extend
When a spawner with remaining depth capacity spawns generic role-less
workers (a task/quick_task spawn without a `role`, or the same agent
cloned >=2x all without roles), TaskTool.execute appends a non-blocking
advisory steering it toward tailored specialists. Gated on DepthCapacity
so a leaf at max recursion is never nudged; the task-tool depth gate is
extracted into a shared `canSpawnAtDepth` helper reused by both the tool
gate and the advisory.
Refs #2469
Op: extend
Document the `role` parameter in the task-tool description (both the
batch and single-spawn shapes) and make tailored specialists the default
rule, not the exception. Direct a recursing worker to pass a `role` for
each sub-specialist. Activates the role field from #2467 for the model.
Refs #2468
Op: extend
Add an optional `role` field to the task spawn contract, threaded end to
end through resolveSpawnItems/spawnParamsFor into the executor. A role
injects a specialization preamble into the subagent system prompt and
becomes the subagent's display name and telemetry identity (label
normalized, length capped), so delegated trees stop being clones of one
generic worker. Empty/absent roles fall back to the agent type name.
Refs #2467
Op: extend
- Changed collapsed progress rendering to keep the most recent live agents visible, adding a summary line for folded-away rows.
- Updated collapsed result rendering to preserve failed and aborted agents in the visible set while trimming other completions.
- Refreshed job polling text to document waiting on all running jobs when `poll` is omitted and added tests for both collapsed progress and result display behavior.
- Updated task call and result headers to use the dispatch glyph during running async calls instead of spinner-style states.
- Switched running and pending agent rows to a static done-dot marker and reused the dot with foreground color settling when rows complete.
- Updated task rendering tests to validate the new glyphs and ensure running/pending rows do not emit spinner or pending symbols.
- Detached async task progress rows stopped running a redraw driver and task progress rendering switched running/pending rows to static task-icon text.
- Background task snapshots were frozen once blocks left the transcript live region, preventing later partial snapshots from repainting commit-eligible rows.
- Updated task-progress and detached-background-task tests to validate static task rows and the new freeze behavior.
Reorders sections in the streaming call preview to match `renderResult` and the
schema's field order. This prevents visual jumps when the preview transitions
to a result and ensures append-only growth of streamed content.
Additionally, omits the agent-list divider when no agent rows are present.
- Updated TaskTool to skip `session.asyncJobManager` and run `task` spawns inline whenever `async.enabled` is false.
- Set `async.enabled` default to `true` and updated task prompts/settings text to reflect async-versus-sync behavior.
- Adjusted task batching tests to cover both async background execution and synchronous batched execution when async is disabled.
- Replaced task-simple-mode with a `task.batch` setting enabled by default.
- Updated task schema to use batch `{agent, context, tasks[]}` payloads.
- Migrated task execution to spawn one async job per task and merge outputs.
- Removed per-call schema passing while preserving legacy flat task calls.
- Removed `resume` from task params and schema, requiring agent and assignment inputs.
- Dropped resume continuation paths in task execution and call rendering, always spawning a new agent.
- Removed the `irc.enabled` setting and computed IRC availability by task-depth rules.
- Updated task follow-up guidance to use IRC messaging/history links instead of `task(resume:)`.
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
- Added a stable progress-ordering helper that moved pending and running agents below completed and failed ones.
- Applied this ordering to top-level and nested live task-progress rendering so finished entries render first.
- Added a renderer test and changelog note covering the finished-before-unfinished progress ordering.
Extension commands (e.g. /sonnet) and TypeScript custom commands that
consume the input without calling the LLM return early from
session.prompt() with no agent turn. In ACP mode this left the pending
prompt promise unresolved, hanging the client forever.
Change session.prompt() to return Promise<boolean>: true when the LLM
was invoked, false when the command was fully handled locally.
#runPromptOrCommand calls #finishPrompt immediately on a false return so
the ACP turn completes.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Fixed help rendering so `--help` no longer triggers unrelated command loaders.
- Fixed startup span logging to emit markers only with PI_DEBUG_STARTUP set.
- Fixed logger startup trace behavior for `:start`, `:done`, and `:fail` phases.
- Fixed prompt template processing with cached raw-template compilation and safer formatting.
- Optimized symbol and tag parsing in prompt templates via manual parsers.