Commit Graph
383 Commits
Author SHA1 Message Date
can1357 3f16615afa fix-task-recent-output-tail-preview 2026-07-01 21:53:20 +02:00
can1357 32f5820d66 Merge PR #4167: fix(tui): cap subagent live progress output (@roboomp) 2026-07-01 21:53:20 +02:00
can1357 87a53cbe0d Merge PR #4137: fix(agent): handle already-applied patch-mode merges (@roboomp) 2026-07-01 21:53:18 +02:00
can1357 aee8b091ee fix(agent): honor closed schema label edge cases 2026-07-01 21:50:54 +02:00
can1357 88e3e77f3e Merge PR #3927: fix(agent): reject stale yield labels for override schemas (@roboomp) 2026-07-01 21:50:54 +02:00
roboomp 4e871a4b51 fix(tui): capped subagent live output
Reused output notice stripping for task live progress and rendered recent subagent output through the viewport-sized preview budget.

Added regression coverage for fixed six-line capping and raw bash footer leakage.

Fixes #4162
2026-07-01 17:20:52 +00:00
roboomp cd692bcb02 fix(agent): required forward-check failure for patch-mode no-op
Reverse-check alone can theoretically succeed via git-apply fuzz when the file carries the postimage at another location, so treating it as sufficient risked silently dropping a task patch whose intended hunk still needed to run.

Required both `--reverse --check` to succeed AND forward `--check` to fail before declaring a no-op. When both check directions succeed (ambiguous), the merge falls through to a forward apply instead of skipping — mirroring the pre-idempotence behavior. Both checks are read-only, so no worktree writes on ambiguity.

Added a spy-based regression that forces the ambiguous case and asserts the merge applies forward instead of silently dropping the change.
2026-07-01 12:17:42 +00:00
roboomp f474fa0e11 fix(agent): detected patch-mode idempotence via reverse-check
git apply --3way --check exits 0 even when the real apply would write conflict markers and unmerged index stages, so the previous fix left the worktree dirty on conflicting patches while only flipping changesApplied to false.

Dropped --3way for patch-mode merge and used a --reverse --check probe instead: it succeeds only when the target state is already present (true no-op) and reads without touching the worktree. Conflicts fall through to the normal --check + apply path, which rejects them before writing anything.

Added regression coverage for the conflict scenario asserting the worktree stays clean, and for the fresh apply path.
2026-07-01 12:06:06 +00:00
roboomp 66052863bb fix(agent): handled already-applied patch merges
Use git apply --3way for patch-mode isolated merge checks and applies so diff-tree patches that are already present are accepted as no-ops.

Added regression coverage for a clean already-applied binary/full-index patch.

Fixes #4135
2026-07-01 11:58:52 +00:00
roboomp 8614b4c086 fix(task): respected restricted spawn defaults
Resolved eval agent() and task tool defaults from the active spawn policy so restricted agents advertise and execute an allowed default.

Fixes #3973
2026-07-01 02:44:08 +00:00
roboomp 8e2dfcc43e fix(agent): routed acquire-time abort through queued-spawn settled path
When a batched task spawn is cancelled while still queued behind task.maxConcurrency the semaphore now rejects acquire(), but the previous patch let the abort throw past the aborted handler so progress.status and onSettled never fired and buildAsyncDetails kept reporting the batch as running. The wrapper now records whether the slot was held, funnels both acquire-time and post-acquire aborts through the same aborted branch (releasing only when held), and a batch regression test pins the contract.

Fixes #3930
2026-06-30 23:46:19 +00:00
roboomp 3bc8f995f6 fix(agent): bounded async job disposal
Made AsyncJobManager.dispose honor its timeout while waiting for cancelled jobs, and passed task abort signals into spawn semaphore waits.

Fixes #3930
2026-06-30 23:37:11 +00:00
roboomp 00ef58f843 fix(agent): steered override-schema subagents
- Marked eval agent schema calls as caller overrides so subagent prompts can revoke native output/yield instructions.\n- Added override-schema prompt guidance telling agents to ignore conflicting native output labels and terminal-yield the caller schema object.\n- Added prompt coverage for the override notice.\n\nRefs #3926
2026-06-30 22:30:15 +00:00
roboomp e4561d64fa fix(coding-agent): restored subagent thinking precedence
Agent frontmatter thinkingLevel now wins over model role suffix thinking when both are configured.

Fixes #3915
2026-06-30 20:15:04 +00:00
can1357 720fb3f120 feat(coding-agent): replaced the oracle subagent with a new tester subagent
- Removed the oracle agent prompt configuration and its references across the codebase.
- Added a new tester agent prompt markdown file with directives for test authoring and techniques.
- Registered the new tester agent in the task runner while removing the oracle agent.
- Renamed the quick-task agent configuration and prompt references to sonic.
2026-06-30 16:16:39 +02:00
can1357 9ccd83a13d feat(coding-agent): made the agent parameter optional with a default value
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
2026-06-30 16:16:39 +02:00
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
can1357 8cd23ebd15 merge #3844: 3-way dirty-context fallback for isolated branch merges
# Conflicts:
#	packages/coding-agent/src/task/worktree.ts
#	packages/coding-agent/test/task/worktree.test.ts
2026-06-30 03:03:43 +02:00
roboomp d120ba6b7d fix(coding-agent): filter baseline wip from preserved agent commits
Dirty isolated baselines can be accidentally committed by subagents that run git add -A. Fetching the raw isolation HEAD then cherry-picking the range would replay that baseline WIP into parent history.

Add a dirty-baseline replay path that rewrites each agent commit against the captured baseline tree, preserving the agent commit message and author while excluding staged, unstaged, and untracked changes that existed before isolation started. Clean baselines still use the raw git fetch path, and nested-only changes keep returning patches without creating an empty root branch.

Add a regression for baseline staged + untracked WIP committed by the agent, asserting the task branch contains only the agent file and parent WIP remains staged/untracked after merge.

Fixes #3842
2026-06-30 00:28:27 +00:00
roboomp da715aae7c fix(coding-agent): preserve agent commits across isolated branch merges
When an isolated task agent commits its own changes before yielding, the
harness used to collapse the captured delta into one AI-summarized commit
and discard the agent's commit messages and authorship entirely. This
violated commit discipline for agentic swarms — multiple logical commits
("fix bug" + "add test") became a single opaque commit, and the
agent's commit object (which lived in isolation/.git/objects under
overlayfs/rcopy) was lost when cleanupIsolation tore down the overlay.

commitToBranch now detects when isolation HEAD moved past baseline.root
.headCommit. When it has, the function git-fetches the agent's HEAD into
the parent repo as omp/task/${taskId} so the commit objects survive
cleanupIsolation, and stamps the captured baselineSha onto the returned
CommitToBranchResult. mergeTaskBranches cherry-picks the inclusive range
baseSha..branchName when baseSha is provided, replaying each agent
commit verbatim with its original message and author. Any uncommitted
leftover (staged, unstaged, untracked) on top of the agent's last commit
becomes one trailing AI-summarized commit on the same branch.

Falls back to the legacy single-commit path when the agent never moved
HEAD (purely dirty working tree); existing patch-mode flow is untouched.

Fixes #3842
2026-06-30 00:13:07 +00:00
roboomp 4b98211c64 fix(coding-agent): fixed dirty isolated branch merges
Applied isolated branch patches with three-way fallback when unrelated parent dirt appears in patch context.

Surfaced branch preparation failures instead of reporting no changes.

Fixes #3841
2026-06-30 00:09:34 +00:00
roboomp 58c0a305d6 fix(agent): shortened isolated task paths
Used compact hashed isolation directory segments and the short m mount dir so long task ids are not copied into subagent working paths.

Kept worktree cleanup compatible with legacy merged task-isolation directories.

Fixes #3756
2026-06-28 21:25:20 +00:00
roboompandcan1357 2042d3b114 fix(task): scoped provider concurrency cap to each LLM turn
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.

Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).

Fixes #3749
2026-06-28 22:50:59 +02:00
can1357 a8dc036c84 feat(coding-agent): corrected payload assembly logic for terminal results
- Refactored payload assembly to distinguish between incremental sections and terminal results.
- Prevented terminal markers from being incorrectly treated as section labels to avoid payload nesting issues.
- Updated section processing to ignore non-incremental terminal items, resolving incorrectly missing data in output-schema validation.
- Improved terminal item resolution to correctly fallback to the last assistant text when no explicit data is provided.
2026-06-28 22:50:31 +02:00
can1357 93f68b75a2 test(coding-agent/modes): updated controller test mocks
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
2026-06-28 16:53:16 +02:00
can1357 51a2a0342f test(coding-agent): implemented guest reconciliation and expanded testing for collaboration
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
2026-06-28 09:52:44 +02:00
can1357 2cf28f972e feat(coding-agent): added logic to assembleYieldResult to
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
2026-06-28 09:05:37 +02:00
can1357 102d6d54ad feat: implemented v2 streaming remote compaction for model history state
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
2026-06-28 07:27:02 +02:00
can1357 d8bf177af7 refactor(coding-agent/task): moved taskToolRenderer to separate module
- Extracted taskToolRenderer to a dedicated renderer file to resolve circular dependencies.
- Updated all references to the renderer to point to the new location.
2026-06-28 07:27:02 +02:00
can1357 4fb1490747 feat(coding-agent/task): updated assistant review prompt and export renderer
- Updated the reviewer prompt instructions for clarity.
- Reordered export statements to address potential circular dependency issues during tool registration.
2026-06-28 07:27:02 +02:00
can1357 289dd770c8 feat(coding-agent): reworked subagent yields for incremental results
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
2026-06-28 07:27:01 +02:00
can1357 5a044dc0da feat: consolidated and automate changelog management
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
2026-06-27 03:08:52 +02:00
can1357 3345085290 feat(coding-agent/task): improved task description previews
- Updated task rendering to use `previewLine` for truncating descriptions consistently.
- Ensured progress and result descriptions are truncated before formatting.
2026-06-27 02:55:52 +02:00
can1357 577d2a8eb8 style: biome format/organize-imports across integrated PRs 2026-06-27 02:06:38 +02:00
can1357 29269d9547 Fix eval agent isolation artifact recovery 2026-06-27 01:39:35 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 2c20a2d368 Fix subagent yield abort cleanup
(cherry picked from commit e22673537f7cc95920cd4ec75d2a4b36d81d7626)
2026-06-26 23:43:03 +02:00
can1357 84ef014faa Merge PR #3407: fix(eval): dispose one-shot eval subagents after run (@korri123) 2026-06-26 23:27:40 +02:00
can1357 938489f3fd feat(coding-agent): added configurable service tier settings for subagents and advisor
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
2026-06-25 22:35:15 +02:00
can1357 06ca03dbd8 fix(agent): release task spawn slot when progress reporting throws
In #registerSpawnJob the markRunning()/reportProgress() calls sat between
semaphore.acquire() and the try whose finally releases the slot. If progress
reporting threw there, the acquired task.maxConcurrency slot leaked and
permanently shrank subagent concurrency. Move those statements inside the try
so finally always releases. The abort-before-execution branch is unchanged
(releases once and throws before the try is entered — not a double release).

Refs #3464
2026-06-25 20:48:38 +02:00
can1357 ecd6f608e5 fix(agent): resize provider concurrency limiter in place instead of replacing it
The per-provider subagent limiter (providers.ollama-cloud.maxConcurrency)
created a fresh Semaphore whenever the configured limit changed, orphaning
in-flight slots on the old instance so a runtime or mixed limit value could
exceed the cap. getProviderSemaphore now always hands out one shared limiter
(Infinity when unlimited, so every run is still counted) and resizes it in
place. Semaphore.release() decrements before admitting, and the new
Semaphore.resize() raises the ceiling by admitting queued waiters while
lowering it drains in-flight holders without admitting past the new cap.

Refs #3464
2026-06-25 20:48:38 +02:00
roboomp 7e90e4d081 fix(agent): released ollama-cloud semaphore slot when waiter aborts
Semaphore.acquire now accepts an AbortSignal so a queued waiter that is cancelled (parent task abort, wall-clock budget elapsing) removes itself from the wait queue instead of being resolved by the next release. The provider semaphore in runSubprocess passes the run's abortSignal through, preventing aborted ollama-cloud subagents from permanently draining the provider concurrency budget.

Fixes #3464
2026-06-25 12:14:45 +00:00
roboomp 80862b79da fix(agent): handled ollama-cloud task backoff
Added ollama-cloud subagent concurrency limiting, role fallback-chain inheritance, and visible empty length errors for native Ollama responses.

Fixes #3464
2026-06-25 11:57:51 +00:00
Kormákur 0a7ec389a0 fix(coding-agent): verify eval subagent cleanup 2026-06-25 09:27:38 +00:00
Kormákur 9118e2cb95 fix(coding-agent): dispose eval subagents after run 2026-06-24 22:15:03 +00:00
roboomp e4c52de24d fix(task): normalize fractional spawn concurrency
Normalize the configured semaphore max before deciding whether the spawn

limit is bounded. Fractional values between 0 and 1 now truncate to 0 and

fall through to the unbounded path instead of storing 0 and deadlocking

the first acquire.

Fixes #3305
2026-06-23 10:41:51 +00:00
roboomp 296125cce9 fix(task): treat maxConcurrency 0 as unbounded in spawn semaphore
The session-scoped spawn Semaphore clamped its max via Math.max(1, max), so

task.maxConcurrency: 0 — labeled 'Unlimited' in the settings UI — serialized

subagent spawns one at a time instead of releasing every eligible seat.

The constructor now treats max <= 0 (and any non-finite input) as unbounded

via Number.POSITIVE_INFINITY, so the existing 'current < max' check naturally

permits every acquire. Mirrors the eval parallel()/pipeline() worker-pool

semantics (runEvalConcurrency in eval/concurrency-bridge.ts), which already

treats 0 as 'run every item at once'.

Fixes #3305
2026-06-23 10:37:34 +00:00
roboompandcan1357 5b6e9f904d fix(eval): paused timeout over baseline capture and surfaced nested stash-restore failures
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:

1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.

2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.

Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.

Fixes #3196
2026-06-22 21:57:35 +02:00
can1357 2f2acaf928 Merge pull request #3205: feat(eval): isolated/apply/merge options for agent() helper
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.

Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
  forwards isolated/apply/merge (as booleans) plus returnHandle, and the
  return_handle node carries isolated/patch_path/branch_name/
  nested_patches/changes_applied/isolation_summary.

Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
  task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
  the stale "defaults track task.isolation.mode" wording to the final
  strict opt-in behavior, and note all four runtimes.

Fixes #3196
2026-06-22 21:56:45 +02:00
roboomp abc9a8f292 fix(task): restored nested stash with index state
git stash pop without --index restores stashed staged changes as unstaged. When a nested repo had staged WIP before the isolated agent ran, the pop in applyNestedPatches() brought the content back but lost the user's index state.

Pass { index: true } so pop uses --index, matching the root merge path that already does the same thing.

Added a regression test that stages a pre-existing edit in the nested repo, runs applyNestedPatches, and asserts the file is still in the index (porcelain "M  " with the trailing space) and the cached diff still shows the staged WIP.

Fixes #3196
2026-06-22 19:41:59 +00:00