Commit Graph

543 Commits

Author SHA1 Message Date
can1357 32f5820d66 Merge PR #4167: fix(tui): cap subagent live progress output (@roboomp) 2026-07-01 21:53:20 +02:00
can1357 87a53cbe0d Merge PR #4137: fix(agent): handle already-applied patch-mode merges (@roboomp) 2026-07-01 21:53:18 +02:00
can1357 aee8b091ee fix(agent): honor closed schema label edge cases 2026-07-01 21:50:54 +02:00
can1357 88e3e77f3e Merge PR #3927: fix(agent): reject stale yield labels for override schemas (@roboomp) 2026-07-01 21:50:54 +02:00
roboomp fb2804247a fix(tui): gated live tool spinner ticks
Added renderer metadata for pending and partial result paths that visibly consume spinner frames.

Stopped headerless bash pending previews from scheduling repaint ticks while preserving eval and shell header animation.
2026-07-01 18:34:12 +00:00
roboomp 4e871a4b51 fix(tui): capped subagent live output
Reused output notice stripping for task live progress and rendered recent subagent output through the viewport-sized preview budget.

Added regression coverage for fixed six-line capping and raw bash footer leakage.

Fixes #4162
2026-07-01 17:20:52 +00:00
roboomp 87e1b03db1 style: bun run fix 2026-07-01 12:49:36 +00:00
roboomp ecbf0760cf fix(coding-agent): guarded leftover WIP seed by filtered-commit landings
Tracking raw agent-commit count meant an agent that committed only its
inherited baseline WIP (via `git add -A`) then left the real edit
uncommitted collapsed every filtered patch to empty, so tmpDir never
advanced past baselineSha yet the leftover path assumed it had.

replayFilteredAgentCommits now counts filtered commits actually applied
and, when zero landed, bypasses the finalFilteredTree/leftoverPatch
synthesis entirely — writeSyntheticTree(HEAD, [rootPatch]) fails hard
whenever rootPatch has HEAD+WIP context for a file missing from HEAD's
index (untracked / staged-new WIP). Instead, commit rootPatch directly
with WIP seed, matching the no-agent-commit path.

Added regression test covering the reviewer's scenario.
2026-07-01 12:49:28 +00:00
roboomp 286e971bfe fix(coding-agent): stopped isolated task merges failing when working tree carries WIP for files the agent also modifies
captureRepoDeltaPatch records the delta against `HEAD + WIP`, so the
patch's context lines and blob SHAs reference the WIP-modified files.
commitPatchToBranchWorktree then tried to apply that patch to a fresh
worktree pinned at HEAD, which failed hard whenever the WIP-side file
was missing from HEAD's index (untracked WIP files, staged-new WIP
files) or when --3way could not resolve an overlap.

commitPatchToBranchWorktree now tries plain apply first, then `--3way`
(which cleanly subtracts WIP via the shared ODB blob for tracked files),
and only when both fail replays baseline WIP into the temp worktree so
the delta's context lines up, rewinding WIP-only files afterward so
they never leak into the branch commit.

Added git.ls.tree helper for the WIP-only-file filter and a set of
regression tests covering the untracked, staged-new, and overlap
scenarios.

Fixes #4136
2026-07-01 12:24:32 +00:00
roboomp cd692bcb02 fix(agent): required forward-check failure for patch-mode no-op
Reverse-check alone can theoretically succeed via git-apply fuzz when the file carries the postimage at another location, so treating it as sufficient risked silently dropping a task patch whose intended hunk still needed to run.

Required both `--reverse --check` to succeed AND forward `--check` to fail before declaring a no-op. When both check directions succeed (ambiguous), the merge falls through to a forward apply instead of skipping — mirroring the pre-idempotence behavior. Both checks are read-only, so no worktree writes on ambiguity.

Added a spy-based regression that forces the ambiguous case and asserts the merge applies forward instead of silently dropping the change.
2026-07-01 12:17:42 +00:00
roboomp f474fa0e11 fix(agent): detected patch-mode idempotence via reverse-check
git apply --3way --check exits 0 even when the real apply would write conflict markers and unmerged index stages, so the previous fix left the worktree dirty on conflicting patches while only flipping changesApplied to false.

Dropped --3way for patch-mode merge and used a --reverse --check probe instead: it succeeds only when the target state is already present (true no-op) and reads without touching the worktree. Conflicts fall through to the normal --check + apply path, which rejects them before writing anything.

Added regression coverage for the conflict scenario asserting the worktree stays clean, and for the fresh apply path.
2026-07-01 12:06:06 +00:00
roboomp 66052863bb fix(agent): handled already-applied patch merges
Use git apply --3way for patch-mode isolated merge checks and applies so diff-tree patches that are already present are accepted as no-ops.

Added regression coverage for a clean already-applied binary/full-index patch.

Fixes #4135
2026-07-01 11:58:52 +00:00
roboomp 8614b4c086 fix(task): respected restricted spawn defaults
Resolved eval agent() and task tool defaults from the active spawn policy so restricted agents advertise and execute an allowed default.

Fixes #3973
2026-07-01 02:44:08 +00:00
roboomp 8e2dfcc43e fix(agent): routed acquire-time abort through queued-spawn settled path
When a batched task spawn is cancelled while still queued behind task.maxConcurrency the semaphore now rejects acquire(), but the previous patch let the abort throw past the aborted handler so progress.status and onSettled never fired and buildAsyncDetails kept reporting the batch as running. The wrapper now records whether the slot was held, funnels both acquire-time and post-acquire aborts through the same aborted branch (releasing only when held), and a batch regression test pins the contract.

Fixes #3930
2026-06-30 23:46:19 +00:00
roboomp 3bc8f995f6 fix(agent): bounded async job disposal
Made AsyncJobManager.dispose honor its timeout while waiting for cancelled jobs, and passed task abort signals into spawn semaphore waits.

Fixes #3930
2026-06-30 23:37:11 +00:00
roboomp 00ef58f843 fix(agent): steered override-schema subagents
- Marked eval agent schema calls as caller overrides so subagent prompts can revoke native output/yield instructions.\n- Added override-schema prompt guidance telling agents to ignore conflicting native output labels and terminal-yield the caller schema object.\n- Added prompt coverage for the override notice.\n\nRefs #3926
2026-06-30 22:30:15 +00:00
roboomp cff8f22d61 fix(task): preserved extension-root order for agent dedup
discoverAgents sorted the listOmpExtensionRoots result by level, which
demoted CLI-injected roots (hard-coded level user) below any project
extensions: settings entry. The sibling skills/hooks/tools surface in
discovery/omp-plugins.ts consumes the returned order verbatim, so an
explicit --extension override could run a different agent than the rest
of its own plugin surface.

Drop the sort and append roots in returned order. Refresh the
discoverAgents doc comment to document the source-precedence chain (CLI
> project settings > user settings > installed plugins) and add a
regression test where a CLI-injected extension wins over a project
extensions: settings extension that defines the same agent name.

Fixes #3920
2026-06-30 21:25:07 +00:00
roboomp b3cda0d921 fix(task): honored disabled omp-plugins for agents
Gate the OMP extension-package agents/ scan on the omp-plugins provider so
disabledProviders suppresses plugin-shipped task agents consistently with
other extension-package surfaces.

Add a regression test covering an installed npm plugin with agents/ while
omp-plugins is disabled.

Fixes #3920
2026-06-30 21:17:18 +00:00
roboomp 285d384ca6 fix(task): scan OMP extension agents/ dirs in discoverAgents
discoverAgents only walked .omp/agents and Claude marketplace plugin
roots, so agents shipped by OMP npm plugins (omp plugin install ...) and
--extension/extensions: settings roots silently disappeared while their
sibling skills/, hooks/, tools/ subdirectories were already discovered.

Route the same listOmpExtensionRoots scan used by discovery/omp-plugins.ts
through discoverAgents and append <root>/agents to the ordered scan list,
project scope before user. listOmpExtensionRoots already filters Claude
marketplace installs by realpath so they continue to flow only through
the claude-plugins provider.

Fixes #3920
2026-06-30 21:11:55 +00:00
roboomp e4561d64fa fix(coding-agent): restored subagent thinking precedence
Agent frontmatter thinkingLevel now wins over model role suffix thinking when both are configured.

Fixes #3915
2026-06-30 20:15:04 +00:00
can1357 720fb3f120 feat(coding-agent): replaced the oracle subagent with a new tester subagent
- Removed the oracle agent prompt configuration and its references across the codebase.
- Added a new tester agent prompt markdown file with directives for test authoring and techniques.
- Registered the new tester agent in the task runner while removing the oracle agent.
- Renamed the quick-task agent configuration and prompt references to sonic.
2026-06-30 16:16:39 +02:00
can1357 9ccd83a13d feat(coding-agent): made the agent parameter optional with a default value
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
2026-06-30 16:16:39 +02:00
metaphorics fa64097992 fix(coding-agent): identify user-invoked skills and expose skill directory
User-invoked skills (typed /skill:, steered, follow-up, interrupted/
resumed via compaction, ACP, RPC) only appended a bare "Skill: <path>"
line, so the model neither learned the user had invoked that specific
skill nor where the skill directory was. Relative paths in skill bodies
(scripts/, templates/) could not be resolved.

Route all user-invoked paths through a self-identifying, baseDir-aware
prompt template; keep hidden autoload skills on the minimal non-user
format. Interactive skillCommands now carries the loaded Skill object
instead of a bare path so baseDir flows through without reconstruction.
The invocation kind defaults to "user" to keep buildSkillPromptMessage
source-compatible.

Op: correct
Restores: spec:user-invoked-skill-prompt-self-identifies-and-exposes-skill-directory
2026-06-30 23:11:05 +09:00
roboomp 1b9c6be129 fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths
Three independent paths bypassed the user's subagent caps:

1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
   task.maxConcurrency only on first use and never re-read the setting,
   so lowering the cap mid-session left every later spawn running
   against the old ceiling. Resize the live semaphore against the
   current setting on each acquire.

2. The task tool prompt threaded MAX_CONCURRENCY through to the
   template but never rendered it. A model with task.maxConcurrency=1
   could still emit oversized tasks[] batches that registered
   immediately and piled up behind the semaphore. Render a 'Concurrency
   cap' directive in task.md whenever the setting is bounded.

3. The eval agent() bridge's assertDepthAllowed gated only against
   the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
   task.maxRecursionDepth, so a user-tightened recursion limit
   (0='None', 1='Single') still let cell-spawned subagents recurse
   to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
   by the hard ceiling.

Fixes #3895
2026-06-30 11:37:58 +00:00
oldschoola 5abde300a1 Merge remote-tracking branch 'origin/main' into perf/streaming-reveal-throughput 2026-06-29 22:08:03 -07:00
oldschoola a1972d6345 perf: cut hot-path quadratic/allocation costs across 5 subsystems
Five hot-path performance fixes + two low-risk allocation reductions.
No behavior change; all derived counts/orderings are identical.

- session-manager pathTo: leaf->root walk used branch.unshift() per node
  (O(n^2) over branch length); now push + single reverse. Backs
  getBranch(), hit at ~17 sites per turn.

- edit/modes/patch: collapseConsecutiveSharedLines / collapseRepeatedBlocks
  / trimCommonContext built shared-line sets via
  new Set(oldLines.filter(l => newLines.includes(l))) -> O(old*new) per
  hunk. Precompute new Set(newLines) and use .has() -> O(old+new).

- task/executor appendRecentOutputTail: re-split + filter + slice + reverse
  of the full (up to 8KB) recentOutputTail on every text_delta token. Fast
  path extends the current last line in place; full recompute only when a
  newline boundary or truncation changes the window. tailLastLineRepresentable
  flag guards the trailing-whitespace-only-line edge case.

- task/render renderResult: header booleans (3x .some) + footer counts
  (3x .filter) + request total (.reduce) re-scanned details.results ~30x/sec
  via the spinner. Single pass derives aborted/failed/mergeFailed/success
  counts + requestTotal; booleans derived from counts.

- task/render extractIncrementalReviewResult: re-called normalizeYieldData
  internally though both callers had already normalized the same yield data.
  Signature now takes pre-normalized RenderYieldItem[].

Honorable mentions (allocation reduction, no algorithmic change):
- config/model-resolver: hoist case-folded pattern out of matchModel filter
  passes; build the O(n) preference context once per role in
  resolveModelRoleValue and reuse across fallback patterns.
- tools/read countTextLines: count newlines directly instead of allocating
  via split("\n"); hashline formatter reuses the line count instead of
  recomputing.
2026-06-29 22:05:09 -07:00
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
can1357 8cd23ebd15 merge #3844: 3-way dirty-context fallback for isolated branch merges
# Conflicts:
#	packages/coding-agent/src/task/worktree.ts
#	packages/coding-agent/test/task/worktree.test.ts
2026-06-30 03:03:43 +02:00
roboomp d120ba6b7d fix(coding-agent): filter baseline wip from preserved agent commits
Dirty isolated baselines can be accidentally committed by subagents that run git add -A. Fetching the raw isolation HEAD then cherry-picking the range would replay that baseline WIP into parent history.

Add a dirty-baseline replay path that rewrites each agent commit against the captured baseline tree, preserving the agent commit message and author while excluding staged, unstaged, and untracked changes that existed before isolation started. Clean baselines still use the raw git fetch path, and nested-only changes keep returning patches without creating an empty root branch.

Add a regression for baseline staged + untracked WIP committed by the agent, asserting the task branch contains only the agent file and parent WIP remains staged/untracked after merge.

Fixes #3842
2026-06-30 00:28:27 +00:00
roboomp da715aae7c fix(coding-agent): preserve agent commits across isolated branch merges
When an isolated task agent commits its own changes before yielding, the
harness used to collapse the captured delta into one AI-summarized commit
and discard the agent's commit messages and authorship entirely. This
violated commit discipline for agentic swarms — multiple logical commits
("fix bug" + "add test") became a single opaque commit, and the
agent's commit object (which lived in isolation/.git/objects under
overlayfs/rcopy) was lost when cleanupIsolation tore down the overlay.

commitToBranch now detects when isolation HEAD moved past baseline.root
.headCommit. When it has, the function git-fetches the agent's HEAD into
the parent repo as omp/task/${taskId} so the commit objects survive
cleanupIsolation, and stamps the captured baselineSha onto the returned
CommitToBranchResult. mergeTaskBranches cherry-picks the inclusive range
baseSha..branchName when baseSha is provided, replaying each agent
commit verbatim with its original message and author. Any uncommitted
leftover (staged, unstaged, untracked) on top of the agent's last commit
becomes one trailing AI-summarized commit on the same branch.

Falls back to the legacy single-commit path when the agent never moved
HEAD (purely dirty working tree); existing patch-mode flow is untouched.

Fixes #3842
2026-06-30 00:13:07 +00:00
roboomp 4b98211c64 fix(coding-agent): fixed dirty isolated branch merges
Applied isolated branch patches with three-way fallback when unrelated parent dirt appears in patch context.

Surfaced branch preparation failures instead of reporting no changes.

Fixes #3841
2026-06-30 00:09:34 +00:00
roboomp 58c0a305d6 fix(agent): shortened isolated task paths
Used compact hashed isolation directory segments and the short m mount dir so long task ids are not copied into subagent working paths.

Kept worktree cleanup compatible with legacy merged task-isolation directories.

Fixes #3756
2026-06-28 21:25:20 +00:00
roboomp 2042d3b114 fix(task): scoped provider concurrency cap to each LLM turn
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.

Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).

Fixes #3749
2026-06-28 22:50:59 +02:00
can1357 a8dc036c84 feat(coding-agent): corrected payload assembly logic for terminal results
- Refactored payload assembly to distinguish between incremental sections and terminal results.
- Prevented terminal markers from being incorrectly treated as section labels to avoid payload nesting issues.
- Updated section processing to ignore non-incremental terminal items, resolving incorrectly missing data in output-schema validation.
- Improved terminal item resolution to correctly fallback to the last assistant text when no explicit data is provided.
2026-06-28 22:50:31 +02:00
can1357 93f68b75a2 test(coding-agent/modes): updated controller test mocks
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
2026-06-28 16:53:16 +02:00
can1357 51a2a0342f test(coding-agent): implemented guest reconciliation and expanded testing for collaboration
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
2026-06-28 09:52:44 +02:00
can1357 2cf28f972e feat(coding-agent): added logic to assembleYieldResult to
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
2026-06-28 09:05:37 +02:00
can1357 102d6d54ad feat: implemented v2 streaming remote compaction for model history state
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
2026-06-28 07:27:02 +02:00
can1357 d8bf177af7 refactor(coding-agent/task): moved taskToolRenderer to separate module
- Extracted taskToolRenderer to a dedicated renderer file to resolve circular dependencies.
- Updated all references to the renderer to point to the new location.
2026-06-28 07:27:02 +02:00
can1357 4fb1490747 feat(coding-agent/task): updated assistant review prompt and export renderer
- Updated the reviewer prompt instructions for clarity.
- Reordered export statements to address potential circular dependency issues during tool registration.
2026-06-28 07:27:02 +02:00
can1357 289dd770c8 feat(coding-agent): reworked subagent yields for incremental results
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
2026-06-28 07:27:01 +02:00
can1357 5a044dc0da feat: consolidated and automate changelog management
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
2026-06-27 03:08:52 +02:00
can1357 3345085290 feat(coding-agent/task): improved task description previews
- Updated task rendering to use `previewLine` for truncating descriptions consistently.
- Ensured progress and result descriptions are truncated before formatting.
2026-06-27 02:55:52 +02:00
can1357 577d2a8eb8 style: biome format/organize-imports across integrated PRs 2026-06-27 02:06:38 +02:00
can1357 29269d9547 Fix eval agent isolation artifact recovery 2026-06-27 01:39:35 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 2c20a2d368 Fix subagent yield abort cleanup
(cherry picked from commit e22673537f7cc95920cd4ec75d2a4b36d81d7626)
2026-06-26 23:43:03 +02:00
can1357 84ef014faa Merge PR #3407: fix(eval): dispose one-shot eval subagents after run (@korri123) 2026-06-26 23:27:40 +02:00
can1357 938489f3fd feat(coding-agent): added configurable service tier settings for subagents and advisor
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
2026-06-25 22:35:15 +02:00
can1357 06ca03dbd8 fix(agent): release task spawn slot when progress reporting throws
In #registerSpawnJob the markRunning()/reportProgress() calls sat between
semaphore.acquire() and the try whose finally releases the slot. If progress
reporting threw there, the acquired task.maxConcurrency slot leaked and
permanently shrank subagent concurrency. Move those statements inside the try
so finally always releases. The abort-before-execution branch is unchanged
(releases once and throws before the try is entered — not a double release).

Refs #3464
2026-06-25 20:48:38 +02:00