Commit Graph
367 Commits
Author SHA1 Message Date
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
can1357 8cd23ebd15 merge #3844: 3-way dirty-context fallback for isolated branch merges
# Conflicts:
#	packages/coding-agent/src/task/worktree.ts
#	packages/coding-agent/test/task/worktree.test.ts
2026-06-30 03:03:43 +02:00
roboomp d120ba6b7d fix(coding-agent): filter baseline wip from preserved agent commits
Dirty isolated baselines can be accidentally committed by subagents that run git add -A. Fetching the raw isolation HEAD then cherry-picking the range would replay that baseline WIP into parent history.

Add a dirty-baseline replay path that rewrites each agent commit against the captured baseline tree, preserving the agent commit message and author while excluding staged, unstaged, and untracked changes that existed before isolation started. Clean baselines still use the raw git fetch path, and nested-only changes keep returning patches without creating an empty root branch.

Add a regression for baseline staged + untracked WIP committed by the agent, asserting the task branch contains only the agent file and parent WIP remains staged/untracked after merge.

Fixes #3842
2026-06-30 00:28:27 +00:00
roboomp da715aae7c fix(coding-agent): preserve agent commits across isolated branch merges
When an isolated task agent commits its own changes before yielding, the
harness used to collapse the captured delta into one AI-summarized commit
and discard the agent's commit messages and authorship entirely. This
violated commit discipline for agentic swarms — multiple logical commits
("fix bug" + "add test") became a single opaque commit, and the
agent's commit object (which lived in isolation/.git/objects under
overlayfs/rcopy) was lost when cleanupIsolation tore down the overlay.

commitToBranch now detects when isolation HEAD moved past baseline.root
.headCommit. When it has, the function git-fetches the agent's HEAD into
the parent repo as omp/task/${taskId} so the commit objects survive
cleanupIsolation, and stamps the captured baselineSha onto the returned
CommitToBranchResult. mergeTaskBranches cherry-picks the inclusive range
baseSha..branchName when baseSha is provided, replaying each agent
commit verbatim with its original message and author. Any uncommitted
leftover (staged, unstaged, untracked) on top of the agent's last commit
becomes one trailing AI-summarized commit on the same branch.

Falls back to the legacy single-commit path when the agent never moved
HEAD (purely dirty working tree); existing patch-mode flow is untouched.

Fixes #3842
2026-06-30 00:13:07 +00:00
roboomp 4b98211c64 fix(coding-agent): fixed dirty isolated branch merges
Applied isolated branch patches with three-way fallback when unrelated parent dirt appears in patch context.

Surfaced branch preparation failures instead of reporting no changes.

Fixes #3841
2026-06-30 00:09:34 +00:00
roboomp 58c0a305d6 fix(agent): shortened isolated task paths
Used compact hashed isolation directory segments and the short m mount dir so long task ids are not copied into subagent working paths.

Kept worktree cleanup compatible with legacy merged task-isolation directories.

Fixes #3756
2026-06-28 21:25:20 +00:00
roboompandcan1357 2042d3b114 fix(task): scoped provider concurrency cap to each LLM turn
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.

Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).

Fixes #3749
2026-06-28 22:50:59 +02:00
can1357 a8dc036c84 feat(coding-agent): corrected payload assembly logic for terminal results
- Refactored payload assembly to distinguish between incremental sections and terminal results.
- Prevented terminal markers from being incorrectly treated as section labels to avoid payload nesting issues.
- Updated section processing to ignore non-incremental terminal items, resolving incorrectly missing data in output-schema validation.
- Improved terminal item resolution to correctly fallback to the last assistant text when no explicit data is provided.
2026-06-28 22:50:31 +02:00
can1357 93f68b75a2 test(coding-agent/modes): updated controller test mocks
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
2026-06-28 16:53:16 +02:00
can1357 51a2a0342f test(coding-agent): implemented guest reconciliation and expanded testing for collaboration
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
2026-06-28 09:52:44 +02:00
can1357 2cf28f972e feat(coding-agent): added logic to assembleYieldResult to
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
2026-06-28 09:05:37 +02:00
can1357 102d6d54ad feat: implemented v2 streaming remote compaction for model history state
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
2026-06-28 07:27:02 +02:00
can1357 d8bf177af7 refactor(coding-agent/task): moved taskToolRenderer to separate module
- Extracted taskToolRenderer to a dedicated renderer file to resolve circular dependencies.
- Updated all references to the renderer to point to the new location.
2026-06-28 07:27:02 +02:00
can1357 4fb1490747 feat(coding-agent/task): updated assistant review prompt and export renderer
- Updated the reviewer prompt instructions for clarity.
- Reordered export statements to address potential circular dependency issues during tool registration.
2026-06-28 07:27:02 +02:00
can1357 289dd770c8 feat(coding-agent): reworked subagent yields for incremental results
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
2026-06-28 07:27:01 +02:00
can1357 5a044dc0da feat: consolidated and automate changelog management
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
2026-06-27 03:08:52 +02:00
can1357 3345085290 feat(coding-agent/task): improved task description previews
- Updated task rendering to use `previewLine` for truncating descriptions consistently.
- Ensured progress and result descriptions are truncated before formatting.
2026-06-27 02:55:52 +02:00
can1357 577d2a8eb8 style: biome format/organize-imports across integrated PRs 2026-06-27 02:06:38 +02:00
can1357 29269d9547 Fix eval agent isolation artifact recovery 2026-06-27 01:39:35 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 2c20a2d368 Fix subagent yield abort cleanup
(cherry picked from commit e22673537f7cc95920cd4ec75d2a4b36d81d7626)
2026-06-26 23:43:03 +02:00
can1357 84ef014faa Merge PR #3407: fix(eval): dispose one-shot eval subagents after run (@korri123) 2026-06-26 23:27:40 +02:00
can1357 938489f3fd feat(coding-agent): added configurable service tier settings for subagents and advisor
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
2026-06-25 22:35:15 +02:00
can1357 06ca03dbd8 fix(agent): release task spawn slot when progress reporting throws
In #registerSpawnJob the markRunning()/reportProgress() calls sat between
semaphore.acquire() and the try whose finally releases the slot. If progress
reporting threw there, the acquired task.maxConcurrency slot leaked and
permanently shrank subagent concurrency. Move those statements inside the try
so finally always releases. The abort-before-execution branch is unchanged
(releases once and throws before the try is entered — not a double release).

Refs #3464
2026-06-25 20:48:38 +02:00
can1357 ecd6f608e5 fix(agent): resize provider concurrency limiter in place instead of replacing it
The per-provider subagent limiter (providers.ollama-cloud.maxConcurrency)
created a fresh Semaphore whenever the configured limit changed, orphaning
in-flight slots on the old instance so a runtime or mixed limit value could
exceed the cap. getProviderSemaphore now always hands out one shared limiter
(Infinity when unlimited, so every run is still counted) and resizes it in
place. Semaphore.release() decrements before admitting, and the new
Semaphore.resize() raises the ceiling by admitting queued waiters while
lowering it drains in-flight holders without admitting past the new cap.

Refs #3464
2026-06-25 20:48:38 +02:00
roboomp 7e90e4d081 fix(agent): released ollama-cloud semaphore slot when waiter aborts
Semaphore.acquire now accepts an AbortSignal so a queued waiter that is cancelled (parent task abort, wall-clock budget elapsing) removes itself from the wait queue instead of being resolved by the next release. The provider semaphore in runSubprocess passes the run's abortSignal through, preventing aborted ollama-cloud subagents from permanently draining the provider concurrency budget.

Fixes #3464
2026-06-25 12:14:45 +00:00
roboomp 80862b79da fix(agent): handled ollama-cloud task backoff
Added ollama-cloud subagent concurrency limiting, role fallback-chain inheritance, and visible empty length errors for native Ollama responses.

Fixes #3464
2026-06-25 11:57:51 +00:00
Kormákur 0a7ec389a0 fix(coding-agent): verify eval subagent cleanup 2026-06-25 09:27:38 +00:00
Kormákur 9118e2cb95 fix(coding-agent): dispose eval subagents after run 2026-06-24 22:15:03 +00:00
roboomp e4c52de24d fix(task): normalize fractional spawn concurrency
Normalize the configured semaphore max before deciding whether the spawn

limit is bounded. Fractional values between 0 and 1 now truncate to 0 and

fall through to the unbounded path instead of storing 0 and deadlocking

the first acquire.

Fixes #3305
2026-06-23 10:41:51 +00:00
roboomp 296125cce9 fix(task): treat maxConcurrency 0 as unbounded in spawn semaphore
The session-scoped spawn Semaphore clamped its max via Math.max(1, max), so

task.maxConcurrency: 0 — labeled 'Unlimited' in the settings UI — serialized

subagent spawns one at a time instead of releasing every eligible seat.

The constructor now treats max <= 0 (and any non-finite input) as unbounded

via Number.POSITIVE_INFINITY, so the existing 'current < max' check naturally

permits every acquire. Mirrors the eval parallel()/pipeline() worker-pool

semantics (runEvalConcurrency in eval/concurrency-bridge.ts), which already

treats 0 as 'run every item at once'.

Fixes #3305
2026-06-23 10:37:34 +00:00
roboompandcan1357 5b6e9f904d fix(eval): paused timeout over baseline capture and surfaced nested stash-restore failures
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:

1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.

2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.

Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.

Fixes #3196
2026-06-22 21:57:35 +02:00
can1357 2f2acaf928 Merge pull request #3205: feat(eval): isolated/apply/merge options for agent() helper
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.

Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
  forwards isolated/apply/merge (as booleans) plus returnHandle, and the
  return_handle node carries isolated/patch_path/branch_name/
  nested_patches/changes_applied/isolation_summary.

Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
  task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
  the stale "defaults track task.isolation.mode" wording to the final
  strict opt-in behavior, and note all four runtimes.

Fixes #3196
2026-06-22 21:56:45 +02:00
roboomp abc9a8f292 fix(task): restored nested stash with index state
git stash pop without --index restores stashed staged changes as unstaged. When a nested repo had staged WIP before the isolated agent ran, the pop in applyNestedPatches() brought the content back but lost the user's index state.

Pass { index: true } so pop uses --index, matching the root merge path that already does the same thing.

Added a regression test that stages a pre-existing edit in the nested repo, runs applyNestedPatches, and asserts the file is still in the index (porcelain "M  " with the trailing space) and the cached diff still shows the staged WIP.

Fixes #3196
2026-06-22 19:41:59 +00:00
roboomp 978d2a76d0 fix(task): preserved nested-repo dirty state across nested patch apply
applyNestedPatches() applied the captured patch then ran git.stage.files(nestedDir), which stages every working-tree change in the nested repo. A nested repo that was already dirty before the agent ran ended up with the user's unrelated work-in-progress committed alongside the agent delta.

Stash any pre-existing dirty state (tracked + untracked) before applying the patch and pop it back in the finally block after the commit, so the agent commit contains only the captured patch and the user's in-flight work is restored on top of it. A failing stash pop logs a warning and leaves the stash entry intact for manual recovery; the broader nested-apply failure path is already non-fatal.

Added a worktree integration test that confirms a pre-existing untracked file in the nested repo is not staged into the agent commit and is still present in the working tree afterwards.

Fixes #3196
2026-06-22 19:34:42 +00:00
roboomp 28137f46d9 refactor(task): deduped nested patch apply + commit-message factory
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.

Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.

Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.

Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.

Fixes #3196
2026-06-22 19:22:24 +00:00
can1357 33e2594f03 feat(coding-agent): added support for Julia and display language icons in code cells
- Added Julia language support to the theme symbol maps.
- Enabled language icons in code cell headers for the eval tool renderer.
2026-06-22 06:13:01 +02:00
can1357 1f3f3cf5d1 feat: added ruby and julia language support to coding-agent
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
2026-06-22 06:13:01 +02:00
roboomp a6ee95df08 fix(eval): preserved returnHandle artifacts and nested branch patches
Eval preludes now forward returnHandle to the bridge so no-session eval runs can preserve the temp artifacts backing returned agent:// handles. The bridge keeps those temporary artifact directories whenever returnHandle is requested, including non-isolated runs and successful isolated applies.

Branch-mode isolation now treats nested-only changes as merge-eligible even when no root branch was produced, letting callers apply nested patches instead of dropping them when the root repo had no diff.

Added regression coverage for returnHandle artifact preservation and nested-only branch isolation.

Fixes #3196
2026-06-22 03:59:17 +00:00
roboomp 6a22499dcd style: bun run fix 2026-06-21 17:20:14 +00:00
roboomp eeed4b93b6 feat(eval): added isolated/apply/merge options to agent() helper
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.

Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.

Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.

Fixes #3196
2026-06-21 17:20:07 +00:00
can1357 f353ae0067 chore: biome format after PR integration 2026-06-21 17:20:22 +02:00
can1357 e8b60408b2 Merge PR #1939: fix(coding-agent): guard Git-mutating automation in pure jj workspaces (@roboomp)
# Conflicts:
#	packages/coding-agent/src/task/worktree.ts
#	packages/coding-agent/test/task/worktree.test.ts
2026-06-21 17:05:15 +02:00
can1357 6bc194b0fb feat(coding-agent): added subagent launch latency instrumentation
- Add `onFirstChatDispatch` hook to `CreateAgentSessionOptions` to track the boundary between session creation and the initial model request.
- Update `runSubprocess` and `TaskTool` to measure and log detailed latency metrics across the subagent lifecycle, including semaphore queue wait, setup time, and dispatch latency.
2026-06-19 16:06:17 +02:00
can1357 29d250fae2 feat(coding-agent): supported advisor transcript persistence
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
2026-06-19 03:50:00 +02:00
can1357 dd69e0e4fe refactor(coding-agent): hardened input schema validation and constraints
- Hardened the `hashline` parameters parsing pipeline to guarantee presence of the `input` field.
- Enforced input length limits on task roles to secure against oversized payloads.
- Configured Arktype schemas to reject or delete extra, undeclared fields in task and inspect-image payloads.
- Updated mock test parameters to align with corrected success exit codes.
2026-06-18 01:57:23 +02:00
can1357 17ce678468 fix(coding-agent): handled provider error finish reasons and preserve subprocess failure codes
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
2026-06-18 01:11:10 +02:00
can1357 e92db73ec6 feat(coding-agent/tools): added explicit ArkType schema descriptions to agent tools
- Added explicit ArkType schema descriptions across all coding agent tool definitions.
- Updated schema definitions in autoresearch and commit tools with descriptive wrappers.
- Documented tool schema enhancements in the packages/coding-agent CHANGELOG.
2026-06-18 00:59:55 +02:00
can1357 a050474af7 feat: migrated validation schemas and tool definitions from Zod to ArkType
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
2026-06-18 00:59:53 +02:00
benandcan1357 c93774f892 fix(coding-agent): implement session stop hook semantics 2026-06-17 12:26:52 +02:00