- Update the eval tool to stream stdout chunks directly into the active cell's output buffer while the process is still running.
- Prevent long-running cells from appearing empty in the UI by surfacing incremental output before the backend resolves.
- Add regression tests to ensure streamed output is captured mid-execution and reconciled with final results.
- Added `off` and `auto` as valid inputs for the `--thinking` CLI flag.
- Centralized thinking level definitions in `CLI_THINKING_LEVELS` to keep flag options, shell completions, and validation in sync.
- Configured CLI parsing to reject `inherit` as an explicit input to prevent unintended configuration suppression.
- Removed deprecated eval prelude helpers `append`, `tree`, `diff`, `sort`, `uniq`, and `counter` from all supported runtimes.
- Cleaned up runtime implementations, protocol definitions, and UI rendering logic associated with the removed helpers.
- Updated project documentation, prompts, and test suites to reflect the reduced helper API surface.
- Recorded functional changes in the package changelog.
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
- Refactored `todo` tool to accept a single operation object instead of an `ops` array.
- Implemented parameter normalization to maintain backward compatibility with legacy array-based tool calls.
- Updated tool instructions, documentation, and UI rendering components to reflect the new interface.
- Added compatibility tests to verify rendering and execution for both legacy and current operation formats.
- Updated streaming diff renderer to use visual line wrapping instead of line count for preview budget.
- Introduced width-aware slicing to ensure displayed content stays within the provided UI bounds.
- Refactored cache salt parameters to incorporate width for consistent rendering across window resizes.
- Update input handler to recognize CSI-u escape sequences using `matchesKey`.
- Ensure legacy bare escape sequences remain supported in environments without the protocol.
- Add test coverage for both CSI-u and legacy escape inputs.
- Introduced a dynamic row budget for the diff preview to prevent overflow during streaming.
- Integrated `previewWindowRows()` into the cache key to ensure proper re-rendering upon viewport resizing.
- Simplified tail window logic to consistently apply the preview budget.
- Added `minDimension` option to ensure images meet minimum size requirements for vision backends.
- Implemented logic to scale up undersized input images while respecting maximum constraints.
- Clamped minimum dimension floor to avoid resolution conflicts with defined maximum bounds.
- Added headless plot configuration for Julia to prevent GUI popup windows during execution.
- Implemented robust mime-bundle serialization in Julia using `invokelatest` to handle runtime-loaded library methods.
- Added IRuby-protocol and magic-byte image sniffing support to the Ruby runner to enable inline rendering for graphics gems like Gruff, ChunkyPNG, and RMagick.
- Elevated `eval` to an essential tool to ensure availability across all discovery modes.
- Updated system and tool prompts to mandate the use of `eval` for non-trivial shell operations like conditionals, loops, heredocs, and complex pipelines.
- Restricted `bash` usage to simple binary invocations and single-fact computation to reduce shell-escaping and execution errors.
The eval agent() helper used `agent_type`/`return_handle` (snake_case) in
Python/Ruby/Julia and `agentType`/`returnHandle` (camelCase) in JS, forcing
the prelude docs to repeat every option twice ("JS same but camelcased").
Both are now single lowercase words identical across all four runtimes, and
`agent` matches the `task` tool's existing agent-selection parameter.
- Renamed across py/js/rb/jl preludes (signatures, forwarding, docstrings).
- Renamed the `__agent__` bridge wire protocol + `EvalAgentArgs` (`agentType`
→ `agent`, `returnHandle` → `handle`) so no prelude-side remap is needed.
- Updated prompt docs (workflow-notice.md, tools/eval.md), repo docs
(docs/tools/eval.md, docs/python-repl.md), and all bridge/prelude tests.
- CHANGELOG: Breaking Changes entry under [Unreleased].
Active goal loops can stay inside one agent run while the model keeps
emitting tool calls, so the normal agent_end threshold maintenance never
runs. That lets context grow past the soft threshold until provider
overflow or user abort.
Run threshold maintenance from the per-turn onTurnEnd hook for active
goals, splice the compacted agent state back into the live loop message
array, and suppress queued continuations because the current run is
already continuing. Cover the mid-run tool-call path and the non-goal
control case.
Refs #3174
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:
1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.
2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.
Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.
Fixes#3196
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.
Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
forwards isolated/apply/merge (as booleans) plus returnHandle, and the
return_handle node carries isolated/patch_path/branch_name/
nested_patches/changes_applied/isolation_summary.
Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
the stale "defaults track task.isolation.mode" wording to the final
strict opt-in behavior, and note all four runtimes.
Fixes#3196
- Clamped tool output preview height to the available viewport rows to stop redundant banner commits.
- Added `outputBlockContentWidth` helper to accurately measure visual lines for scrollback budget calculations.
- Updated `bash` and `eval-render` output wrapping to account for block padding and inner content width.
- Added regression test confirming streaming tool output maintains a stable line count without duplicating headers.
git stash pop without --index restores stashed staged changes as unstaged. When a nested repo had staged WIP before the isolated agent ran, the pop in applyNestedPatches() brought the content back but lost the user's index state.
Pass { index: true } so pop uses --index, matching the root merge path that already does the same thing.
Added a regression test that stages a pre-existing edit in the nested repo, runs applyNestedPatches, and asserts the file is still in the index (porcelain "M " with the trailing space) and the cached diff still shows the staged WIP.
Fixes#3196
applyNestedPatches() applied the captured patch then ran git.stage.files(nestedDir), which stages every working-tree change in the nested repo. A nested repo that was already dirty before the agent ran ended up with the user's unrelated work-in-progress committed alongside the agent delta.
Stash any pre-existing dirty state (tracked + untracked) before applying the patch and pop it back in the finally block after the commit, so the agent commit contains only the captured patch and the user's in-flight work is restored on top of it. A failing stash pop logs a warning and leaves the stash entry intact for manual recovery; the broader nested-apply failure path is already non-fatal.
Added a worktree integration test that confirms a pre-existing untracked file in the nested repo is not staged into the agent commit and is still present in the working tree afterwards.
Fixes#3196
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.
Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.
Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.
Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.
Fixes#3196
The two julia-prelude tests pay a ~11-12s Julia kernel cold-start
(JIT + package precompile) when run with reset: true, exceeding
Bun's default 5000ms per-test timeout. CI exposed this since 33e2594f0
(Julia eval support) ran on PR runners (ubuntu-22.04, Julia
preinstalled) rather than the omp-kata self-hosted runners that lack
Julia and skip the suite.
Bumped both tests to 30_000ms, matching the precedent in
agent-bridge.test.ts:540 for similar persistent-kernel tests.
Fixes#3274
Previously withBridgeTimeoutPause only wrapped the subagent subprocess; mergeIsolatedChanges, applyNestedPatches, nested commit-message generation, and artifact cleanup ran with the eval watchdog re-armed. A cherry-pick or large patch apply could trip the cell timeout and abort successful post-processing.
Moved the entire bridge work (subprocess + merge + nested apply + cleanup + usage recording) inside one withBridgeTimeoutPause block. The pause helper still resumes via its finally on success and on throw, so existing failure paths are unchanged.
Added a regression that captures the emitted op order and asserts merge fires after timeout-pause and before timeout-resume.
Fixes#3196
Per maintainer ruling on #3196, eval agent() now defaults to non-isolated regardless of task.isolation.mode, mirroring the task tool. isolated=true is the only way to turn it on; isolated=true while task.isolation.mode === "none" still throws the same clear error.
Updated tests, workflow-notice.md, and Python agent() docstring to reflect the strict opt-in contract. Existing isolation tests now pass isolated:true explicitly; the inherit-from-settings assertion is replaced with a default-off + isolated=true opt-in regression.
Fixes#3196
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
Raced queued Exa throttle waits against the caller abort signal so requests cancelled behind an earlier throttle wait reject immediately without breaking the serialized throttle chain.
Added regression coverage for cancelling a third Exa request queued behind another delayed request.
Fixes#3271
Made Exa request pacing observe cancellation during the configured delay instead of waiting for the full delay before checking the signal.
Added regression coverage for a queued Exa request cancelled while throttled.
Fixes#3271
Added configurable Exa search request pacing via exa.searchDelayMs so repeated web_search calls no longer burst directly into Exa rate limits.
Covered the provider contract with a focused Exa test and recorded the back-to-back request repro.
Fixes#3271
Keep the ask tool on the current question when the custom answer editor is dismissed, so Escape from Other returns to the option selector instead of aborting or recording an empty answer.
Fixes#3269
API-key /login appends sibling keys for every provider (#3265); the stale test still asserted MiniMax replace-on-relogin (#156). Assert append, same-key idempotency, and individual removal via removeCredential (the logout path).
- Added `Shell.liveBackgroundJobCount` to query active background processes.
- Retained per-call `:async:` shells if background jobs are still running upon turn completion.
- Reaped shells automatically once their last background process exits to prevent lingering processes.
- Added a `share.store` configuration option (`blob` | `gist`) that allows users to choose between the default share server or a GitHub gist for storing exported session data.
- Changed the default upload target from secret GitHub gists to the share server to avoid GitHub API rate limits for shared sessions.
- Enabled fallback to the share server when a gist upload fails or the GitHub CLI is unavailable.
The completions provider stores session state under the request-time resolved base URL, which can differ from the catalog baseUrl for Moonshot, Alibaba Coding Plan, Azure deployments, and similar provider overrides. The model-switch cleanup now evicts the previous provider prefix whenever the switch leaves that completions backend, so those resolved-url keys cannot survive the switch.
`AgentSession.#closeProviderSessionsForModelSwitch` only handled
`openai-codex-responses` and `openai-responses:<provider>` keys. The
`openai-completions:<provider>:<baseUrl>:<modelId>` entries — which cache
strict-tools disable scopes and reasoning-effort fallbacks tied to the
upstream backend — survived /model switches between different providers or
base URLs, so the next request to that backend (e.g. on /model toggle
back) replayed stale decisions made against an entirely different
transport.
Switching to a model whose `(provider, baseUrl)` differs from the current
openai-completions model now evicts every cached entry sharing the old
prefix. Same-backend model toggles keep their cached state, matching the
existing codex/responses semantics.
Fixes#3260