- Elevated `eval` to an essential tool to ensure availability across all discovery modes.
- Updated system and tool prompts to mandate the use of `eval` for non-trivial shell operations like conditionals, loops, heredocs, and complex pipelines.
- Restricted `bash` usage to simple binary invocations and single-fact computation to reduce shell-escaping and execution errors.
The eval agent() helper used `agent_type`/`return_handle` (snake_case) in
Python/Ruby/Julia and `agentType`/`returnHandle` (camelCase) in JS, forcing
the prelude docs to repeat every option twice ("JS same but camelcased").
Both are now single lowercase words identical across all four runtimes, and
`agent` matches the `task` tool's existing agent-selection parameter.
- Renamed across py/js/rb/jl preludes (signatures, forwarding, docstrings).
- Renamed the `__agent__` bridge wire protocol + `EvalAgentArgs` (`agentType`
→ `agent`, `returnHandle` → `handle`) so no prelude-side remap is needed.
- Updated prompt docs (workflow-notice.md, tools/eval.md), repo docs
(docs/tools/eval.md, docs/python-repl.md), and all bridge/prelude tests.
- CHANGELOG: Breaking Changes entry under [Unreleased].
Two Codex P2 findings landed against 978d2a76d0 that were not in the previously delivered review event:
1) prepareIsolationContext() (which runs captureBaseline → walks nested repos and untracked diffs) was running OUTSIDE withBridgeTimeoutPause; on dirty/large repos the baseline walk can exceed the eval idle timeout while the runtime is blocked. Moved the prep call into the pause closure so the watchdog is suspended for the whole bridge call from prep through cleanup.
2) applyNestedPatches() swallowed git stash pop failures with only a logger.warn, so a stash-pop conflict after a successful agent commit was invisible to the workflow. Changed the helper to return Promise<string[]> of warnings; applyEligibleNestedPatches now wraps them in a <system-notification> appended to the merge summary so the caller actually sees the partial-success case.
Added regression tests:
- bridge: prepare fires after timeout-pause and before timeout-resume.
- runner: applyEligibleNestedPatches surfaces stash-restore warnings as a system-notification.
- worktree (real git): a pre-existing dirty edit on the same file the agent patches causes stash pop to conflict; the helper returns a warning naming the nested repo and the stash entry is preserved for manual recovery.
Fixes#3196
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.
Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
forwards isolated/apply/merge (as booleans) plus returnHandle, and the
return_handle node carries isolated/patch_path/branch_name/
nested_patches/changes_applied/isolation_summary.
Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
the stale "defaults track task.isolation.mode" wording to the final
strict opt-in behavior, and note all four runtimes.
Fixes#3196
- Clamped tool output preview height to the available viewport rows to stop redundant banner commits.
- Added `outputBlockContentWidth` helper to accurately measure visual lines for scrollback budget calculations.
- Updated `bash` and `eval-render` output wrapping to account for block padding and inner content width.
- Added regression test confirming streaming tool output maintains a stable line count without duplicating headers.
git stash pop without --index restores stashed staged changes as unstaged. When a nested repo had staged WIP before the isolated agent ran, the pop in applyNestedPatches() brought the content back but lost the user's index state.
Pass { index: true } so pop uses --index, matching the root merge path that already does the same thing.
Added a regression test that stages a pre-existing edit in the nested repo, runs applyNestedPatches, and asserts the file is still in the index (porcelain "M " with the trailing space) and the cached diff still shows the staged WIP.
Fixes#3196
applyNestedPatches() applied the captured patch then ran git.stage.files(nestedDir), which stages every working-tree change in the nested repo. A nested repo that was already dirty before the agent ran ended up with the user's unrelated work-in-progress committed alongside the agent delta.
Stash any pre-existing dirty state (tracked + untracked) before applying the patch and pop it back in the finally block after the commit, so the agent commit contains only the captured patch and the user's in-flight work is restored on top of it. A failing stash pop logs a warning and leaves the stash entry intact for manual recovery; the broader nested-apply failure path is already non-fatal.
Added a worktree integration test that confirms a pre-existing untracked file in the nested repo is not staged into the agent commit and is still present in the working tree afterwards.
Fixes#3196
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.
Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.
Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.
Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.
Fixes#3196
The two julia-prelude tests pay a ~11-12s Julia kernel cold-start
(JIT + package precompile) when run with reset: true, exceeding
Bun's default 5000ms per-test timeout. CI exposed this since 33e2594f0
(Julia eval support) ran on PR runners (ubuntu-22.04, Julia
preinstalled) rather than the omp-kata self-hosted runners that lack
Julia and skip the suite.
Bumped both tests to 30_000ms, matching the precedent in
agent-bridge.test.ts:540 for similar persistent-kernel tests.
Fixes#3274
Previously withBridgeTimeoutPause only wrapped the subagent subprocess; mergeIsolatedChanges, applyNestedPatches, nested commit-message generation, and artifact cleanup ran with the eval watchdog re-armed. A cherry-pick or large patch apply could trip the cell timeout and abort successful post-processing.
Moved the entire bridge work (subprocess + merge + nested apply + cleanup + usage recording) inside one withBridgeTimeoutPause block. The pause helper still resumes via its finally on success and on throw, so existing failure paths are unchanged.
Added a regression that captures the emitted op order and asserts merge fires after timeout-pause and before timeout-resume.
Fixes#3196
Per maintainer ruling on #3196, eval agent() now defaults to non-isolated regardless of task.isolation.mode, mirroring the task tool. isolated=true is the only way to turn it on; isolated=true while task.isolation.mode === "none" still throws the same clear error.
Updated tests, workflow-notice.md, and Python agent() docstring to reflect the strict opt-in contract. Existing isolation tests now pass isolated:true explicitly; the inherit-from-settings assertion is replaced with a default-off + isolated=true opt-in regression.
Fixes#3196
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
API-key /login appends sibling keys for every provider (#3265); the stale test still asserted MiniMax replace-on-relogin (#156). Assert append, same-key idempotency, and individual removal via removeCredential (the logout path).
- Added `Shell.liveBackgroundJobCount` to query active background processes.
- Retained per-call `:async:` shells if background jobs are still running upon turn completion.
- Reaped shells automatically once their last background process exits to prevent lingering processes.
- Added a `share.store` configuration option (`blob` | `gist`) that allows users to choose between the default share server or a GitHub gist for storing exported session data.
- Changed the default upload target from secret GitHub gists to the share server to avoid GitHub API rate limits for shared sessions.
- Enabled fallback to the share server when a gist upload fails or the GitHub CLI is unavailable.
The completions provider stores session state under the request-time resolved base URL, which can differ from the catalog baseUrl for Moonshot, Alibaba Coding Plan, Azure deployments, and similar provider overrides. The model-switch cleanup now evicts the previous provider prefix whenever the switch leaves that completions backend, so those resolved-url keys cannot survive the switch.
`AgentSession.#closeProviderSessionsForModelSwitch` only handled
`openai-codex-responses` and `openai-responses:<provider>` keys. The
`openai-completions:<provider>:<baseUrl>:<modelId>` entries — which cache
strict-tools disable scopes and reasoning-effort fallbacks tied to the
upstream backend — survived /model switches between different providers or
base URLs, so the next request to that backend (e.g. on /model toggle
back) replayed stale decisions made against an entirely different
transport.
Switching to a model whose `(provider, baseUrl)` differs from the current
openai-completions model now evicts every cached entry sharing the old
prefix. Same-backend model toggles keep their cached state, matching the
existing codex/responses semantics.
Fixes#3260
Bun.spawn -> CreateProcess only appends `.exe` to extensionless names; `.cmd`/`.bat` are never tried. Bare MCP commands like `npx` (which exists only as `npx.cmd` on Windows) crashed the subprocess ~140ms after spawn with ENOENT/EINVAL whenever our own PATH walk couldn't pin the file down (empty `Bun.env.PATH` under a restricted parent process, UNC mounts that reject `fs.access`, locked-down shells).
`resolveStdioSpawnCommand` now routes any unresolved bare command through `cmd.exe /d /s /c` so Windows's PATHEXT search runs Windows-native. Direct-spawn fast path is preserved for resolved `.exe`/`.com` files; the existing `.cmd`/`.bat` wrap is unchanged. The reporter's stated cause ("process.env not merged") was incorrect — the merge has been in place at `transports/stdio.ts:316-319` since well before 16.1.14 — but the symptom they hit is real.
Fixes#3250
Require bracketed image-path paste detection to see an explicit local path separator or file URI before routing .png-like text to image attachment handling.\n\nFixes #3253
PerplexityProvider.isAvailable() accepted authStorage.hasAuth("openrouter")
as a valid credential, so any user with an OpenRouter key configured (for
LLM access) had every webSearch: auto request silently routed through
OpenRouter's perplexity/sonar-pro endpoint. Since Perplexity sits first in
SEARCH_PROVIDER_ORDER, downstream providers like Gemini were never reached
and users saw unexpected charges on their OpenRouter billing.
Auto-chain admission now requires a direct Perplexity credential
(PERPLEXITY_COOKIES, Perplexity OAuth, or PERPLEXITY_API_KEY).
isExplicitlyAvailable still returns true, so users who want the
OpenRouter-backed perplexity/sonar-pro path can opt in by setting
webSearch: perplexity explicitly — the existing OpenRouter fallback in
getApiConfigs handles that case unchanged.
Fixes#3251
Configured provider discovery now treats models.yml/models.json edits as a cache staleness boundary, forcing online-if-uncached refreshes instead of reusing fresh rows written before the config change.
Added regression coverage for Ollama metadata overrides with a pre-existing models.db row.
Fixes#3242
Walked custom/hook `details` recursively through the obfuscator so nested renderer fields (e.g. async-result `jobs[].label`) cannot leak configured secrets into the advisor prompt.
Fixes#3237
Rewrote file-mention path and content through the configured obfuscator before the advisor delta is formatted, matching the primary provider's hide-secrets behavior.
Fixes#3237
brush-core's alias expander resolves aliases via
`value.split_ascii_whitespace()` (`crates/brush-core-vendored/src/interp.rs:1500`,
upstream brush issue reubeno/brush#57): each whitespace piece is dropped
into argv as-is, completely bypassing the shell parser. Any alias body
containing `(`, `)`, `|`, `&`, `;`, `<`, `>`, or `\`` therefore
turns the first piece into the command name, so Fedora's default
`alias which='(alias; declare -f) | /usr/bin/which …'` produces
`error: command not found: (alias;` for every `which` invocation.
The user's shell snapshot is generated by sourcing their real rc-file
under `/bin/bash` or `/bin/zsh` (so we can capture functions, options,
PATH) and then sourced by brush per-session. `sanitizeSnapshotForBrush`
now scans the emitted `alias -- NAME='VALUE'` lines after generation,
drops any whose decoded body contains those metacharacters, and rewrites
the file in place before caching. Compatible aliases (`ll='ls -l'`,
`gc='git --color=auto commit'`, embedded-quote `say='echo '\\''hi'\\'''`)
are preserved untouched; dropped names are logged at debug. brush then
falls through to whatever lives on `PATH`, which is what the user
expected when they ran `which` in the first place.
Covered by unit tests for the sanitizer (Fedora-which case, every
incompatible-metachar shape, every preserve case) and an integration
test that loads a poisoned snapshot and verifies `which sh` now exits
`0` with a real path.
Fixes#3234
- Updated system and tool prompts to explicitly forbid using shell utilities like grep, rg, awk, and find for tasks better suited to specialized tools.
- Clarified that bash should be reserved for terminal operations and computational pipelines that produce facts not available through existing specialized tool outputs.
- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
- Extracted common kernel and executor logic into `BaseKernel` and `executor-base` to eliminate duplicated implementations for Julia, Python, and Ruby.
- Migrated shared operational workflows--including session namespacing, environment filtering, and result mapping--to centralized backend helpers.
- Consolidated runtime discovery and resolution logic into a unified `runtime-env` utility module.
- Simplified language-specific modules by delegating subprocess lifecycle, IPC, and configuration management to the newly established base classes.
- Centralized Parallel API utilities and parsing logic into a single module.
- Exported constants and helper functions from `parallel.ts` to replace duplicated definitions in the search provider.
- Updated the search provider to leverage the unified `parseParallelSearchPayload` function with metadata parsing toggled off.
- Assert the eval tool hides disabled backends from the model-facing wire
schema (language enum + field descriptions), summary, and description by
default (rb/jl off), and advertises them once enabled — including the
enabled-subset case.
- Update the env-flag fallback test for the new rb/jl opt-in defaults and
extend the env guard to PI_RB/PI_JL so the suite is shell-independent.
- Note the opt-in default and dynamic advertising in the changelog.
When an isolated apply fails the bridge throws a ToolError and never returns details, so the nested-patch payload that previously lived in details.nestedPatches was unrecoverable after the isolation worktree was torn down.
The bridge now writes each captured nested patch to a file under the per-call artifacts dir (e.g. <agentId>.nested-<index>-<slug>.patch) before throwing and includes the resolved paths in the error message so the caller can apply them manually.
Added a regression test verifying the persisted file exists with the original patch contents and that the path is surfaced in the thrown error.
Fixes#3196
A failed isolated apply (changesApplied === false) previously only set details.isolationSummary and returned the subagent text. Schema-backed agent() calls then parsed the JSON and returned the object, so workflows saw a successful structured result while none of the edits had landed.
Throw a ToolError when mergeIsolatedChanges reports a failed apply, with the merge summary plus a recovery hint pointing at the preserved patch/branch/nested artifacts so the caller can apply manually.
Added regression tests for the schema and non-schema apply-failure paths.
Fixes#3196
- Changed default evaluation backend configuration to only enable Python and JavaScript by default.
- Implemented dynamic tool parameter generation to hide Ruby and Julia from the model's schema when they are disabled in settings.
- Updated tool summary and field descriptions to reflect the currently enabled runtime backends.