Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.
Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.
Fixes#3356
- Replaced global timer mocks with local `YieldGate` instances to avoid test flakiness from concurrent environment interference.
- Introduced an injected clock and counting sleep pattern to test gating logic without relying on `process` globals.
- Added a test case to ensure the gate correctly handles negative clock jumps without stalling.
- Extracted yield logic into a configurable YieldGate class to avoid process-global state.
- Injected time and sleep dependencies to support deterministic testing.
- Handled potential negative time progression by forcing a re-anchor instead of gating indefinitely.
- Maintained existing behavior for the public yieldIfDue export via a shared instance.
- Prevented infinite blocking in `yieldIfDue` caused by backward system clock adjustments.
- Implemented time-delta re-anchoring to detect and recover from negative clock drift.
- Decoupled the yield gate into an injectable component to allow deterministic testing.
- Corrected TTY property restoration to properly delete injected properties instead of redefining them as truthy values.
- Updated environment restoration to modify process.env keys individually to prevent breaking environment object reference bindings.
- Prevented pollution of TTY-gated code and environment variable state across test suites.
- Pin colorMode to none in golden test suites to prevent nondeterministic ANSI escape sequences.
- Normalize render calls across test files to ensure stable output comparison.
- Prevent test flakes caused by environment-specific TTY auto-detection in the renderer.
- Constrained git configuration for smart-HTTP requests by explicitly overriding proxy, sslVerify, and credential helpers across all relevant path suffixes.
- Prevented potential credential capture by disabling repo-configured credential helpers that could otherwise execute malicious commands during authentication challenges.
- Hardened git operations against attacker-injected proxies by exhaustively blanking configuration keys for all identifiable git request endpoints.
- Excluded sslCAInfo/sslCAPath from overrides to prevent premature TLS negotiation failure while maintaining security via mandatory proxy neutralization.
- Added `_require_fetch_ref` validator to enforce strict alphanumeric character sets and disallow special git characters (e.g., `:`, `--`, `..`, `*`).
- Integrated validation into `git_fetch_ref_endpoint` to block malicious refspec inputs before git execution.
- Added test cases in `test_proxy_server.py` to verify rejection of attempted shell and refspec injections.
- Added test suite to verify git configuration overrides for auth tokens.
- Validated that repo-local git configurations cannot override security-critical settings such as proxy, sslVerify, and credential helpers.
- Confirmed that smart-HTTP path-specific overrides are correctly applied to prevent proxy-based MITM attacks.
- Ensured system CA locations remain untouched to prevent disruption of TLS verification.
- Added a check to reject remote URLs that begin with a hyphen to prevent command-line option injection.
- Updated the test suite to verify that option-shaped URLs are correctly blocked.
- Extracted `_pat_safe_remote` and rejected HTTP(S) origins with embedded credentials or mismatched host/repo.
- Guarded `clone` via `_assert_clone_url_safe` on the caller-supplied `clone_url` (pool has no `origin` yet).
- Asserted origin safety before `fetch`, `fetch_ref`, and `fetch_pr_head` inject the PAT header.
- Appended POSIX `--` separator in `omp_local` so prompts starting with `-` aren't parsed as flags.
- Added proxy tests covering attacker-origin fetch rejection and unsafe `clone_url` refusal.
Co-authored-by: can1357 <me@can.ac>
`__computeBunfsPackageRoot` now returns `//root/packages` for the Bun 1.3.14
`//root/<binary>` import.meta.dir shape, but production immediately joined that
root with shim and package segments through `path.join`, which collapses the
POSIX double-slash bunfs mount back to `/root`. That still made override
validation miss the embedded shim files.
Added a bunfs join helper that preserves the `//root` mount prefix after joining
production descendants, wired `bunfsPath` through it, and extended the #3329
regression test to assert the full typebox shim path stays under
`//root/packages/...`.
Fixes#3329
The reporter clarified that the failing binary is the pre-built
`omp-darwin-arm64` release asset from GitHub Releases; Homebrew is only a
local-tap wrapper that downloads that asset. The fix already covers every
cross-compiled `<bunfs-root>/<binary>` shape, but the source/test docstrings
and changelog blurb framed it as a Homebrew-build-specific bug. Updated those
three call sites to name the release asset and note the Homebrew tap as a
downstream consumer of the same binary; no code change.
Bun 1.3.14 reports `import.meta.dir` as `<bunfs-mount>/<binary-basename>` for
the compiled entry on some hosts — e.g. the Homebrew darwin-arm64 build sees
`//root/omp-darwin-arm64` instead of the bunfs root alone. The pre-fix path
joined `metaDir` with `"packages"` and baked the binary basename into every
bunfs path, so the typebox / legacy-pi shim overrides failed `existsSync`
validation, `resolveCanonicalPiSpecifier` fell through to a bunfs
`Bun.resolveSync` that also could not find the module, and every third-party
`@oh-my-pi/pi-*` extension was silently dropped.
`__computeBunfsPackageRoot` now detects the trailing binary-basename segment
(`path.basename(path.dirname(metaDir)) === "root"`) and strips it off the
original `metaDir` via string slicing rather than `path.join`, so Bun's
bunfs-native `//root` and `B:\~BUN\root` prefixes survive verbatim
(`path.posix.join` would collapse `//root` to `/root`). The single-segment
`<bunfs-root>` and deep `<bunfs>/packages/coding-agent/src/extensibility/plugins`
paths keep their existing branches.
Regression test added in `legacy-pi-bunfs-root.test.ts` for the POSIX
`//root/<bin>`, POSIX `/$bunfs/root/<bin>`, and Win32 `<drive>:\~BUN\root\<bin>.exe`
shapes.
Fixes#3329
Use a non-resolving discovery context for selected-model llama.cpp metadata refresh so command-backed and OAuth credentials stay lazy during model switches.\n\nFixes #3310
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
Normalize the configured semaphore max before deciding whether the spawn
limit is bounded. Fractional values between 0 and 1 now truncate to 0 and
fall through to the unbounded path instead of storing 0 and deadlocking
the first acquire.
Fixes#3305
The session-scoped spawn Semaphore clamped its max via Math.max(1, max), so
task.maxConcurrency: 0 — labeled 'Unlimited' in the settings UI — serialized
subagent spawns one at a time instead of releasing every eligible seat.
The constructor now treats max <= 0 (and any non-finite input) as unbounded
via Number.POSITIVE_INFINITY, so the existing 'current < max' check naturally
permits every acquire. Mirrors the eval parallel()/pipeline() worker-pool
semantics (runEvalConcurrency in eval/concurrency-bridge.ts), which already
treats 0 as 'run every item at once'.
Fixes#3305
Pruned empty optional MCP argument placeholders before tools/call while preserving required fields and meaningful falsy values.
Added regression coverage for active and deferred MCP tools.
Fixes#3302
`nohup cmd &` is now a transparent background wrapper that double-forks the
operand so it reparents to init (commit 00dcd54597). The shell only tracks
the short-lived intermediate fork, so `$!` is no longer the surviving
process — the prior test read `$!`, then `process.kill(pid, 0)` checked an
already-reaped pid and failed on Linux (the failing CI job).
Split into two contracts:
- plain `&` retention: stays a child of the shell, counted by
`liveBackgroundJobCount`, kept alive by the retain map; `$!` is the real
child pid we assert on.
- nohup reparenting: the operand writes its own pid before `exec`ing the
long sleep, and that (post-exec-stable) pid is asserted to survive across
turns — independent of `$!`.
- Added `sanitizeOpenAIResponsesReasoningItemForReplay` to process reasoning-type items by stripping unique identifiers and filtering properties.
- Updated the main sanitization utility to route reasoning items through the new logic.
- Introduce `detach_reparent` parameter to command execution to support process reparenting.
- Add `detach_session_reparent` to Unix command extensions using a double-fork technique to orphan processes from the shell descendant tree.
- Update background pipeline logic to automatically apply reparenting when unwrapping transparent wrappers like `nohup`.
- Remove unused `command_is_resolvable` helper.
- Standardized `local://` image processing to prevent file corruption during decoding.
- Refactored local path resolution logic to enforce safety constraints and path containment.
- Implemented an image fast-path in `ReadTool` to correctly render local images before text decoding.
- Added comprehensive test coverage for image rendering, text compatibility, and path security.
- Resolved an event loop hang associated with `omp --resume` operations.
- Force process exit when the startup session picker is cancelled instead of returning.
- Prevent hanging the event loop caused by long-lived startup handles such as theme listeners and timers.
- Add regression test case to verify clean process termination upon picker cancellation.
Vertex Claude rawPredict expects Anthropic beta flags in the JSON body as anthropic_beta. The Anthropic stream path can only add context-management-2025-06-27 as an HTTP header there, so sending context_management would make reasoning requests fail.
Omit context_management and its beta header for google-vertex Anthropic models while preserving thinking itself, and add regression coverage to the rawPredict routing test.
Fixes#3288
Injected Anthropic clients bypass buildAnthropicClientOptions, so this package cannot add the context-management beta header their SDK instance would need before accepting context_management.clear_thinking_20251015.
Omit context_management for options.client requests while preserving thinking itself, and add regression coverage for injected-client payload shaping.
Fixes#3288
Anthropic rejects context_management.clear_thinking_20251015 without the context-management-2025-06-27 beta header. OAuth requests carried it via claudeCodeAgentBetaDefaults; API-key requests pushed the new context_management field without the beta, so the server rejected them.
Push the beta into extraBetas alongside the field for every API-key thinking request, and exclude the GitHub Copilot proxy (which strips Anthropic betas and demotes thinking blocks upstream) from emitting the field.
Fixes#3288
Sent context_management.keep=all for all enabled Anthropic thinking requests so API-key and Anthropic-compatible providers preserve replayed reasoning blocks across turns.
Added regression coverage for budget and adaptive thinking payloads.
Fixes#3288
Resolve the session-selector.ts conflict by integrating the delete-dialog
content-slot fix (#3283) on top of the fullscreen mouse-picker refactor.
The branch swapped the delete-confirmation dialog INTO a single content
slot (replacing the SessionList) so the picker is always
`chrome + max(list, dialog) + chrome` and never overflows the viewport.
Adjustments baked into this merge:
- Wrap the SessionList in `#contentSlot` and keep the dialog swapping into
that slot, but preserve the new fullscreen path: mouse hit-testing,
the pinned footer (`#footerLines`/`#footerStart`), and fill-height
trimming all still work because the render offset now tracks
`#contentSlot` (the list lives one level down).
- Keep both CHANGELOG entries (picker mouse/fullscreen + #3283 fix) and
the ported scroll-stability regression test.
- Enabled fullscreen overlay rendering for the terminal session picker.
- Implemented full mouse support including wheel-based scrolling and click-to-select functionality.
- Anchored the session picker footer to the bottom of the viewport to correct UI flickering.
- Added comprehensive unit tests for mouse interaction and layout constancy during resizing.
Earlier rounds shrank the SessionList by the dialog's row count to keep
the picker inside the viewport, but the SessionList could only claw back
whole session rows and bottomed out at zero entries. On a narrow
terminal with a long session title the dialog still wrapped past what
the SessionList could free, the picker overflowed the viewport, and the
TUI committed the header into native scrollback.
The picker now hosts the SessionList inside a single contentSlot
Container. Opening the delete confirmation swaps the dialog INTO that
slot (replacing the SessionList); closing it swaps the SessionList back.
The dialog therefore competes only with the SessionList's rendered
budget, not with the SessionList AND the picker chrome, so the picker
frame stays bounded by terminalRows even when the dialog wraps to many
rows. SessionList's external-reserve plumbing is no longer needed and is
removed.
Addresses PR #3285 second-round review feedback.
The first round of the issue #3283 fix reserved a fixed 12 SessionList
rows for the delete confirmation dialog. On a narrow terminal or against
a long session name, HookSelectorComponent's Markdown title and help
text wrap past 12 rows; the picker would still overflow even after the
SessionList shrank to zero entries, and the TUI committed the picker
header into native scrollback again.
SessionSelectorComponent now overrides render() to measure the dialog's
actual rendered height at the live width before super.render() walks
the children, and pushes that as the SessionList's external-row reserve.
The dialog's own Container memoization makes the extra pre-render
essentially free.
Addresses PR #3285 review feedback.
- Update the eval tool to stream stdout chunks directly into the active cell's output buffer while the process is still running.
- Prevent long-running cells from appearing empty in the UI by surfacing incremental output before the backend resolves.
- Add regression tests to ensure streamed output is captured mid-execution and reconciled with final results.
- Added `off` and `auto` as valid inputs for the `--thinking` CLI flag.
- Centralized thinking level definitions in `CLI_THINKING_LEVELS` to keep flag options, shell completions, and validation in sync.
- Configured CLI parsing to reject `inherit` as an explicit input to prevent unintended configuration suppression.