- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
- Preserved newest-first ordering for startup changelog markdown so collapsed notices report the current release.
- Kept default changelog rendering oldest-first for explicit recent/full views.
- Added regression assertions for both startup and full-history heading order.
- Treated missing or invalid changelog markers as first install and persisted the current version without replaying historical notes.
- Shared bounded changelog rendering between startup and recent changelog views, with a 64 KiB startup cap and full-history hint on truncation.
- Added marker, truncation, recent/full rendering, and PTY startup regression coverage.
Fixes#5135
- Ran commit host completion before commit-agent session disposal so mnemopi/autolearn teardown cannot preempt a valid proposal.
- Converted missing commit-agent host outputs and split-plan gaps into thrown errors so omp commit cannot resolve into exit 0 without creating a commit.
- Preserved caller GPG_TTY state instead of forcing a bogus signing TTY in git and non-interactive subprocess environments.
Fixes#4794
- Applied the known reasoning envelope filter when deciding whether a title marker is visible.
- Added title extraction regressions for reasoning tag and reasoning fence envelopes before the visible title.
Fixes#5122
- Limited markerless fallback cleanup to leading leaked-thinking envelopes so literal reasoning syntax in plain titles survives.
- Added markerless regression coverage for think tags and thinking-fence titles.
Fixes#5122
- Parse only title markers that remain visible after leaked-thinking cleanup so markers inside leaked reasoning are skipped.
- Preserve literal reasoning tag syntax inside the chosen title and cover it with a regression test.
Fixes#5122
- Reused the leaked-thinking healer before parsing title markers so visible reasoning envelopes cannot win extraction.
- Added regression coverage for <thinking> and <think> envelopes that contain internal title tags before the real title.
Fixes#5122
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
- Removed literal HTML comment sentinels (`<!-- -->`) from thinking block displays.
- Added logic to hide blocks that consist entirely of reasoning noise and updated display validation to omit empty formatted output.
- Refactored the memoization cache to maintain separate slots for prose and raw modes.
- Replaced tool-based `set_title` invocation with XML-style `<title>` marker tags for session title discovery.
- Implemented robust JSON-unwrapping logic to handle and sanitize title generation outputs.
- Updated model registry in catalog with new model support, provider prefixes, and metadata adjustments.
- Synchronized system prompt documentation and test suites to reflect the new marker-based generation flow.
Follow-up to the #4420 opener hardening: absolute-path rundll32 fixes the
stripped-PATH spawn throw, but rundll32 exits 0 unconditionally, so the
delayed-failure telemetry added there can never observe a Windows launch
failure. Replace it with %SystemRoot%-resolved PowerShell Start-Process
via -EncodedCommand:
- failures ShellExecute itself reports (missing target, no handler
executable, access denied) surface as exit code 1 and reach the
existing non-zero-exit logging (verified live on Windows 11: missing
file exits 1; unregistered schemes exit 0 on any opener because the
OS hands them to the app-picker — documented limitation);
- the UTF-16LE/base64 payload keeps OAuth query strings opaque to
cmd/PowerShell metacharacter parsing; embedded single quotes are
doubled into a PS literal;
- %SystemRoot% anchoring with a bare-name PATH fallback preserves the
stripped-PATH resilience from #4420.
Also pins the WSL-mount test's path.resolve against Windows dev hosts
so the mocked linux platform stays deterministic.
Refs #4418
An intermediate commit whose net effect is already on HEAD (redundant
change, or 3-way merged to HEAD by "theirs == ours") stopped the
sequencer with "The previous cherry-pick is now empty" and was
treated as a hard conflict. mergeTaskBranches aborted the whole range,
marked the branch failed, and dropped every remaining non-overlapping
commit.
Add cherryPick.skip and cherryPick.isEmptyError to the git namespace,
then in mergeTaskBranches' catch classify the failure before aborting:
loop --skip while the error stderr matches the "now empty" phrase so
consecutive empties advance the sequencer; fall through to abort/fail
on the first non-empty error (genuine conflict with unmerged files).
Fixes#4438
Two independent defects broke /mcp reauth against S256-only providers on
Windows boxes whose PATH no longer references System32:
1. openPath spawned bare rundll32 and swallowed the
`Executable not found in $PATH` throw with a bare `catch {}`, so the MCP
controller's outer try/catch was dead and the transcript unconditionally
claimed "Opening browser automatically...".
2. TUI#prepareLine silently truncates any composed row wider than the
viewport. MCPAuthorizationLinkPrompt rendered `Copy URL: <full URL>` as a
single ~271-column line whose trailing parameter is
code_challenge_method=S256. On the reporter's 270-col terminal the cut
landed inside that parameter, dropping the method while keeping
code_challenge — which RFC 7636 §4.3 treats as plain PKCE, which Linear
correctly rejects with "The plain PKCE method is not allowed. Use S256
instead."
OAuthCallbackFlow now hosts a `GET /launch` route on the same loopback
callback server it already runs; the route 302-redirects to the pending
authorization URL and is advertised as `OAuthAuthInfo.launchUrl` — a
~30-char copy target no viewport can meaningfully truncate. The MCP OAuth
fallback, /login, setup wizard, auth-broker CLI, and login-dialog all
prefer the launch URL for the visible copy target, keep the full URL in
the OSC 8 hyperlink for click-through, and the MCP flow additionally
stages the copy target on the clipboard via OSC 52 (same pattern the
setup wizard uses).
openPath now resolves rundll32.exe through %SystemRoot%\System32 (with a
C:\Windows fallback when SystemRoot is unset) and logs both synchronous
spawn throws and non-zero exits via the shared logger, so silent
misconfigurations show up in ~/.omp/logs/omp.*.log. The dead try/catch
around openPath in the MCP controller is removed.
Fixes#4418
Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.
Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.
Sites fixed:
- utils/title-generator.ts (30 -> 1024)
- utils/commit-message-generator (60 -> 1024)
- tts/speech-enhancer (512 -> 1536)
- auto-thinking/classifier online path (8 -> 1024); classifyLocal
keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
classifyLocal keeps ANSWER_MAX_TOKENS
Fixes#4355
readTextFromClipboard called execSync for pbpaste, termux-clipboard-get,
wl-paste, and xclip; readMacFileUrlsFromClipboard did the same for
osascript; copyToClipboard for termux-clipboard-set. execSync parks the
event loop until the child exits or the 2000ms timeout fires, so a hung
clipboard daemon froze the TUI render loop for the whole budget on every
paste and copy chord (input-controller handleImagePaste and
handleClipboardTextRawPaste).
A new spawnCapture helper wraps Bun.spawn with the same 2000ms guard,
stdout-to-string decoding, and non-zero-exit/timeout throw semantics the
outer try/catch already assumed. Every synchronous clipboard shell-out
now yields to the event loop while the child runs. Regression test
under readTextFromClipboard runs a slow fake pbpaste and asserts a
concurrent setInterval keeps ticking; the pre-fix code delivered zero
ticks.
Fixes#4235
Added timeout-backed AbortSignals to update, Hindsight, and Smithery fetch calls so stalled endpoints abort instead of hanging indefinitely.
Added regression coverage for the timeout signals on the exposed command/client paths.
Fixes#4229
getOrCreateSnapshot in shell-snapshot.ts pre-creates snapshotPath as an empty temp file and returns null on spawn failure, timeout, or nonzero exit without deleting it, leaving stale files in os.tmpdir()/omp-shell-snapshots/. Track whether snapshot creation succeeded and remove the pre-created file in a finally block on every failure path using fs.rmSync(snapshotPath, { force: true }), with best-effort error suppression so cleanup failures never propagate. Verified with `bunx tsc --noEmit -p packages/coding-agent/tsconfig.json` and `bun test packages/coding-agent/test/shell-snapshot.test.ts` (22 pass).
Closes#4236
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
- Added GIT_NETWORK_TIMEOUT_MS (30 min) for clone/fetch with an overridable timeoutMs option; local plumbing keeps the 5-minute cap.
- Migrated fetch() from a positional AbortSignal to an options object.
Failed stash-pop cleanup now invokes git clean with literal pathspecs for
stash-derived untracked paths. Filenames such as `:(glob)*` are valid POSIX
filenames and valid Git pathspec magic; passing them as ordinary pathspecs with
`-x` could delete unrelated ignored artifacts that were never stashed and are
not recoverable from the preserved stash.
Extend the fallback regression with a literal `:(glob)*` stash file and an
ignored `build.log` that must survive cleanup.
Fixes#4175
When a task branch adds ignore rules for a path that was untracked in the
user's stashed WIP, a failed stash pop can restore the file and then leave it
hidden from normal status after reset. Default `git clean -fd -- <path>` does
not remove ignored files, so the partial restore could still leak into later
isolated task baselines.
Add an includeIgnored clean mode and use `git clean -fdx -- <stash path>` for
failed stash-pop cleanup. Extend the fallback regression so the task branch adds
.gitignore for the restored untracked path and verify both normal and ignored
status return clean.
Fixes#4175
A failed `git stash pop --index` can restore unrelated untracked files before
exiting on a tracked conflict while still preserving the stash entry. The
previous fallback only reset tracked/index state, leaving those untracked files
in the working tree for subsequent task baselines.
Record the top stash entry's untracked paths before popping and clean exactly
those paths if the pop fails after preflight. Add a regression that forces the
fallback branch and verifies the worktree returns clean with the stash preserved.
Fixes#4175
mergeTaskBranches and applyNestedPatches both stashed dirty WIP, cherry-picked
task branches, then called `git stash pop` in a finally block. On conflict git
left stage 1/2/3 unmerged entries in .git/index with no MERGE_HEAD to abort;
neither call cleaned up. The corrupted index persisted indefinitely, and every
subsequent overlay-isolated task inherited it through the lower layer —
captureRepoDeltaPatch then emitted `diff --cc` (combined merge format) that
git apply rejects with "No valid patches in input", failing every downstream
task merge with 'Branch merge failed before a task branch could be created'.
Fix at the git API level: git.stash.tryPop now runs `git apply --3way --check`
on `git stash show -p --binary stash@{0}` before popping (`--3way` matches
what git stash pop does internally, so context that drifted after cherry-pick
is still accepted). Preflight failure short-circuits — stash entry preserved,
index untouched. Preflight pass falls through to pop; if pop still leaves
unmerged entries (mode-only or delete/modify conflicts the preflight can miss),
a `reset --hard HEAD" fallback restores the merged HEAD without losing the
cherry-picked commits (stash is preserved by git on failed pop, so the user's
WIP stays recoverable).
Both call sites now share this contract via git.stash.tryPop.
Fixes#4175
Ports only the thinking double-format fix: resolveThinkingDisplay reuses block.thinking when rawThinking is set (buildDisplayMessage already formatted it), plus a single-entry memo in formatThinkingForDisplay and a rawThinking regression test. The PR's incremental reveal slicing is superseded by the already-merged #3848 (memoized grapheme slicing).
captureRepoDeltaPatch records the delta against `HEAD + WIP`, so the
patch's context lines and blob SHAs reference the WIP-modified files.
commitPatchToBranchWorktree then tried to apply that patch to a fresh
worktree pinned at HEAD, which failed hard whenever the WIP-side file
was missing from HEAD's index (untracked WIP files, staged-new WIP
files) or when --3way could not resolve an overlap.
commitPatchToBranchWorktree now tries plain apply first, then `--3way`
(which cleanly subtracts WIP via the shared ODB blob for tracked files),
and only when both fail replays baseline WIP into the temp worktree so
the delta's context lines up, rewinding WIP-only files afterward so
they never leak into the branch commit.
Added git.ls.tree helper for the WIP-only-file filter and a set of
regression tests covering the untracked, staged-new, and overlap
scenarios.
Fixes#4136
git apply --3way --check exits 0 even when the real apply would write conflict markers and unmerged index stages, so the previous fix left the worktree dirty on conflicting patches while only flipping changesApplied to false.
Dropped --3way for patch-mode merge and used a --reverse --check probe instead: it succeeds only when the target state is already present (true no-op) and reads without touching the worktree. Conflicts fall through to the normal --check + apply path, which rejects them before writing anything.
Added regression coverage for the conflict scenario asserting the worktree stays clean, and for the fresh apply path.
The model selector's persistence path dropped the `:auto` selector when parsing role values, producing a warning ('Invalid thinking level "auto"') and rendering the badge as `inherit` instead of `auto`. Reload of the default role also lost the auto state whenever the role value carried an explicit `:auto` suffix instead of relying on `defaultThinkingLevel`.
Widen the resolver chain (`parseThinkingSuffix`, `splitThinkingSuffix`, `parseModelString`, `parseModelPattern*`, `ResolvedModelRoleValue`, `ResolvedRoleModel`, `ResolveCliModelResult`) to carry the `AUTO_THINKING` sentinel end to end, and coerce it back to `undefined` at concrete-only boundaries (glob scope patterns, retry fallback, advisor, commit pipeline, guided-goal, bench).
Regression tests cover:
- `resolveModelRoleValue("provider/model:auto")` returns explicit auto without a warning.
- `ModelSelector` renders `DEFAULT (auto)` and `SMOL (auto)` when the role value has `:auto`.
- `cycleRoleModels` activates auto thinking on entering a `:auto` role.
- Startup resume activates auto thinking when `modelRoles.default` carries `:auto`.
Fixes#4128
Forced non-interactive credential env for git and gh subprocesses, added a default timeout, and capped captured stdout/stderr with a truncation marker. Added regression coverage for prompt env, output capping, and timeout cleanup.
Fixes#4072