- Added `TokenTaskBudget` type and `taskBudget` option to `StreamOptions`.
- Forwarded `taskBudget` as `output_config.task_budget` with the `task-budgets-2026-03-13` beta header.
- Fixed `disableThinkingIfToolChoiceForced` to preserve `task_budget` when clearing `effort`.
- Accepted `output_config.task_budget` from Anthropic gateway requests.
- Extended ResolveContext / WriteContext with localProtocolOptions so the
internal-URL router can thread the calling session's local-root mapping
through to handlers.
- LocalProtocolHandler.resolveOptions now prefers context.localProtocolOptions
before consulting the process-global override or the first main-kind session
in AgentRegistry, fixing multi-session ACP hosts (cmux) where reads of
local://PLAN.md were routing to a sibling session's artifacts dir even
though plan-mode writes succeeded against the calling session.
- read, find, ast_grep, ast_edit, and search now thread
this.session.localProtocolOptions into the router so local://, memory://,
agent://, and other handlers see the right caller.
- Added regression tests covering the override-vs-context priority and the
ENOENT-against-caller-root path.
Fixes#1608
Guarded find renderer path summaries so raw pre-validation string paths render instead of throwing. Added coverage for pending, fallback, empty, and detailed result render paths.\n\nFixes #1622
The catch around the subagent yield-reminder prompt previously logged
every exception at ERROR. User cancel (^C) and compaction-driven aborts
both surface as ToolAbortError through awaitAbortable, so benign control
flow generated 9 spurious 'Subagent prompt failed' errors in 2 days on
the reporter's instance.
Gate the ERROR branch on '!abortSignal.aborted && !(err instanceof
ToolAbortError)' and route the abort path to logger.debug. The outer
catch + finally still mark the run aborted, so observable behaviour is
unchanged.
Fixes#1623
The first cut at the subprocess isolation swallowed every signal exit (`exitCode === null`) on the assumption it was the intentional SIGKILL from `terminate()`. That misclassifies real worker deaths — SIGSEGV from a native crash, SIGKILL from the OOM killer, an operator `kill -9` — so any in-flight title/completion/download promise would await forever while `#worker` still pointed at a dead process.
Added an `intentionalExit` flag flipped by `wrapSubprocess.terminate()` right before its SIGKILL. `onExit` swallows only the flagged exit; every other signal exit now fires the `errors` channel with a "signal SIGFOO" message so `TinyTitleClient.#handleWorkerError` clears `#pending` and dumps the dead worker handle. Added two regression tests pinning both branches.
Reported by chatgpt-codex-connector on #1607.
Moved the tiny title/memory worker from a Bun Worker thread into a child process spawned via Bun.spawn IPC. The agent CLI gains a hidden --tiny-worker dispatch the parent invokes through process.execPath; the parent SIGKILLs the child on dispose so onnxruntime-node's NAPI finalizer never runs in any address space the agent owns. On Windows that finalizer was segfaulting Bun at shutdown after the tiny title model loaded (issue #1606). Drops the now-dead 'close'/'closed' handshake and the unused parentPort bootstrap, and removes tiny/worker.ts from --compile worker entries in both build scripts plus the regression test that pinned them.
Fixes#1606
- Added `repairDoubleEncodedJsonString` to unescape fields double-encoded by the model (e.g. literal `\n`, `\"`, `\uXXXX` in `context`/`assignment`/`description`).
- Scoped repair to natural-language fields only, leaving code-bearing tools untouched.
- Applied repair on both render and execution paths in `TaskTool`.
- Converted [SECTION]...[/SECTION] markers to "SECTION\n===" format in system prompt templates.
- Updated system conventions doc to reference the new marker style.
- Updated tests to match against the new header pattern.
Included the reviews field in comments-enabled PR view fetches so pr:// output can show formal review submissions and approvals.
Added protocol coverage that emulates gh --json field selection before asserting rendered approval output.
Fixes#1600
Send the LSP exit notification after a successful shutdown response before falling back to process termination. Add a regression test that fails when a server receives shutdown but not exit.\n\nFixes #1593
- Replaced perimeter-based border animation with a bottom-edge `borderSegmentHeadCol` bounce cycle.
- Constrained animated segment rendering so only the bottom border can be darkened while other edges stay flat accent.
- Updated tests to validate non-teleporting width-based motion and bottom-edge easing behavior.
git clone --depth 1 --single-branch only fetches the tip of the
requested branch, so any subsequent git checkout <sha> for a non-tip
commit fails with 'reference is not a tree'. The error was caught and
rethrown as 'shallow clone may not contain this commit', but the clone
arguments were never adjusted.
Drop --depth 1 (and --single-branch when no ref is requested) when the
caller supplies options.sha so the desired commit is present in the
local object store. The ref-only path remains shallow.
Fixes#1589
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
- Fixed Anthropic stream idle-timeout errors incorrectly triggering provider retries after streaming had begun.
- Fixed darwin-x64 `bun build --compile` failure by guarding `onnxruntime-node` preload behind a `process.platform === "win32"` literal for dead-code elimination.
- Added `prefill` and `stop` parameters to the tiny-model worker's `complete` message type to pin output format without biasing content.
- Updated toggleToolOutputExpansion to pass allowUnknownViewportMutation when requesting render.
- Added a test that verifies tool output toggling sets expansion state and calls requestRender with the new flag.
- Documented the Ctrl+O POSIX tool-result expansion scrollback fix in the changelog.
- Added a post-dispatch refresh hook to recompute foreground tool render mode after key turn/tool events.
- Computed whether any non-background pending tool is active and toggled eager native scrollback rebuild accordingly.
- Added a regression test confirming eager rebuild is enabled while foreground tools are pending and disabled when none remain.
- Added a system prompt instruction to never re-audit applied edits.
- Added guidance to avoid routine `git status` and `git diff` checks, with exceptions for explicit requests and selective repo operations.
- Preserved hidden tmux overlays in live viewport while keeping native scrollback intact during forced renders.
- Adjusted forced-render handling so pure appends preserve full scrollback and return viewport repaint on equal-size diffs.
- Capped streaming edit diff previews to a fixed trailing window and reported hidden hunk counts.
- Updated tmux and regression tests to validate hidden overlays, nativeText frame sources, and scrollback retention.
- Removed stream preview height tracking and simplified tool call rendering to always invalidate the content box.
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
- Replaced composed keyword decorators in `CustomEditor.decorateText` with `highlightMagicKeywords`.
- Adjusted `UserMessageComponent` rendering to use `highlightMagicKeywords` with `keywordReset` for consistent foreground.
- Added `highlightMagicKeywords(text: string, resetTo?: string): string` and chained ultrathink, orchestrate, workflow glow.
- Updated ultrathink, orchestrate, and workflow matching to lowercase whitespace-delimited patterns with prose-only checks.
- Added `maskNonProse` and `keywordInProse` to skip fenced/inline code and HTML/XML segments while highlighting keywords.
- Added `KeywordHighlighter` with optional `resetTo`, and updated gradient highlighting to use masked match slicing.
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
- Replaced ad-hoc ANSI/VT stripping regexes with `stripVTControlCharacters` in status text handling and related tests.
- Updated status footer rendering to truncate using `truncateToWidth` and visible width after VT stripping.
- Extended tui cursor handling and rendering to strip markers from all lines and fit repaint/append-tail lines to width.
- Expanded deterministic render tests with overlay-aware assertions and recorded the truncation/cursor-marker behavior in changelogs.
- Added fuzzy token matching in agent-dashboard, state-manager, and tree-selector, replacing lowercased checks.
- Added search-query state and fuzzy-filter helpers to hook, oauth, and user-message selectors for query filtering.
- Updated filtered selectors to render match results, status lines, no-match text, and move selection within results.
- Added `overflowSearch` and filter state to `SelectList`, switching overflowing list matching to fuzzy checks.
- Configured `SelectList` input flow and fixed cancel so Escape/Ctrl+C closes lists when no matches exist.
- Updated changelogs and added tests for fuzzy-filter behavior in hook, oauth, user-message, and list selectors.
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
- Stored mcpManager and localProtocolOptions on ToolSession so nested subagents inherit them without relying on process-global singletons.
- TaskTool now uses the session's localProtocolOptions and mcpManager when spawning sub-tasks, falling back to defaults if absent.
- Previously the slider started at the current cycle index, so execution would inherit whichever model drove planning.
- Now finds the `default` role in the cycle and anchors the slider there, falling back to `currentIndex` if no default exists.
- Explicit `executionModel` is set whenever the chosen tier differs from the restored cycle position, covering the case where the slider stays on `default` but planning ran on another model.
Follow-up to #1503. When an extension registered a flag whose name collides
with a value-taking built-in — e.g. plan-mode's boolean `--plan` vs the
built-in `--plan <plan-model>` selector — the extension-aware reparse still
took the built-in branch. `omp --extension plan-mode --plan "review the diff"`
consumed "review the diff" as the plan-model value, leaving parsed.messages
empty and overwriting result.plan with the prompt text. recoverFlagValue only
patched the extension flag value, not the corrupted parsed object that
applyExtensionFlags returns as initialArgs.
Fix at the source: parseArgs now checks the registered extension-flag set
BEFORE the built-in branches, so a registered flag is parsed with the
extension's semantics (boolean toggle / string value) and surfaces in
unknownFlags without consuming the following token or touching the built-in
field. This makes recoverFlagValue dead, so applyExtensionFlags is simplified
to read resolved values straight from unknownFlags.
Tests: parseArgs-level shadowing guard (boolean --plan keeps the message and
leaves result.plan unset); applyExtensionFlags message/built-in-field
preservation for colliding boolean (--plan) and string (--model) flags;
non-colliding flag-looking-value rule retained. Verified the new guards fail
without the shadowing fix.
- Dropped `onShowHotkeys` callback and its binding from `CustomEditor` and `InputController`.
- `?` now inserts a literal question mark regardless of editor state; use `/hotkeys` explicitly.
- Added regression test confirming `?` is treated as plain input when the editor is empty.
Two review fixes for the extension-flag/initial-prompt work:
1. @file ordering — `processFileArguments` runs `process.exit(1)` on a
missing/unreadable file. It had been moved after `createSession`, which
writes the terminal breadcrumb eagerly (SessionManager.create →
#newSessionSync), so `omp @missing.md "x"` left a junk session/breadcrumb
behind before exiting.
Resolve extension-registered CLI flags BEFORE creating the session: load the
session's extensions up front (new `loadSessionExtensions` helper, the single
source of createAgentSession's discovery-branch logic), build an
ExtensionFlagSink straight from the loaded extensions + runtime, re-parse
argv, then process @file args — all before any session exists. The loaded
result is handed back to createAgentSession via `preloadedExtensions` (now
checked before `disableExtensionDiscovery`, so it can't double-load) and the
same EventBus is shared, so no extra work. This keeps the P1#1 fix
(`--flag @value` is the flag's value, not a file) while failing fast with no
session side effects.
2. "Can we avoid the big list of names?" — removed the hand-maintained
`BUILTIN_FLAG_NAMES` set (and its stale "rejected at registration" doc).
`applyExtensionFlags` now always falls back to recovering a flag's value from
argv when parseArgs didn't surface it; the recovery scan mirrors parseArgs's
consumption rules (flag-looking space-form values stay their own flag) and is
a no-op for flags that were absent or already surfaced, so no list of
built-in names is needed.
Adds `ExtensionRunner.aggregateFlags` (static) so getFlags and the CLI's
pre-session sink share one implementation.
Tests: pre-session flag resolution via the exact main.ts sink pattern;
list-free recovery of an arbitrary colliding built-in (`--model`); and the
flag-looking-value rule. Verified typecheck + extension/runner/acp suites.
- Restricted `setModel` to persist settings only when `persist: true` is passed; all runtime switches (Ctrl+P, `--model`, `/model`, model picker temp selections) no longer overwrite `modelRoles.default`.
- Changed `cycleRoleModels` to accept a direction ("forward"/"backward") instead of a `temporary` flag; both directions now use `applyRoleModel` without persisting.
- Added `persist: true` exclusively to the model picker's "Set as default" action in `SelectorController`.
- Added test suite covering persistence behavior for `setModel`, `cycleRoleModels`, and `cycleModel`.