- Updated toggleToolOutputExpansion to pass allowUnknownViewportMutation when requesting render.
- Added a test that verifies tool output toggling sets expansion state and calls requestRender with the new flag.
- Documented the Ctrl+O POSIX tool-result expansion scrollback fix in the changelog.
- Updated optional-embeddings tests to use MNEMOPI_* environment variable names instead of MNEMOSYNE_*.
- Updated withEnv calls and ENV_KEYS entries so optional embeddings test configuration matches current env var naming.
- Added a post-dispatch refresh hook to recompute foreground tool render mode after key turn/tool events.
- Computed whether any non-background pending tool is active and toggled eager native scrollback rebuild accordingly.
- Added a regression test confirming eager rebuild is enabled while foreground tools are pending and disabled when none remain.
- Added `TUI#setEagerNativeScrollbackRebuild` to opt into native scrollback rebuilds during unknown-viewport rendering.
- Hooked the new flag into render mutation handling so offscreen and structural updates rebuild immediately when enabled.
- Documented the behavior and added a regression test validating clean, non-duplicated scrollback after streaming layout edits on POSIX.
- Adjusted TUI render logic to force a native history rebuild when terminal width changes, so POSIX sessions rewrap committed history even when viewport position is unknown.
- Preserved the existing defer path for known scrolled-away viewports, avoiding unnecessary destructive rebuilds outside explicit resize events.
- Added a regression test that resizes an unknown-viewport terminal from 20 to 40 columns and verifies old wrapped rows are replaced by rewrapped lines.
- Added a system prompt instruction to never re-audit applied edits.
- Added guidance to avoid routine `git status` and `git diff` checks, with exceptions for explicit requests and selective repo operations.
- Replaced the streaming edit diff rendering strategy with a fixed-height trailing window so partial re-diff streams stayed anchored and stopped oscillating.
- Updated streaming preview tests to assert the window remained saturated without trailing blank padding across chunked updates and still produced a real diff after finalization.
- Tuned TUI viewport repaint heuristics for pure appends and multiplexer sessions, and adjusted render-stress handling for preserved tmux scrollback.
- Preserved hidden tmux overlays in live viewport while keeping native scrollback intact during forced renders.
- Adjusted forced-render handling so pure appends preserve full scrollback and return viewport repaint on equal-size diffs.
- Capped streaming edit diff previews to a fixed trailing window and reported hidden hunk counts.
- Updated tmux and regression tests to validate hidden overlays, nativeText frame sources, and scrollback retention.
- Removed stream preview height tracking and simplified tool call rendering to always invalidate the content box.
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
- Fixed scrolled-up readers being yanked to the tail on POSIX terminals when streaming content arrived; unknown native viewport position was treated as "at bottom", triggering destructive history rebuilds.
- Fixed appended rows slipping down by the height delta when a resize and new content coalesced into one frame; height-grow repaint now only fires when content fits within the new viewport.
- Added regression test covering the height-grow-with-new-content case.
- Changed the render planner so height increases with new overflow now trigger a viewport repaint in non-Termux and non-multiplexer sessions.
- The new path marks native scrollback dirty and defers full rebuild, preventing anchor-loss while streaming inserts are offscreen.
- Added a POSIX unknown-viewport regression test that verifies scrolled-up readers stay on their rows during streaming and recover after an explicit native refresh.
- Added markdown prose tests to verify maskNonProse preserves plain prose and masks fenced/tilde blocks.
- Added keyword-in-prose tests that enforced lowercase token matching and ignored code spans, fences, XML, and comments.
- Added highlight-mode tests ensuring only prose keywords are ANSI-highlighted and inline/fenced/XML regions stay unmodified.
- Added workflow/orchestrate assertions that uppercase or path-like forms remain unchanged while lowercase token boundaries are handled correctly.
- Replaced composed keyword decorators in `CustomEditor.decorateText` with `highlightMagicKeywords`.
- Adjusted `UserMessageComponent` rendering to use `highlightMagicKeywords` with `keywordReset` for consistent foreground.
- Added `highlightMagicKeywords(text: string, resetTo?: string): string` and chained ultrathink, orchestrate, workflow glow.
- Updated ultrathink, orchestrate, and workflow matching to lowercase whitespace-delimited patterns with prose-only checks.
- Added `maskNonProse` and `keywordInProse` to skip fenced/inline code and HTML/XML segments while highlighting keywords.
- Added `KeywordHighlighter` with optional `resetTo`, and updated gradient highlighting to use masked match slicing.
- Added overlayRebuild intent, refactored forced-frame prep around base lines, and added frame dump support.
- Updated forced-render prep to preserve forced flags, allow unknown viewport mutation, and clear scrollback on resize.
- Introduced #syncChildOrder for child attach/reorder and skipped duplicate checks when after.atBottom was false.
- Added streaming-preview test scaffolding with VirtualTerminal, settleTerminal draining, and normalized row helpers.
- Preserved preexisting terminal scrollback during forced and structural rerenders in TUI.#prepareForcedRender.
- Gated historyRebuild triggers on #scrollbackHighWater in TUI.#canReplayNativeScrollbackAtCheckpoint.
- Expanded render-stress coverage with child mutations, viewport variants, and replay-mode scenario parsing.
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
- Tracked forced-render line drops with a dedicated flag and used it to force viewport repaints without treating valid empty frames as non-diffable.
- Refined render-kind selection to skip viewport repaints when appended content increases overflow and to rebuild history when line counts change while native scrollback replay is possible.
- Added a high-water preview-collapse stress operation and tightened native scrollback replay checks across geometry-mutation transitions.
- Updated shrink-path logic in `tui.ts` to choose checkpoint-based history rebuilds when bottom-anchored content changes and the viewport is not scrolled into history.
- Added a regression test ensuring a high-water preview collapse fully rebuilds native scrollback and clears stale preview rows from the buffer.
- Expanded strict scrollback stress tests with collapse operations and replay-fidelity assertions for both full-buffer and viewport matching at sampled scroll positions.
- Expanded offscreen edit handling in `TUI` to rebuild native scrollback when replay is safe and the terminal is not multiplexed.
- Marked native scrollback dirty and returned to viewport-only repainting when replay was not safe.
- Adjusted stress assertions to skip clean-buffer checks during geometry-changing operations.
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
- Asserts scrollback buffer growth does not exceed logical frame growth during dirty live rendering.
- Validates that newly appended scrollback lines match the tail of the logical frame.
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
- Replaced ad-hoc ANSI/VT stripping regexes with `stripVTControlCharacters` in status text handling and related tests.
- Updated status footer rendering to truncate using `truncateToWidth` and visible width after VT stripping.
- Extended tui cursor handling and rendering to strip markers from all lines and fit repaint/append-tail lines to width.
- Expanded deterministic render tests with overlay-aware assertions and recorded the truncation/cursor-marker behavior in changelogs.
- Added fuzzy token matching in agent-dashboard, state-manager, and tree-selector, replacing lowercased checks.
- Added search-query state and fuzzy-filter helpers to hook, oauth, and user-message selectors for query filtering.
- Updated filtered selectors to render match results, status lines, no-match text, and move selection within results.
- Added `overflowSearch` and filter state to `SelectList`, switching overflowing list matching to fuzzy checks.
- Configured `SelectList` input flow and fixed cancel so Escape/Ctrl+C closes lists when no matches exist.
- Updated changelogs and added tests for fuzzy-filter behavior in hook, oauth, user-message, and list selectors.
- Detects appended tail vs. offscreen row edits within a single render frame.
- Ambiguous appended tails now trigger a history rebuild instead of splicing stale rows into the scrollback buffer.
- Pure viewport-suffix changes above the viewport top bypass replay with a direct repaint.
- Replaced flat `protectedTools: string[]` with `ProtectedToolMatcher[]` supporting predicate functions.
- Regular file/URL `read` calls are now eligible for pruning and shake compaction.
- `read` calls whose `path` starts with `skill://` remain protected like native `skill` results.
- Added `collectToolCallsById` to correlate tool results with their originating call arguments.
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
- Stored mcpManager and localProtocolOptions on ToolSession so nested subagents inherit them without relying on process-global singletons.
- TaskTool now uses the session's localProtocolOptions and mcpManager when spawning sub-tasks, falling back to defaults if absent.
- Previously the slider started at the current cycle index, so execution would inherit whichever model drove planning.
- Now finds the `default` role in the cycle and anchors the slider there, falling back to `currentIndex` if no default exists.
- Explicit `executionModel` is set whenever the chosen tier differs from the restored cycle position, covering the case where the slider stays on `default` but planning ran on another model.
Follow-up to #1503. When an extension registered a flag whose name collides
with a value-taking built-in — e.g. plan-mode's boolean `--plan` vs the
built-in `--plan <plan-model>` selector — the extension-aware reparse still
took the built-in branch. `omp --extension plan-mode --plan "review the diff"`
consumed "review the diff" as the plan-model value, leaving parsed.messages
empty and overwriting result.plan with the prompt text. recoverFlagValue only
patched the extension flag value, not the corrupted parsed object that
applyExtensionFlags returns as initialArgs.
Fix at the source: parseArgs now checks the registered extension-flag set
BEFORE the built-in branches, so a registered flag is parsed with the
extension's semantics (boolean toggle / string value) and surfaces in
unknownFlags without consuming the following token or touching the built-in
field. This makes recoverFlagValue dead, so applyExtensionFlags is simplified
to read resolved values straight from unknownFlags.
Tests: parseArgs-level shadowing guard (boolean --plan keeps the message and
leaves result.plan unset); applyExtensionFlags message/built-in-field
preservation for colliding boolean (--plan) and string (--model) flags;
non-colliding flag-looking-value rule retained. Verified the new guards fail
without the shadowing fix.
- Changed `EmbeddingOutput` to `AsyncIterable<number[][]>` in runtime options and removed `EmbeddingRow` from the public provider contract.
- Reworked embedding result handling to stream-collect async batches into `Float32Array` rows before caching and querying.
- Updated embedding tests to provide providers as async generators yielding row batches to match the new contract.
- Dropped `onShowHotkeys` callback and its binding from `CustomEditor` and `InputController`.
- `?` now inserts a literal question mark regardless of editor state; use `/hotkeys` explicitly.
- Added regression test confirming `?` is treated as plain input when the editor is empty.
Two review fixes for the extension-flag/initial-prompt work:
1. @file ordering — `processFileArguments` runs `process.exit(1)` on a
missing/unreadable file. It had been moved after `createSession`, which
writes the terminal breadcrumb eagerly (SessionManager.create →
#newSessionSync), so `omp @missing.md "x"` left a junk session/breadcrumb
behind before exiting.
Resolve extension-registered CLI flags BEFORE creating the session: load the
session's extensions up front (new `loadSessionExtensions` helper, the single
source of createAgentSession's discovery-branch logic), build an
ExtensionFlagSink straight from the loaded extensions + runtime, re-parse
argv, then process @file args — all before any session exists. The loaded
result is handed back to createAgentSession via `preloadedExtensions` (now
checked before `disableExtensionDiscovery`, so it can't double-load) and the
same EventBus is shared, so no extra work. This keeps the P1#1 fix
(`--flag @value` is the flag's value, not a file) while failing fast with no
session side effects.
2. "Can we avoid the big list of names?" — removed the hand-maintained
`BUILTIN_FLAG_NAMES` set (and its stale "rejected at registration" doc).
`applyExtensionFlags` now always falls back to recovering a flag's value from
argv when parseArgs didn't surface it; the recovery scan mirrors parseArgs's
consumption rules (flag-looking space-form values stay their own flag) and is
a no-op for flags that were absent or already surfaced, so no list of
built-in names is needed.
Adds `ExtensionRunner.aggregateFlags` (static) so getFlags and the CLI's
pre-session sink share one implementation.
Tests: pre-session flag resolution via the exact main.ts sink pattern;
list-free recovery of an arbitrary colliding built-in (`--model`); and the
flag-looking-value rule. Verified typecheck + extension/runner/acp suites.
- Restricted `setModel` to persist settings only when `persist: true` is passed; all runtime switches (Ctrl+P, `--model`, `/model`, model picker temp selections) no longer overwrite `modelRoles.default`.
- Changed `cycleRoleModels` to accept a direction ("forward"/"backward") instead of a `temporary` flag; both directions now use `applyRoleModel` without persisting.
- Added `persist: true` exclusively to the model picker's "Set as default" action in `SelectorController`.
- Added test suite covering persistence behavior for `setModel`, `cycleRoleModels`, and `cycleModel`.
- Clarified that reordering imports, re-indenting, and other mechanical restyling must be handled by the project formatter, not hand-edited through the patch tool.