- Preserved preexisting terminal scrollback during forced and structural rerenders in TUI.#prepareForcedRender.
- Gated historyRebuild triggers on #scrollbackHighWater in TUI.#canReplayNativeScrollbackAtCheckpoint.
- Expanded render-stress coverage with child mutations, viewport variants, and replay-mode scenario parsing.
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
- Tracked forced-render line drops with a dedicated flag and used it to force viewport repaints without treating valid empty frames as non-diffable.
- Refined render-kind selection to skip viewport repaints when appended content increases overflow and to rebuild history when line counts change while native scrollback replay is possible.
- Added a high-water preview-collapse stress operation and tightened native scrollback replay checks across geometry-mutation transitions.
- Updated shrink-path logic in `tui.ts` to choose checkpoint-based history rebuilds when bottom-anchored content changes and the viewport is not scrolled into history.
- Added a regression test ensuring a high-water preview collapse fully rebuilds native scrollback and clears stale preview rows from the buffer.
- Expanded strict scrollback stress tests with collapse operations and replay-fidelity assertions for both full-buffer and viewport matching at sampled scroll positions.
- Expanded offscreen edit handling in `TUI` to rebuild native scrollback when replay is safe and the terminal is not multiplexed.
- Marked native scrollback dirty and returned to viewport-only repainting when replay was not safe.
- Adjusted stress assertions to skip clean-buffer checks during geometry-changing operations.
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
- Asserts scrollback buffer growth does not exceed logical frame growth during dirty live rendering.
- Validates that newly appended scrollback lines match the tail of the logical frame.
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
- Replaced ad-hoc ANSI/VT stripping regexes with `stripVTControlCharacters` in status text handling and related tests.
- Updated status footer rendering to truncate using `truncateToWidth` and visible width after VT stripping.
- Extended tui cursor handling and rendering to strip markers from all lines and fit repaint/append-tail lines to width.
- Expanded deterministic render tests with overlay-aware assertions and recorded the truncation/cursor-marker behavior in changelogs.
- Added fuzzy token matching in agent-dashboard, state-manager, and tree-selector, replacing lowercased checks.
- Added search-query state and fuzzy-filter helpers to hook, oauth, and user-message selectors for query filtering.
- Updated filtered selectors to render match results, status lines, no-match text, and move selection within results.
- Added `overflowSearch` and filter state to `SelectList`, switching overflowing list matching to fuzzy checks.
- Configured `SelectList` input flow and fixed cancel so Escape/Ctrl+C closes lists when no matches exist.
- Updated changelogs and added tests for fuzzy-filter behavior in hook, oauth, user-message, and list selectors.
- Detects appended tail vs. offscreen row edits within a single render frame.
- Ambiguous appended tails now trigger a history rebuild instead of splicing stale rows into the scrollback buffer.
- Pure viewport-suffix changes above the viewport top bypass replay with a direct repaint.
- Replaced flat `protectedTools: string[]` with `ProtectedToolMatcher[]` supporting predicate functions.
- Regular file/URL `read` calls are now eligible for pruning and shake compaction.
- `read` calls whose `path` starts with `skill://` remain protected like native `skill` results.
- Added `collectToolCallsById` to correlate tool results with their originating call arguments.
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
- Stored mcpManager and localProtocolOptions on ToolSession so nested subagents inherit them without relying on process-global singletons.
- TaskTool now uses the session's localProtocolOptions and mcpManager when spawning sub-tasks, falling back to defaults if absent.
- Previously the slider started at the current cycle index, so execution would inherit whichever model drove planning.
- Now finds the `default` role in the cycle and anchors the slider there, falling back to `currentIndex` if no default exists.
- Explicit `executionModel` is set whenever the chosen tier differs from the restored cycle position, covering the case where the slider stays on `default` but planning ran on another model.
Follow-up to #1503. When an extension registered a flag whose name collides
with a value-taking built-in — e.g. plan-mode's boolean `--plan` vs the
built-in `--plan <plan-model>` selector — the extension-aware reparse still
took the built-in branch. `omp --extension plan-mode --plan "review the diff"`
consumed "review the diff" as the plan-model value, leaving parsed.messages
empty and overwriting result.plan with the prompt text. recoverFlagValue only
patched the extension flag value, not the corrupted parsed object that
applyExtensionFlags returns as initialArgs.
Fix at the source: parseArgs now checks the registered extension-flag set
BEFORE the built-in branches, so a registered flag is parsed with the
extension's semantics (boolean toggle / string value) and surfaces in
unknownFlags without consuming the following token or touching the built-in
field. This makes recoverFlagValue dead, so applyExtensionFlags is simplified
to read resolved values straight from unknownFlags.
Tests: parseArgs-level shadowing guard (boolean --plan keeps the message and
leaves result.plan unset); applyExtensionFlags message/built-in-field
preservation for colliding boolean (--plan) and string (--model) flags;
non-colliding flag-looking-value rule retained. Verified the new guards fail
without the shadowing fix.
- Changed `EmbeddingOutput` to `AsyncIterable<number[][]>` in runtime options and removed `EmbeddingRow` from the public provider contract.
- Reworked embedding result handling to stream-collect async batches into `Float32Array` rows before caching and querying.
- Updated embedding tests to provide providers as async generators yielding row batches to match the new contract.
- Dropped `onShowHotkeys` callback and its binding from `CustomEditor` and `InputController`.
- `?` now inserts a literal question mark regardless of editor state; use `/hotkeys` explicitly.
- Added regression test confirming `?` is treated as plain input when the editor is empty.
Two review fixes for the extension-flag/initial-prompt work:
1. @file ordering — `processFileArguments` runs `process.exit(1)` on a
missing/unreadable file. It had been moved after `createSession`, which
writes the terminal breadcrumb eagerly (SessionManager.create →
#newSessionSync), so `omp @missing.md "x"` left a junk session/breadcrumb
behind before exiting.
Resolve extension-registered CLI flags BEFORE creating the session: load the
session's extensions up front (new `loadSessionExtensions` helper, the single
source of createAgentSession's discovery-branch logic), build an
ExtensionFlagSink straight from the loaded extensions + runtime, re-parse
argv, then process @file args — all before any session exists. The loaded
result is handed back to createAgentSession via `preloadedExtensions` (now
checked before `disableExtensionDiscovery`, so it can't double-load) and the
same EventBus is shared, so no extra work. This keeps the P1#1 fix
(`--flag @value` is the flag's value, not a file) while failing fast with no
session side effects.
2. "Can we avoid the big list of names?" — removed the hand-maintained
`BUILTIN_FLAG_NAMES` set (and its stale "rejected at registration" doc).
`applyExtensionFlags` now always falls back to recovering a flag's value from
argv when parseArgs didn't surface it; the recovery scan mirrors parseArgs's
consumption rules (flag-looking space-form values stay their own flag) and is
a no-op for flags that were absent or already surfaced, so no list of
built-in names is needed.
Adds `ExtensionRunner.aggregateFlags` (static) so getFlags and the CLI's
pre-session sink share one implementation.
Tests: pre-session flag resolution via the exact main.ts sink pattern;
list-free recovery of an arbitrary colliding built-in (`--model`); and the
flag-looking-value rule. Verified typecheck + extension/runner/acp suites.
- Restricted `setModel` to persist settings only when `persist: true` is passed; all runtime switches (Ctrl+P, `--model`, `/model`, model picker temp selections) no longer overwrite `modelRoles.default`.
- Changed `cycleRoleModels` to accept a direction ("forward"/"backward") instead of a `temporary` flag; both directions now use `applyRoleModel` without persisting.
- Added `persist: true` exclusively to the model picker's "Set as default" action in `SelectorController`.
- Added test suite covering persistence behavior for `setModel`, `cycleRoleModels`, and `cycleModel`.
- Clarified that reordering imports, re-indenting, and other mechanical restyling must be handled by the project formatter, not hand-edited through the patch tool.
- Dropped `Iterable` variant; providers now return a flat row array or an async iterable only.
- Simplified `normalizeEmbeddingResult` to handle the two remaining shapes.
- Updated test to cover non-finite value rejection instead of non-array shape.
- Changed `addedCount > tailAppendCount` to `addedCount !== tailAppendCount` to also catch over-counted tails that scroll an extra row and duplicate the viewport-top line into scrollback.
- Restricted `appendFrom` in the offscreen-edit path to clean tail boundaries only; marks scrollback dirty otherwise.
- Added regression test covering the duplicate-line scenario with an offscreen edit plus mis-located tail append.
- Updated `EmbeddingOutput` to require matrix-shaped outputs and removed single-row return shapes from the provider contract.
- Normalized embedding vectors by coercing iterable rows to `Float32Array`, rejecting non-array batches and non-finite values before appending.
- Adjusted embedding result parsing to consume matrix batches from sync/async streams and destructure response data directly.
- Removed `ModelRegistry.create` and updated construction sites to use `new ModelRegistry` directly.
- Dropped the async `ConfigFile` migration warmup path and switched migration handling to the unified `#ensureMigrated` logic used by relocate and load flows.
- Adjusted model registry tests to instantiate `ModelRegistry` through the constructor.
The per-delta streaming-parse throttle (parseStreamingJsonThrottled) can
skip the last partial re-parse of a tool-call's arguments. For OpenAI
Responses and Codex streams that finalize a function_call via
response.output_item.done WITHOUT a trailing
response.function_call_arguments.done event, the handler built a fresh
toolCall with the authoritative full-buffer parse for the emitted
toolcall_end event but never wrote those args back to the block stored in
output.content. The persisted assistant message therefore retained the
last throttled partial parse (often {}), so subsequent tool execution and
history replay used stale/empty arguments.
Write the authoritative final args onto the persisted block before
dropping the streaming bookkeeping in both handlers.
Also update the stale codex idle-timeout test: bc7afc143 began stripping
partialJson/lastParseLen on finalization, but its expectation still
required partialJson: "" on a block finalized via output_item.done.
Tests: add regression coverage in both providers driving throttled deltas
that finalize via output_item.done with no args.done event, asserting the
persisted block carries the full parsed arguments and no streaming
internals. Verified both fail without the source fix.
- Defined `EmbeddingRow` and `EmbeddingOutput` in runtime options and exported them from core embeddings.
- Updated `EmbeddingProvider`, `MnemosyneEmbeddingProvider`, and `provider` runtime option types to return `EmbeddingOutput` instead of `unknown`.
- Refactored embedding result normalization to accept typed rows and sync/async batches and coerce them into validated `Float32Array` vectors.
- Consolidated `cosineSimilarity` in `vector-math`, removed duplicate local impls, and treated mismatch/non-finite as zero.
- Switched beam, query-cache, recall, shmr, and binary-vectors to consume shared `cosineSimilarity` from `vector-math`.
- Changed embeddings to parse `$env/$flag`, return `Float32Array`, lazily import model deps, and retry via `fetchWithRetry`.
- Added `@oh-my-pi/pi-utils` and replaced hardcoded FastEmbed cache paths with `getFastembedCacheDir`.
- Expanded truthy parsing for lowercase `y`, `true`, `yes`, and `on`, then updated optional-embedding tests for cache behavior.
- Updated binary-vectors and optional-embedding tests for NaN-safe cosine, `Float32Array`, and cache-dir expectations.
- Updated each package's test script to run `bun test` with the `--parallel` flag.
- Updated the tui package test command to preserve its `test/*.test.ts` file filter while enabling parallel execution.
- Mocked `watchBranch` in LSP startup test to prevent real `fs.watch` from triggering Bun SIGTRAP in parallel workers.
- Moved `refreshBaseSystemPrompt` assertion inside the `await` block where it belongs.
Codex review on #1527 flagged that the documented forms
`git+https://github.com/user/repo` and `git@github.com:user/repo` still
fell through to the npm install path. `git+https` was rejected by the
package-name validator; scp-style `git@…` passed the validator but then
resolved `actualName` via `extractPackageName` to `git` (everything before
the `@`), causing the post-install package.json lookup to fail at
`node_modules/git/package.json`.
- `parseGitUrl`: strip leading `git+` (forwarded to bun/git as-is) and
extend the protocol gate to also accept scp-like `git@host:user/repo`.
The scp form is unambiguous — no local path starts with `git@` — and
matches what `git clone` itself takes.
- `isGitSpec` now returns true for both forms, routing them through the
snapshot/diff path in `PluginManager.install` so the real package name
is discovered correctly.
- Tests: flip the two cases that asserted rejection, add ref and
`git+ssh` coverage. Verified end-to-end:
`PluginManager.install('git+https://github.com/oldschoola/omp-insights')`
installs `@oldschoola/omp-insights@1.2.3`.
Extends `omp plugin install` to accept git sources alongside npm specs and
marketplace refs. Bun's installer already understands git URLs; the blocker
was `PluginManager.install`'s strict npm-name validator and the assumption
that the actual package name could be derived from the spec.
- `git-url.ts`: `parseGitUrl` now recognizes npm-style namespaced shorthand
(`github:user/repo`, `gitlab:`, `bitbucket:`, `codeberg:`, `sourcehut:` /
`srht:`), with optional `#ref` and `.git` suffix. Exposes `isGitSpec` as
`parseGitUrl(s) !== null`. Existing protocol-URL and `git:` shorthand paths
are untouched.
- `manager.ts`: `install()` branches on `isGitSpec`. Git specs go through a
separate `validateGitSpec` (shell-metachar rejection only — `/`, `:`, `@`,
`#`, `+` are legal) and the real package name is discovered by snapshotting
`plugins/package.json` deps before `bun install` and diffing afterwards.
Falls back to value-match on force-reinstall where the key already exists.
- Help text in `plugin-cli` documents the new sources and adds a github:
example.
Smoke tested end-to-end on Windows with both forms against the test repo:
PluginManager.install('github:oldschoola/omp-insights')
PluginManager.install('https://github.com/oldschoola/omp-insights')
both resolve `@oldschoola/omp-insights@1.2.3` and write a correct lock entry.
Shell-injection probe (`github:foo/bar; rm -rf /`) is rejected.
F6: defer the JSON -> YAML migration out of ConfigFile's constructor and
add an async path so the boot sequence stops blocking the event loop on
sync I/O. New ConfigFile.tryLoadAsync/loadAsync/loadOrDefaultAsync/
getMtimeMsAsync and static ConfigFile.warmup. The migration is now
idempotent (per-process cache) so relocate() does not re-run it.
ModelRegistry.create(authStorage, modelsPath?) is a new async factory
that runs the warmup before the sync constructor's bundled-model load.
Production call sites (main.ts, sdk.ts, task/executor.ts, commit
pipelines, SDK example) all switched. Sync new ModelRegistry(...)
constructor is still supported for tests.
F2: rewrite MemorySessionStorage's mirror as { chunks: string[]; byteLen;
mtimeMs } so writeLineSync appends a single chunk in O(1) instead of
read-modify-writing the entire file (which was O(N) per append, O(N^2)
per session). statSync now reports true UTF-8 byte length instead of
character count. readTextPrefix walks chunks until the byte budget is
exhausted instead of materialising the full mirror.
Also rolls in per-package CHANGELOG entries for F1-F8.