- Replaced composed keyword decorators in `CustomEditor.decorateText` with `highlightMagicKeywords`.
- Adjusted `UserMessageComponent` rendering to use `highlightMagicKeywords` with `keywordReset` for consistent foreground.
- Added `highlightMagicKeywords(text: string, resetTo?: string): string` and chained ultrathink, orchestrate, workflow glow.
- Updated ultrathink, orchestrate, and workflow matching to lowercase whitespace-delimited patterns with prose-only checks.
- Added `maskNonProse` and `keywordInProse` to skip fenced/inline code and HTML/XML segments while highlighting keywords.
- Added `KeywordHighlighter` with optional `resetTo`, and updated gradient highlighting to use masked match slicing.
- Added overlayRebuild intent, refactored forced-frame prep around base lines, and added frame dump support.
- Updated forced-render prep to preserve forced flags, allow unknown viewport mutation, and clear scrollback on resize.
- Introduced #syncChildOrder for child attach/reorder and skipped duplicate checks when after.atBottom was false.
- Added streaming-preview test scaffolding with VirtualTerminal, settleTerminal draining, and normalized row helpers.
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
- Replaced ad-hoc ANSI/VT stripping regexes with `stripVTControlCharacters` in status text handling and related tests.
- Updated status footer rendering to truncate using `truncateToWidth` and visible width after VT stripping.
- Extended tui cursor handling and rendering to strip markers from all lines and fit repaint/append-tail lines to width.
- Expanded deterministic render tests with overlay-aware assertions and recorded the truncation/cursor-marker behavior in changelogs.
- Added fuzzy token matching in agent-dashboard, state-manager, and tree-selector, replacing lowercased checks.
- Added search-query state and fuzzy-filter helpers to hook, oauth, and user-message selectors for query filtering.
- Updated filtered selectors to render match results, status lines, no-match text, and move selection within results.
- Added `overflowSearch` and filter state to `SelectList`, switching overflowing list matching to fuzzy checks.
- Configured `SelectList` input flow and fixed cancel so Escape/Ctrl+C closes lists when no matches exist.
- Updated changelogs and added tests for fuzzy-filter behavior in hook, oauth, user-message, and list selectors.
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
- Stored mcpManager and localProtocolOptions on ToolSession so nested subagents inherit them without relying on process-global singletons.
- TaskTool now uses the session's localProtocolOptions and mcpManager when spawning sub-tasks, falling back to defaults if absent.
- Previously the slider started at the current cycle index, so execution would inherit whichever model drove planning.
- Now finds the `default` role in the cycle and anchors the slider there, falling back to `currentIndex` if no default exists.
- Explicit `executionModel` is set whenever the chosen tier differs from the restored cycle position, covering the case where the slider stays on `default` but planning ran on another model.
Follow-up to #1503. When an extension registered a flag whose name collides
with a value-taking built-in — e.g. plan-mode's boolean `--plan` vs the
built-in `--plan <plan-model>` selector — the extension-aware reparse still
took the built-in branch. `omp --extension plan-mode --plan "review the diff"`
consumed "review the diff" as the plan-model value, leaving parsed.messages
empty and overwriting result.plan with the prompt text. recoverFlagValue only
patched the extension flag value, not the corrupted parsed object that
applyExtensionFlags returns as initialArgs.
Fix at the source: parseArgs now checks the registered extension-flag set
BEFORE the built-in branches, so a registered flag is parsed with the
extension's semantics (boolean toggle / string value) and surfaces in
unknownFlags without consuming the following token or touching the built-in
field. This makes recoverFlagValue dead, so applyExtensionFlags is simplified
to read resolved values straight from unknownFlags.
Tests: parseArgs-level shadowing guard (boolean --plan keeps the message and
leaves result.plan unset); applyExtensionFlags message/built-in-field
preservation for colliding boolean (--plan) and string (--model) flags;
non-colliding flag-looking-value rule retained. Verified the new guards fail
without the shadowing fix.
- Dropped `onShowHotkeys` callback and its binding from `CustomEditor` and `InputController`.
- `?` now inserts a literal question mark regardless of editor state; use `/hotkeys` explicitly.
- Added regression test confirming `?` is treated as plain input when the editor is empty.
Two review fixes for the extension-flag/initial-prompt work:
1. @file ordering — `processFileArguments` runs `process.exit(1)` on a
missing/unreadable file. It had been moved after `createSession`, which
writes the terminal breadcrumb eagerly (SessionManager.create →
#newSessionSync), so `omp @missing.md "x"` left a junk session/breadcrumb
behind before exiting.
Resolve extension-registered CLI flags BEFORE creating the session: load the
session's extensions up front (new `loadSessionExtensions` helper, the single
source of createAgentSession's discovery-branch logic), build an
ExtensionFlagSink straight from the loaded extensions + runtime, re-parse
argv, then process @file args — all before any session exists. The loaded
result is handed back to createAgentSession via `preloadedExtensions` (now
checked before `disableExtensionDiscovery`, so it can't double-load) and the
same EventBus is shared, so no extra work. This keeps the P1#1 fix
(`--flag @value` is the flag's value, not a file) while failing fast with no
session side effects.
2. "Can we avoid the big list of names?" — removed the hand-maintained
`BUILTIN_FLAG_NAMES` set (and its stale "rejected at registration" doc).
`applyExtensionFlags` now always falls back to recovering a flag's value from
argv when parseArgs didn't surface it; the recovery scan mirrors parseArgs's
consumption rules (flag-looking space-form values stay their own flag) and is
a no-op for flags that were absent or already surfaced, so no list of
built-in names is needed.
Adds `ExtensionRunner.aggregateFlags` (static) so getFlags and the CLI's
pre-session sink share one implementation.
Tests: pre-session flag resolution via the exact main.ts sink pattern;
list-free recovery of an arbitrary colliding built-in (`--model`); and the
flag-looking-value rule. Verified typecheck + extension/runner/acp suites.
- Restricted `setModel` to persist settings only when `persist: true` is passed; all runtime switches (Ctrl+P, `--model`, `/model`, model picker temp selections) no longer overwrite `modelRoles.default`.
- Changed `cycleRoleModels` to accept a direction ("forward"/"backward") instead of a `temporary` flag; both directions now use `applyRoleModel` without persisting.
- Added `persist: true` exclusively to the model picker's "Set as default" action in `SelectorController`.
- Added test suite covering persistence behavior for `setModel`, `cycleRoleModels`, and `cycleModel`.
- Removed `ModelRegistry.create` and updated construction sites to use `new ModelRegistry` directly.
- Dropped the async `ConfigFile` migration warmup path and switched migration handling to the unified `#ensureMigrated` logic used by relocate and load flows.
- Adjusted model registry tests to instantiate `ModelRegistry` through the constructor.
- Defined `EmbeddingRow` and `EmbeddingOutput` in runtime options and exported them from core embeddings.
- Updated `EmbeddingProvider`, `MnemosyneEmbeddingProvider`, and `provider` runtime option types to return `EmbeddingOutput` instead of `unknown`.
- Refactored embedding result normalization to accept typed rows and sync/async batches and coerce them into validated `Float32Array` vectors.
- Updated each package's test script to run `bun test` with the `--parallel` flag.
- Updated the tui package test command to preserve its `test/*.test.ts` file filter while enabling parallel execution.
- Mocked `watchBranch` in LSP startup test to prevent real `fs.watch` from triggering Bun SIGTRAP in parallel workers.
- Moved `refreshBaseSystemPrompt` assertion inside the `await` block where it belongs.
Codex review on #1527 flagged that the documented forms
`git+https://github.com/user/repo` and `git@github.com:user/repo` still
fell through to the npm install path. `git+https` was rejected by the
package-name validator; scp-style `git@…` passed the validator but then
resolved `actualName` via `extractPackageName` to `git` (everything before
the `@`), causing the post-install package.json lookup to fail at
`node_modules/git/package.json`.
- `parseGitUrl`: strip leading `git+` (forwarded to bun/git as-is) and
extend the protocol gate to also accept scp-like `git@host:user/repo`.
The scp form is unambiguous — no local path starts with `git@` — and
matches what `git clone` itself takes.
- `isGitSpec` now returns true for both forms, routing them through the
snapshot/diff path in `PluginManager.install` so the real package name
is discovered correctly.
- Tests: flip the two cases that asserted rejection, add ref and
`git+ssh` coverage. Verified end-to-end:
`PluginManager.install('git+https://github.com/oldschoola/omp-insights')`
installs `@oldschoola/omp-insights@1.2.3`.
Extends `omp plugin install` to accept git sources alongside npm specs and
marketplace refs. Bun's installer already understands git URLs; the blocker
was `PluginManager.install`'s strict npm-name validator and the assumption
that the actual package name could be derived from the spec.
- `git-url.ts`: `parseGitUrl` now recognizes npm-style namespaced shorthand
(`github:user/repo`, `gitlab:`, `bitbucket:`, `codeberg:`, `sourcehut:` /
`srht:`), with optional `#ref` and `.git` suffix. Exposes `isGitSpec` as
`parseGitUrl(s) !== null`. Existing protocol-URL and `git:` shorthand paths
are untouched.
- `manager.ts`: `install()` branches on `isGitSpec`. Git specs go through a
separate `validateGitSpec` (shell-metachar rejection only — `/`, `:`, `@`,
`#`, `+` are legal) and the real package name is discovered by snapshotting
`plugins/package.json` deps before `bun install` and diffing afterwards.
Falls back to value-match on force-reinstall where the key already exists.
- Help text in `plugin-cli` documents the new sources and adds a github:
example.
Smoke tested end-to-end on Windows with both forms against the test repo:
PluginManager.install('github:oldschoola/omp-insights')
PluginManager.install('https://github.com/oldschoola/omp-insights')
both resolve `@oldschoola/omp-insights@1.2.3` and write a correct lock entry.
Shell-injection probe (`github:foo/bar; rm -rf /`) is rejected.
F6: defer the JSON -> YAML migration out of ConfigFile's constructor and
add an async path so the boot sequence stops blocking the event loop on
sync I/O. New ConfigFile.tryLoadAsync/loadAsync/loadOrDefaultAsync/
getMtimeMsAsync and static ConfigFile.warmup. The migration is now
idempotent (per-process cache) so relocate() does not re-run it.
ModelRegistry.create(authStorage, modelsPath?) is a new async factory
that runs the warmup before the sync constructor's bundled-model load.
Production call sites (main.ts, sdk.ts, task/executor.ts, commit
pipelines, SDK example) all switched. Sync new ModelRegistry(...)
constructor is still supported for tests.
F2: rewrite MemorySessionStorage's mirror as { chunks: string[]; byteLen;
mtimeMs } so writeLineSync appends a single chunk in O(1) instead of
read-modify-writing the entire file (which was O(N) per append, O(N^2)
per session). statSync now reports true UTF-8 byte length instead of
character count. readTextPrefix walks chunks until the byte budget is
exhausted instead of materialising the full mirror.
Also rolls in per-package CHANGELOG entries for F1-F8.
F3 (agent): mutate state.messages and state.pendingToolCalls in place on
appendMessage/popMessage/clearMessages/reset/tool_execution_start/_end
instead of allocating a fresh array/Set on every transition. Subscribers
that capture state.messages by reference now observe updates directly.
Public type signature unchanged.
F5 (ai): add parseStreamingJsonThrottled to utils/json-parse — a per-delta
wrapper around parseStreamingJson that skips the re-parse until the
buffer has grown by minGrowthBytes (default 256). Wired into every
provider's tool-call argument accumulator (anthropic, amazon-bedrock,
openai-completions, openai-codex-responses, openai-responses-shared) so
per-delta cost becomes O(N) in total buffer length instead of O(N²).
Every provider's toolcall_end still runs a final unthrottled parse, so
the published block.arguments is unchanged.
F8 (coding-agent): drop the per-delta structuredClone of streaming tool
arguments in ToolExecutionComponent.updateArgs. event-controller.ts and
ui-helpers.ts already spread their input into a fresh object on each
delta, so cloning here was dead work on the rendering hot path. Added a
reference-equality short-circuit so repeat calls with the same args
object skip the preview-diff and display refresh.
F1 (coding-agent): close the withFileLock mkdir-vs-writeLockInfo race that
let a losing contender wipe the winner's freshly-created lock directory.
Every lock now carries a per-process UUID token; releaseLock verifies the
token before fs.rm, and isLockStale no longer treats an info-less but
fresh dir (or a dir that vanished mid-check) as stale.
F4 (coding-agent): sanitize tabs and truncate oversized error strings in
formatErrorMessage so error renderings that embed file content
(apply_patch, hashline, etc.) cannot break terminal alignment or
overflow the line width.
F7 (ai): support named-tool routing on Google providers. Widens
GoogleSharedStreamOptions.toolChoice and GoogleGeminiCliOptions.toolChoice
to accept { mode: 'ANY'; allowedFunctionNames }. mapGoogleToolChoice
now converts ToolChoice { type: 'tool'|'function', name } to the wire
shape (mirroring mapAnthropicToolChoice). buildGoogleGenerateContentParams
and the gemini-cli request serializer honor the allow-list.
The previous collision guard threw in registerFlag, which broke loading the
bundled plan-mode example extension (it registers `--plan`, also a built-in)
even when `--plan` was never passed — making a documented extension unusable.
Registering a built-in-named flag is a supported pattern: `--plan` is both the
built-in plan-model selector and plan-mode's boolean mode toggle, and the value
must reach both. So instead of rejecting, preserve delivery: remove the guard,
and in applyExtensionFlags recover a colliding flag's value from argv
(resolveCollidingFlag) when parseArgs routed it to the built-in branch and it
never reached unknownFlags. Non-colliding flags are unchanged (peer-* etc.).
Verified the real bundled plan-mode.ts loads with --plan registered and
delivered; replaced the reject-test with a loads-without-throwing regression
plus colliding-flag delivery coverage.
Three issues from an adversarial review, all rooted in the startup argv parse
running before extensions load:
1. Flag-looking string values (`--name --print`): the extension-aware reparse
consumed the following token as the value, disagreeing with the startup
parse that treated `--print` as the built-in flag — so the reparse could
silently flip command shape. Extension string flags now consume a following
token only in `--flag=value` form or when it is not flag-looking; pass a
flag-looking value as `--flag=value`. Keeps both parses consistent.
2. `@file` string values (`--target @notes.md`): file args were processed from
the startup parse, which misreads the value as a file and reads it into the
prompt. processFileArguments now runs on the extension-aware parse
(initialArgs.fileArgs); pipedInput stays early for mode detection.
3. Built-in collisions: an extension flag named like a built-in (e.g. `model`)
was consumed by the built-in branch and never delivered to the runner.
registerFlag now rejects names in BUILTIN_FLAG_NAMES with a clear error
(isolated per-extension by loadExtension's try/catch).
Adds tests for all three plus the documented startup-parse misclassification.
Addresses review: a boolean flag in equals form still leaked its value. parseArgs
splices `--headless=true` into `--headless`, `true` so value-consuming flags can
pick the value up via `args[++i]`; a boolean flag sets itself without consuming
it, leaving `true` to fall through as a positional message — and since
applyExtensionFlags feeds this parse into buildInitialMessage, `omp
--headless=true "do the task"` sent `true` as the prompt.
Track the spliced value's index and, if no branch advanced past it (i.e. the
matched flag did not consume a value), drop it after the dispatch. Closes the
whole equals-form class — boolean extension flags and built-in non-consuming
flags (`--no-tools=true`, `--print=1`) alike — at the single parsing site.
Adds tests for boolean extension + built-in flags in equals form.
Addresses review: a string extension flag in equals form (--spawn-peer=reviewer)
was still leaking its value into the initial prompt. Root cause was a second,
hand-rolled argv parser in applyExtensionFlagValues that recognized only
`--flag` and `--flag value`, not `--flag=value`; it looked up the literal name
`spawn-peer=reviewer`, set nothing, and (because the reparse was gated on
"were values set") skipped the reparse entirely, so the extension-unaware
startup parse won — leaving `reviewer` as the first message. The extension
itself also never received the value.
Replace the duplicate parser with a single source of truth: extract
applyExtensionFlags() into cli/extension-flags.ts, which re-parses argv through
the same parseArgs() the startup pass uses (now seeded with the registered
flags) and pushes the resulting values onto the runner. parseArgs already
normalizes `--flag`, `--flag value`, and `--flag=value` identically, so no flag
form can be handled by one parser and missed by the other. The reparse is now
gated on registered-flag presence, not on values having been set.
Wires parseArgs's previously-unused `unknownFlags` output to the runner, and
removes the now-redundant parseArgs import from main.ts. Adds unit tests for
applyExtensionFlags across all flag forms (including equals form) plus the
no-runner / no-flags / no-args-passed gate cases.
The `--option=value` handling splices the value into the argv to reuse the
`args[++i]` path, mutating the caller's array. The post-extension reparse in
runRootCommand then ran on that already-mutated argv, so
omp --model=sonnet --spawn-peer reviewer "review"
re-spliced `sonnet` and leaked it into the initial prompt before "review".
parseArgs now copies its input and never mutates the caller's array, so
launch, acp, and the reparse are all safe. Drops the now-redundant
`[...rawArgs]` copy at the reparse site, and adds regression coverage for the
--option=value + extension-flag combo plus input non-mutation.
The root command parses argv twice — once at startup before extensions
load (so their flag set is unknown) and once after the extension runner
is ready. buildInitialMessage was reading the first, extension-unaware
parse, so a string-valued extension flag's value leaked into the prompt:
omp --spawn-peer reviewer "review the diff"
sent "reviewer" as the first message instead of "review the diff" — the
--spawn-peer token is dropped (it starts with "-"), but its bare value
is mis-read as the first positional message.
Build the initial message from args re-parsed with the extension flag map
(session.extensionRunner.getFlags()) whenever any extension flag was
applied, so the flag and its value are consumed before the prompt is
assembled. Generic across any flag-registering extension; no behavior
change when no extension flags are present.
Adds regression coverage for the parse/build pipeline: string + boolean
extension flags are consumed correctly, and the pre-fix leak the second
parse corrects is pinned.
- Updated the internal URL resolution test input pattern to use a new search phrase.
- Updated the expected assertion string to match the revised phrase in command output.
- Changed canonical config path from `keybindings.json` to `keybindings.yml`.
- Added automatic migration of legacy JSON to YAML on first load.
- Retained read support for `keybindings.yaml` without promoting it to canonical.
- Added tests for yml, yaml, and JSON migration scenarios.
- Expanded search scope handling for virtual multi-file targets and executed grep only when searchable paths existed.
- Merged `searchVirtualResources` results into main output and rendered internal URL matches as accent lines.
- Updated grouped-file output to detect URL-like paths and keep full URL headers for root grouping.
- Documented URL path/range behavior and added tests for doc routing, missing-content errors, and `omp://` expansion.
- Added virtual internal URL path resolution in `SearchTool` via `InternalUrlRouter` for in-memory search.
- Added `omp://` root expansion so `search` resolves and scans each completion target.
- Fixed `SearchTool` handling of internal URLs without `sourcePath` by returning virtual matches instead of `Path not found`.
- Extended `edit-renderer.test.ts` coverage for normalizing raw streamed text in custom text renderers.
- Renamed the partial-json helper and removed edit-mode checks so raw streamed text in `__partialJson` is converted to `input` whenever `input` is absent.
- Updated preview, render, and fallback argument preparation to use the shared helper for consistent `input` fallback handling.
- Updated interactive-mode plan review tests to capture shared fixtures, clear references, and run cleanup with explicit garbage collection before disposal.
- Increased the MCP HTTP transport test connection timeout from 200ms to 1,000ms.
- Adjusted the tool streaming command delay and tightened a start-pending-submission spy type in tests for better stability and type accuracy.
- Added helper logic to treat raw non-JSON `__partialJson` values as edit `input` for hashline and apply_patch args.
- Updated edit argument preparation, preview generation, and streaming fallback rendering to use that derived input.
- Added renderer tests verifying raw hashline and apply_patch partial streams render target paths and patch content.
- Added an amend option to the save selector that prompts for feedback and returns amend, rejected, or aborted outcomes.
- When the user chooses amend, the controller regenerated the candidate using prior rule content plus feedback before attempting to save.
- Updated prompt text and interaction tests to pass amendment context into candidate generation and validate the new save/amend flow.