Three issues from an adversarial review, all rooted in the startup argv parse
running before extensions load:
1. Flag-looking string values (`--name --print`): the extension-aware reparse
consumed the following token as the value, disagreeing with the startup
parse that treated `--print` as the built-in flag — so the reparse could
silently flip command shape. Extension string flags now consume a following
token only in `--flag=value` form or when it is not flag-looking; pass a
flag-looking value as `--flag=value`. Keeps both parses consistent.
2. `@file` string values (`--target @notes.md`): file args were processed from
the startup parse, which misreads the value as a file and reads it into the
prompt. processFileArguments now runs on the extension-aware parse
(initialArgs.fileArgs); pipedInput stays early for mode detection.
3. Built-in collisions: an extension flag named like a built-in (e.g. `model`)
was consumed by the built-in branch and never delivered to the runner.
registerFlag now rejects names in BUILTIN_FLAG_NAMES with a clear error
(isolated per-extension by loadExtension's try/catch).
Adds tests for all three plus the documented startup-parse misclassification.
Addresses review: a boolean flag in equals form still leaked its value. parseArgs
splices `--headless=true` into `--headless`, `true` so value-consuming flags can
pick the value up via `args[++i]`; a boolean flag sets itself without consuming
it, leaving `true` to fall through as a positional message — and since
applyExtensionFlags feeds this parse into buildInitialMessage, `omp
--headless=true "do the task"` sent `true` as the prompt.
Track the spliced value's index and, if no branch advanced past it (i.e. the
matched flag did not consume a value), drop it after the dispatch. Closes the
whole equals-form class — boolean extension flags and built-in non-consuming
flags (`--no-tools=true`, `--print=1`) alike — at the single parsing site.
Adds tests for boolean extension + built-in flags in equals form.
Addresses review: a string extension flag in equals form (--spawn-peer=reviewer)
was still leaking its value into the initial prompt. Root cause was a second,
hand-rolled argv parser in applyExtensionFlagValues that recognized only
`--flag` and `--flag value`, not `--flag=value`; it looked up the literal name
`spawn-peer=reviewer`, set nothing, and (because the reparse was gated on
"were values set") skipped the reparse entirely, so the extension-unaware
startup parse won — leaving `reviewer` as the first message. The extension
itself also never received the value.
Replace the duplicate parser with a single source of truth: extract
applyExtensionFlags() into cli/extension-flags.ts, which re-parses argv through
the same parseArgs() the startup pass uses (now seeded with the registered
flags) and pushes the resulting values onto the runner. parseArgs already
normalizes `--flag`, `--flag value`, and `--flag=value` identically, so no flag
form can be handled by one parser and missed by the other. The reparse is now
gated on registered-flag presence, not on values having been set.
Wires parseArgs's previously-unused `unknownFlags` output to the runner, and
removes the now-redundant parseArgs import from main.ts. Adds unit tests for
applyExtensionFlags across all flag forms (including equals form) plus the
no-runner / no-flags / no-args-passed gate cases.
The `--option=value` handling splices the value into the argv to reuse the
`args[++i]` path, mutating the caller's array. The post-extension reparse in
runRootCommand then ran on that already-mutated argv, so
omp --model=sonnet --spawn-peer reviewer "review"
re-spliced `sonnet` and leaked it into the initial prompt before "review".
parseArgs now copies its input and never mutates the caller's array, so
launch, acp, and the reparse are all safe. Drops the now-redundant
`[...rawArgs]` copy at the reparse site, and adds regression coverage for the
--option=value + extension-flag combo plus input non-mutation.
The root command parses argv twice — once at startup before extensions
load (so their flag set is unknown) and once after the extension runner
is ready. buildInitialMessage was reading the first, extension-unaware
parse, so a string-valued extension flag's value leaked into the prompt:
omp --spawn-peer reviewer "review the diff"
sent "reviewer" as the first message instead of "review the diff" — the
--spawn-peer token is dropped (it starts with "-"), but its bare value
is mis-read as the first positional message.
Build the initial message from args re-parsed with the extension flag map
(session.extensionRunner.getFlags()) whenever any extension flag was
applied, so the flag and its value are consumed before the prompt is
assembled. Generic across any flag-registering extension; no behavior
change when no extension flags are present.
Adds regression coverage for the parse/build pipeline: string + boolean
extension flags are consumed correctly, and the pre-fix leak the second
parse corrects is pinned.
- Updated the internal URL resolution test input pattern to use a new search phrase.
- Updated the expected assertion string to match the revised phrase in command output.
- Changed canonical config path from `keybindings.json` to `keybindings.yml`.
- Added automatic migration of legacy JSON to YAML on first load.
- Retained read support for `keybindings.yaml` without promoting it to canonical.
- Added tests for yml, yaml, and JSON migration scenarios.
- Expanded search scope handling for virtual multi-file targets and executed grep only when searchable paths existed.
- Merged `searchVirtualResources` results into main output and rendered internal URL matches as accent lines.
- Updated grouped-file output to detect URL-like paths and keep full URL headers for root grouping.
- Documented URL path/range behavior and added tests for doc routing, missing-content errors, and `omp://` expansion.
- Added virtual internal URL path resolution in `SearchTool` via `InternalUrlRouter` for in-memory search.
- Added `omp://` root expansion so `search` resolves and scans each completion target.
- Fixed `SearchTool` handling of internal URLs without `sourcePath` by returning virtual matches instead of `Path not found`.
- Extended `edit-renderer.test.ts` coverage for normalizing raw streamed text in custom text renderers.
- Renamed the partial-json helper and removed edit-mode checks so raw streamed text in `__partialJson` is converted to `input` whenever `input` is absent.
- Updated preview, render, and fallback argument preparation to use the shared helper for consistent `input` fallback handling.
- Updated interactive-mode plan review tests to capture shared fixtures, clear references, and run cleanup with explicit garbage collection before disposal.
- Increased the MCP HTTP transport test connection timeout from 200ms to 1,000ms.
- Adjusted the tool streaming command delay and tightened a start-pending-submission spy type in tests for better stability and type accuracy.
- Added helper logic to treat raw non-JSON `__partialJson` values as edit `input` for hashline and apply_patch args.
- Updated edit argument preparation, preview generation, and streaming fallback rendering to use that derived input.
- Added renderer tests verifying raw hashline and apply_patch partial streams render target paths and patch content.
- Added an amend option to the save selector that prompts for feedback and returns amend, rejected, or aborted outcomes.
- When the user chooses amend, the controller regenerated the candidate using prior rule content plus feedback before attempting to save.
- Updated prompt text and interaction tests to pass amendment context into candidate generation and validate the new save/amend flow.
- Added `/omfg <complaint>` command integration from slash registry to mode context and input handling.
- Added OMFG panel and controller for live draft streaming, up to three retries, and confirmed save flow.
- Added strict OMFG rule extraction and validation with JSON parsing, alias checks, and history-based repair.
- Fixed auto-thinking restore to preserve resolved effort instead of reverting to pending auto sessions.
- Auto classification now writes the concrete effort to the session log after the first real user turn.
- Resumed sessions restore the last resolved effort instead of reverting to pending auto.
- Added `dedupeReply` opt-out flag for ephemeral turn reply deduplication.
- Updated hashline preview diffing to route section edits through shared apply/resolve logic and partial streaming parsing.
- Allowed previews to accept live content-hash matches immediately when snapshot records are missing.
- Enabled stale-tag recovery from the snapshot store for anchor-scoped edits and added coverage for snapshot-capture behavior and session-mismatch failures.
- Added a `skipHashValidation` flag to hashline diff options and skipped section hash checks when it is set.
- Updated streaming edit preview generation to set `skipHashValidation` while streaming is active.
- Added streaming preview tests asserting stale hash errors are suppressed during streaming and surfaced after completion.
- Updated footer and status-line rendering so resolved auto-thinking status displays just the thinking level.
- Kept unresolved auto-thinking status unchanged by retaining the existing pending `auto` marker text.
- Changed tiny-device preference resolution to always default to CPU instead of platform-specific DirectML/CUDA heuristics.
- Updated settings schema, documentation, and changelog text to describe the CPU default while keeping accelerated providers behind explicit `providers.tinyModelDevice`/`PI_TINY_DEVICE` choices.
- Updated ACP thought-level resolution to use a helper that prefers session.configuredThinkingLevel and falls back to session.thinkingLevel when unavailable.
- Added a private AcpAgent method to safely compute configured thinking level without assuming a method exists on every session.
- Extended the ACP event-mapper replay test session with configuredThinkingLevel so it follows the new lookup path.
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.
The JS eval kernel's LocalModuleLoader linked and evaluated every local
module individually inside the recursive vm.SourceTextModule linker
callback. On any import cycle that re-enters Bun's node:vm linker
mid-instantiation and segfaults JSC (getImportedModule on a null record,
SIGTRAP at 0xFFFFFFFFFFFFFFF8) — e.g. `await import(".../edit/streaming.ts")`,
whose relative-import subtree is cyclic.
Construct the entire local module graph first, then drive a single
link() + evaluate() from the graph root so cyclic graphs instantiate in
one pass. The shared static linker only constructs dependencies;
external (node_modules) modules stay eagerly loaded since they carry no
imports and cannot form a cycle. Failed loads invalidate any
not-fully-evaluated modules so a retry reconstructs them.
Upstream Bun bug: https://github.com/oven-sh/bun/issues/31623
- Removed `trimTrailingPartialLine` from hashline preview path, as the streaming-tolerant parser already handles partial ops correctly; trimming stripped the sole payload of single-op patches, collapsing the preview to "No changes."
- Broadened trailing-section error suppression to apply during streaming regardless of section count, preventing transient mid-typed ops from wiping stable prior frames.
- Replaced fixed calendar timestamps with `minutesAgo()` so beam-store items stay within the 24h working-memory TTL.
- Broadened background-color assertion to match both truecolor (`48;2;`) and 256-palette (`48;5;`) terminals.
- Added persistent settings in the Providers tab for ONNX execution provider and quantization/precision, replacing env-var-only configuration.
- PI_TINY_DEVICE and PI_TINY_DTYPE env vars still override the matching setting at spawn time.
- Added tinyWorkerEnvOverlay to map settings onto worker env without clobbering explicit env vars.
- Updated docs and tests to reflect the new setting-first resolution order.
- Added 14 bundled TTSR rules (TypeScript and Rust conventions) embedded into the binary via the lowest-priority `builtin-defaults` provider.
- Extracted rule bucketing into `bucketRules` with support for `disabledRules` and `builtinRules` settings.
- Added `ttsr.builtinRules` and `ttsr.disabledRules` settings to control which rules are active per session.
- Replaced single-candidate resolution with `enumeratePythonRuntimes`, returning venv, managed env, and system interpreter in priority order.
- Availability check now probes each candidate and falls through to the first that executes, so a stale `uv`-managed Python no longer fails the whole session.
- Kernel spawn reuses the probed runtime from the availability result instead of re-resolving independently.
- Expanded test coverage for enumeration, fallback, and env isolation between candidates.
- Added `PI_TINY_DEVICE` env var to control ONNX execution provider (`gpu` default, `cpu`, `metal`, `cuda`, `dml`, `coreml`).
- Local tiny-model inference now tries accelerated GPU provider first and retries on CPU if initialization fails.
- Added `device.ts` module with normalization, preference resolution, and load-order helpers.
- Updated docs and model descriptions to drop CPU-specific language.
- Added clockwise sweeping dark segment animation to output block borders while bash/eval tool calls are pending/running.
- Changed bash renderCall to immediately render a full bordered block instead of a one-liner status preview, so silent commands show the framed block for their entire runtime.
- Added shimmerEnabled() helper and wired animate flag through OutputBlockOptions, CodeCellOptions, and shell/eval renderers.
- Added unit tests for renderSegmentTrack verifying per-segment coloring, active chip fill formatting, and active-index movement.
- Added tests for theme.getContrastFgAnsi to ensure selected foreground contrast remains clearly readable against RGB fills.
- Removed the OAuth selector keybinding test and its local auth-storage scaffolding, and documented the new model-tier/role-cycle chip-track behavior in the changelog.
- Added a shared `renderSegmentTrack` helper to render colored segment tracks with a filled active chip style.
- Updated hook selector and role-cycle status-line rendering to use the shared track renderer and aligned role-cycle output.
- Added contrast text-color logic to `Theme` so active chip labels stay legible across fill colors.
- Added `replace block N:` syntax in grammar/parser/tokenizer/types, emitting `kind: "block"` edits with empty-body checks.
- Added block-resolution in hashline apply/recovery, wiring `BlockResolver` to resolve edits with throw/drop unresolved behavior.
- Added native `blockRangeAt` support and exported block range/options types through pi-ast, pi-natives, and JS bindings.
- Added coding-agent resolver wiring plus internal visibility/exports updates and new parser, patcher, and setup-wizard tests.
- Added `SetupTab` abstractions and wired Sign-In/Web-Search tabs under the new providers onboarding scene.
- Added `providersSetupScene` and swapped `providerSetupScene` for `providers` in setup-wizard scene registration.
- Added `WebSearchTab` option loading, provider availability checks, and persisted selected search-provider preference.
- Changed sign-in flow handling to keep the scene active across providers and support cancellation during authentication.
- Changed glyph-mode picker initialization and rendering to preselect current preset and show live symbol previews.
- Updated setup wizard tests to use the `providers` scene id and cover the new web search tab path.
- Added setup wizard with provider login, glyph mode, and theme scenes shown once per setup version.
- Wired `omp setup` (no args) to trigger the wizard in a TTY; `--check`/`--json` still show help.
- Extracted `gradientEscape` and exported `PI_LOGO`/`ShineConfig` from welcome for shared use in splash/outro.
- Fixed race condition in `setSymbolPreset`/`setColorBlindMode` by tracking load request IDs.
- Mirrors the alt+p keybinding for switching provider mid-session.
- Added test verifying showModelSelector is called with temporaryOnly: true.
- Updated tips.txt to document the /switch shortcut alongside alt+p.
- Added `stripCodeBlocks` to remove fenced code blocks before titling, preventing literal noise (e.g. version strings in UI mockups) from becoming the session title.
- Added `prepareTitleInput` composing strip and truncate steps, updated `formatTitleUserMessage` to use it.
- Added unit tests for stripping logic and an integration test verifying the model never receives code block contents.
- Removed `scrubProcessEnv` from procmgr and its explicit call in cli.ts.
- Scrubbing now happens automatically when `dirs.ts` is imported, eliminating the need for a manual call at startup.
- Added `omp completions ` command generating scripts from live command/flag metadata.
- Added hidden `omp __complete` helper for dynamic model and session candidates.
- Completions never drift from the CLI: flags, enums, and subcommands are derived from static descriptors.
- Updated read tool prompt to show alternate image description when inspect_image is disabled.
- Passed INSPECT_IMAGE_ENABLED flag into prompt rendering context.
- Added test verifying description omits inspect_image references when disabled.
Allowed live input renders to bypass unknown Windows viewport deferral while preserving the default scrollback protection for background mutations.
Fixes#1550
- Added pending extraction tracking to BeamMemory and a flush loop that waits for queued extraction promises.
- Added Mnemosyne.flushExtractions export and awaited it after forced session retention in coding-agent backend.
- Implemented remember/rememberBatch extraction scheduling with runFactExtraction, skipping empty content and swallowing extraction errors.
- Replaced extractAndStoreFacts insertion logic with storeFactStrings and updated entity counts from its return value.
- Added extraction wiring tests and changelog notes for extraction shutdown and todo animation visibility behavior.
- Added todo_write strike-frame animation timing and updated execution flow after completion finalizes.
- Added strike animation cancellation in cleanup to clear todo timer and reset frames when spinner is idle.
- Removed todo-closing state and timeout handling from interactive mode and simplified empty todo-list rendering.
- Reworked todo-write to compute completion transitions, track completedTasks, and render strike-through frames.
- Added test coverage for completedTasks, theme setup, and strike-through progression at hold-frame thresholds.