- Fixed OAuth credentials to keep unknown fields in schema while preserving existing shape checks.
- Fixed MCP OAuth IDs to be profile-scoped and avoid deleting credentials from non-active profiles.
- Fixed string-flag parsing so PROFILE_BOOTSTRAP_BOUNDARY tokens are not consumed as values.
- Fixed active-profile directory resolution to refresh after env updates so profile .env overrides apply.
Without an after-separator state, the new unknown-flag guard rejected
flag-shaped prompts (`omp -p -- --explain-this`): the loop dropped the
`--` token and then re-validated `--explain-this`, recorded it in
`unrecognizedFlags`, and exited 2.
parseArgs now flips a `sawSeparator` latch on `--` and short-circuits
the remaining tokens straight into `messages` — no built-in dispatch,
no extension match, no `@file` expansion. Two regression tests cover
the flag-shaped and `@`-prefixed cases.
Refs #2459
Bare `omp --list-models` (or any other stale/typoed --flag) was silently
consumed by `parseArgs` and the agent went on to start a real session,
connect to the configured MCP servers, and hang waiting on the model.
Any positional after the unknown flag was reinterpreted as the initial
prompt, so a documentation drift turned into an unintended LLM invocation.
`parseArgs` now tracks flag-shaped tokens that did not match any built-in
or extension-registered flag in a new `unrecognizedFlags: string[]` field,
and `reportUnrecognizedFlags` prints a clean `Error: unknown flag(s): …`
line plus the `--help` hint. `runRootCommand` invokes the helper right
after the post-extension reparse and `process.exit(2)`s before any
session, MCP, or initial-message work runs.
The validation is gated on the extension-aware reparse, so extension
flags (`--spawn-peer`, `--headless`, `--plan`, …) still pass through
the same way `applyExtensionFlags` already handles them. `-` (stdin
marker) and `--` (POSIX separator) are deliberately allowed through.
Fixes#2459
- Added `omp models` command with `ls`, `find`, `canonical`, and `refresh` actions.
- Removed top-level `--list-models` parsing from CLI args, launch, and main command flow.
- Implemented action-driven model listing with provider filtering, extension loading, and `--json` output.
- Updated unknown provider/model errors and tests to direct users to `omp models` guidance.
Upstream force-rewrote history; this branch carried old-SHA twins of the
rewritten commits. All non-goal conflicts resolved to upstream (verified
ours == old upstream tip). Goal-side reconciliation:
- cli.ts: profile bootstrap woven into the new lazy-import/resolveCliArgv
structure; worker-host entry declaration deferred until after profile
selection (pi-utils/env eagerly snapshots the agent dir .env); the
floating runCli call guarded with import.meta.main || !Bun.isMainThread
so importing runCli stays side-effect free while Worker re-entry works.
- args.ts/flag-tables.ts: kept profile/alias branches; upstream's new
repeatable --config overlay flag moved into STRING_SETTERS.
- Changelogs: upstream-released bullets deduped out of Unreleased; profile
entries restored under Unreleased.
- task/index.ts: removed duplicated validateTaskIds block from auto-merge.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Changed the `github-copilot` service provider to resolve credentials only from `COPILOT_GITHUB_TOKEN`.
- Updated CLI extra help text to document `COPILOT_GITHUB_TOKEN` as the GitHub Copilot environment variable.
- Reworded environment variable docs to reflect the revised Copilot/GitHub token usage and order.
- Migrated Effort and THINKING_EFFORTS imports to @oh-my-pi/pi-ai/effort in CLI args and launch command files.
- Split model-registry dependencies across focused @oh-my-pi/pi-ai submodules instead of the root barrel export.
- Added anonymous Perplexity authentication mode for unauthenticated web searches.
- Switched web-search setup checks to use `isExplicitlyAvailable` and removed key enforcement in doctor.
- Updated Perplexity OAuth flow to reuse auth handling for all non-key searches and anonymous responses.
- Updated CLI and provider option help text to mark the Perplexity key optional with fallback.
- Renamed `TodoWriteTool` to `TodoTool` and its source/prompt files.
- Updated tool registration, schema, renderers, and gating to `todo`.
- Adjusted cursor provider native tool names and tests to match.
- Renamed strike-animation constants and `todo-error-reminder` type.
The Anthropic web search path built request headers via buildAnthropicSearchHeaders, which never threaded model headers through buildAnthropicHeaders, so ANTHROPIC_CUSTOM_HEADERS was dropped from every web-search request regardless of mode. The streaming path's resolveAnthropicCustomHeaders also gated on isFoundryEnabled(), so users with a corporate ANTHROPIC_BASE_URL + ANTHROPIC_CUSTOM_HEADERS (e.g. X-Gateway-Key) got 401s on web_search unless they set CLAUDE_CODE_USE_FOUNDRY=true.
Loosen the resolver to also apply when ANTHROPIC_BASE_URL points to a non-Anthropic host, export the baseUrl-keyed variant, and have buildAnthropicSearchHeaders pass the resolved custom headers as modelHeaders so search and streaming paths behave identically. Stock api.anthropic.com (no Foundry) still omits the headers.
Fixes#1693
- Expanded the existing entries in docs/environment-variables.md so the override-semantics ('search-only, isolates from main ANTHROPIC_API_KEY / ANTHROPIC_BASE_URL / FOUNDRY_BASE_URL') are spelled out, and added a usage note for enterprise-gateway split routing.
- Surfaced the search-only env vars (ANTHROPIC_SEARCH_API_KEY / ANTHROPIC_SEARCH_BASE_URL / ANTHROPIC_SEARCH_MODEL) in the Anthropic provider section of docs/tools/web_search.md, where users were already looking.
- Added ANTHROPIC_SEARCH_BASE_URL alongside ANTHROPIC_SEARCH_API_KEY in 'omp --help' so the pair shows up together in the CLI env-var summary.
Fixes#1694
Follow-up to #1503. When an extension registered a flag whose name collides
with a value-taking built-in — e.g. plan-mode's boolean `--plan` vs the
built-in `--plan <plan-model>` selector — the extension-aware reparse still
took the built-in branch. `omp --extension plan-mode --plan "review the diff"`
consumed "review the diff" as the plan-model value, leaving parsed.messages
empty and overwriting result.plan with the prompt text. recoverFlagValue only
patched the extension flag value, not the corrupted parsed object that
applyExtensionFlags returns as initialArgs.
Fix at the source: parseArgs now checks the registered extension-flag set
BEFORE the built-in branches, so a registered flag is parsed with the
extension's semantics (boolean toggle / string value) and surfaces in
unknownFlags without consuming the following token or touching the built-in
field. This makes recoverFlagValue dead, so applyExtensionFlags is simplified
to read resolved values straight from unknownFlags.
Tests: parseArgs-level shadowing guard (boolean --plan keeps the message and
leaves result.plan unset); applyExtensionFlags message/built-in-field
preservation for colliding boolean (--plan) and string (--model) flags;
non-colliding flag-looking-value rule retained. Verified the new guards fail
without the shadowing fix.
Two review fixes for the extension-flag/initial-prompt work:
1. @file ordering — `processFileArguments` runs `process.exit(1)` on a
missing/unreadable file. It had been moved after `createSession`, which
writes the terminal breadcrumb eagerly (SessionManager.create →
#newSessionSync), so `omp @missing.md "x"` left a junk session/breadcrumb
behind before exiting.
Resolve extension-registered CLI flags BEFORE creating the session: load the
session's extensions up front (new `loadSessionExtensions` helper, the single
source of createAgentSession's discovery-branch logic), build an
ExtensionFlagSink straight from the loaded extensions + runtime, re-parse
argv, then process @file args — all before any session exists. The loaded
result is handed back to createAgentSession via `preloadedExtensions` (now
checked before `disableExtensionDiscovery`, so it can't double-load) and the
same EventBus is shared, so no extra work. This keeps the P1#1 fix
(`--flag @value` is the flag's value, not a file) while failing fast with no
session side effects.
2. "Can we avoid the big list of names?" — removed the hand-maintained
`BUILTIN_FLAG_NAMES` set (and its stale "rejected at registration" doc).
`applyExtensionFlags` now always falls back to recovering a flag's value from
argv when parseArgs didn't surface it; the recovery scan mirrors parseArgs's
consumption rules (flag-looking space-form values stay their own flag) and is
a no-op for flags that were absent or already surfaced, so no list of
built-in names is needed.
Adds `ExtensionRunner.aggregateFlags` (static) so getFlags and the CLI's
pre-session sink share one implementation.
Tests: pre-session flag resolution via the exact main.ts sink pattern;
list-free recovery of an arbitrary colliding built-in (`--model`); and the
flag-looking-value rule. Verified typecheck + extension/runner/acp suites.
Three issues from an adversarial review, all rooted in the startup argv parse
running before extensions load:
1. Flag-looking string values (`--name --print`): the extension-aware reparse
consumed the following token as the value, disagreeing with the startup
parse that treated `--print` as the built-in flag — so the reparse could
silently flip command shape. Extension string flags now consume a following
token only in `--flag=value` form or when it is not flag-looking; pass a
flag-looking value as `--flag=value`. Keeps both parses consistent.
2. `@file` string values (`--target @notes.md`): file args were processed from
the startup parse, which misreads the value as a file and reads it into the
prompt. processFileArguments now runs on the extension-aware parse
(initialArgs.fileArgs); pipedInput stays early for mode detection.
3. Built-in collisions: an extension flag named like a built-in (e.g. `model`)
was consumed by the built-in branch and never delivered to the runner.
registerFlag now rejects names in BUILTIN_FLAG_NAMES with a clear error
(isolated per-extension by loadExtension's try/catch).
Adds tests for all three plus the documented startup-parse misclassification.
Addresses review: a boolean flag in equals form still leaked its value. parseArgs
splices `--headless=true` into `--headless`, `true` so value-consuming flags can
pick the value up via `args[++i]`; a boolean flag sets itself without consuming
it, leaving `true` to fall through as a positional message — and since
applyExtensionFlags feeds this parse into buildInitialMessage, `omp
--headless=true "do the task"` sent `true` as the prompt.
Track the spliced value's index and, if no branch advanced past it (i.e. the
matched flag did not consume a value), drop it after the dispatch. Closes the
whole equals-form class — boolean extension flags and built-in non-consuming
flags (`--no-tools=true`, `--print=1`) alike — at the single parsing site.
Adds tests for boolean extension + built-in flags in equals form.
The `--option=value` handling splices the value into the argv to reuse the
`args[++i]` path, mutating the caller's array. The post-extension reparse in
runRootCommand then ran on that already-mutated argv, so
omp --model=sonnet --spawn-peer reviewer "review"
re-spliced `sonnet` and leaked it into the initial prompt before "review".
parseArgs now copies its input and never mutates the caller's array, so
launch, acp, and the reparse are all safe. Drops the now-redundant
`[...rawArgs]` copy at the reparse site, and adds regression coverage for the
--option=value + extension-flag combo plus input non-mutation.
- Add POSIX `--` end-of-options handling in argument parser
- Stop consuming flags after `--` and pass remaining tokens as messages
- Stop global `--profile`/alias extraction at first registered subcommand
- Add focused tests for parseArgs and profile bootstrap boundary behavior
Wafer (https://wafer.ai) exposes a single OpenAI-compatible endpoint
(`https://pass.wafer.ai/v1`) for two SKUs whose entitlement differs
server-side, so we model them as two parallel providers — mirroring the
firepass/fireworks split so a user with both subscriptions can switch
without re-pasting:
- `wafer-pass` — flat-rate. `/v1/models` is filtered to entries whose
`wafer.tier === "pass_included"`.
- `wafer-serverless` — pay-as-you-go superset of Pass.
Both issue `wfr_…` keys. `/login wafer-pass` and `/login wafer-serverless`
paste-and-validate via `/v1/models`. `WAFER_PASS_API_KEY` and
`WAFER_SERVERLESS_API_KEY` are wired through `getEnvApiKey`.
Bundled catalog:
- `wafer-pass`: GLM-5.1, Qwen3.5-397B-A17B.
- `wafer-serverless`: GLM-5.1, Qwen3.5-397B-A17B, Kimi-K2.6, Qwen3.6-35B-A3B.
Dynamic discovery via `/v1/models` overlays additional models at runtime
and folds the `wafer` envelope (tier, capabilities, cents/M pricing) into
the canonical `Model<"openai-completions">` shape. GLM-family entries
carry the zai-style thinking compat (`thinkingFormat: "zai"`,
`reasoningContentField: "reasoning_content"`) so reasoning tokens land in
the right field. Cents-per-million → dollars-per-million via /100.
Tests (`packages/ai/test/wafer.test.ts`, 5 cases): bundled catalog
contract for both providers and wire-id pass-through (case-sensitive,
no rewrite — `GLM-5.1` must round-trip verbatim or upstream 404s).
Optional `packages/ai/test/wafer.live.ts` exercises a real round-trip
against `pass.wafer.ai` when `WAFER_PASS_API_KEY` is set.
Refactored the optional-value flag handling so per-flag quirks live in
shared metadata instead of the args.ts dispatch loop. Added
OPTIONAL_FLAGS in cli/flag-tables.ts with rejectEmpty and
rejectAtPrefix controls, then updated both parseArgs and the profile
bootstrap to consult the same source of truth.
This restores the pre-refactor behavior for `--resume`, `-r`, and
`--session`: an empty-string argv token is treated as “no value
provided”, leaving resume=true and the empty string to fall through as a
positional message on the next iteration. `--list-models` intentionally
keeps its existing empty-string behavior.
Added regressions for parseArgs(["--resume", ""]), parseArgs(["-r",
""]), parseArgs(["--session", ""]), the preserved
`--list-models` empty-string behavior, and the bootstrap path where an
empty-string resume value precedes `--profile`.
Added named OMP profiles that isolate agent state (auth credentials,
sessions, settings, model cache, history, memories, blobs, plus
config root subdirs) under `~/.omp/profiles/<name>/agent/`. Activated
via `--profile <name>` or `OMP_PROFILE=<name>`; `default` maps back to
the regular `~/.omp/agent/` tree.
Added `--alias <command>` to generate a shell shortcut (e.g.
`omp-work`) that forwards `omp --profile <name>`. Detects the active
shell (bash, zsh, fish, PowerShell, pwsh), writes a wrapper into the
correct rc file, and preserves subcommands like `update`, `--version`,
and `--model` because the wrapper passes through argv unchanged.
The `--profile`/`--alias` bootstrap pre-parser lives in
`packages/coding-agent/src/cli/profile-bootstrap.ts` and runs before
any module that touches `getAgentDir()` (notably `@oh-my-pi/pi-utils/env`,
which eagerly loads `.env` from the agent directory at its own import
time). The pre-parser mirrors `parseArgs` value-consumption rules and
honors `--`, so commands like `omp --system-prompt --profile foo` pass
the literal `--profile` through as the prompt body instead of silently
activating profile `foo`.
XDG resolution for named profiles is keyed on the profile-specific
XDG path (`$XDG_*_HOME/omp/profiles/<name>`), never the base app root,
so a profile's location is decided once at first activation and stays
stable even after `omp config init-xdg` materializes the base later.
The default profile keeps its existing base-app-root check.
`setProfile(undefined)` (and `setProfile("default")`) restores the
pre-profile `PI_CODING_AGENT_DIR` snapshot taken at first activation
instead of unconditionally deleting it. `setAgentDir` refreshes the
snapshot since that call is the user explicitly redefining the
baseline.
Validation rejects profile names that match `.`/`..`, fail
`/^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/`, or hit a Windows reserved
device name (`CON`, `PRN`, `AUX`, `NUL`, `COM0-9`, `LPT0-9`, including
dotted variants like `CON.txt`) — those would let `setProfile` accept
the input only for directory creation to fail later with confusing
errors on Windows.
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
- Added a new `approvalMode` argument to CLI parsing with validation for `auto`, `prompt`, and `custom` values.
- Registered `--approval-mode` on the launch command so it appears in generated help output.
- Applied the parsed approval mode as a runtime override on `Settings`, ensuring downstream `tools.approvalMode` reads reflect the CLI value.
Re-introduces the per-tool approval system from luzidd's commit 39124f3 (which
is no longer reachable from main) and improves it before re-landing.
What's restored:
- ApprovalPolicy (allow/deny/prompt) plus DEFAULT_APPROVAL_POLICIES.
- ACTION_EXCEPTIONS registry (LSP read-only, bash critical patterns).
- getApprovalPolicy() six-level resolution order.
- ExtensionToolWrapper.execute() gate before extension handlers.
- --auto-approve / --yolo CLI flag and tools.approval.<tool> user config.
- docs/approval-mode.md user guide.
What's improved over the original:
- Replaced unchecked 'as any' casts with typed unknown narrowing helpers.
- Validate userConfig values: invalid strings, numbers, etc. fall through to
the built-in default instead of being silently honoured (typo no longer
locks a tool out or grants implicit approval).
- Expanded CRITICAL_BASH_PATTERNS: chmod -R /, chown -R /, bash <(curl ...),
writes to /etc/passwd|shadow|sudoers, shutdown/reboot/halt/init 0,
kill -9 1, nc -e / nc -c reverse shells. Pattern shapes require a
command-position boundary so 'npm run reboot-tests' and 'echo "shutdown the
queue"' don't false-positive.
- Added DEBUG_READONLY_ACTIONS exception so DAP inspection actions (threads,
stack_trace, variables, scopes, read_memory, …) auto-allow while
execution-side actions (launch, attach, continue, evaluate, write_memory,
set_breakpoint, …) still prompt.
- formatApprovalPrompt: labels mcp__<server>__<tool> calls as MCP server
tools, surfaces ssh host + command, recognises the modern § hashline header
for edit, and truncates >240-char fields so a heredoc-sized body cannot
blow out the confirmation dialog.
- Test suite grown from 40 to 57 cases — new coverage for invalid user
config, the extended critical-bash patterns, benign-keyword negatives,
debug exceptions, MCP/ssh prompt formatting, and command truncation.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts -> 57 pass
- bun x biome check . -> clean
- bun run check:ts across all 9 workspaces -> clean
Wire --hide-thinking launch flag that sets hideThinkingBlock before TUI
init. Display-only: does not disable model reasoning, just hides the
thinking output in the terminal.
- Add hideThinking to Args interface and parseArgs
- Add --hide-thinking flag definition in launch command
- Apply setting via settingsInstance.override in main
Closes#1313
Adds a new `rpc-ui` mode that extends the existing headless RPC mode with
interactive tool support (ask tool, extension UI dialogs, etc.).
In plain `rpc` mode the session has `hasUI=false` and no UI context is
wired, so interactive tools are disabled. `rpc-ui` mode sets `hasUI=true`
and wires a single shared `RpcExtensionUIContext` instance into both the
tool context store and the extension runner. Both consumers share the same
`pendingExtensionRequests` map and output closure, so `extension_ui_response`
messages received on stdin are routed to the correct waiting promise
regardless of which code path (tool or extension) created the request.
Changes:
- `args.ts`: add `rpc-ui` to the `Mode` union and the parse guard
- `launch.ts`: expose `rpc-ui` in the OCLIF flag definition and help text
- `main.ts`: propagate `rpc-ui` through all RPC-mode guard conditions and
pass `setToolUIContext` to `runRpcMode` when the mode is `rpc-ui`
- `rpc-mode.ts`: accept optional `setToolUIContext` callback; create one
shared `RpcExtensionUIContext` instance and pass it to both the tool
context store and the extension runner
- Canonicalized file and CLI defaults from `read` to `open` across tool registration and prompts.
- Added `resolveToolAlias()` and applied alias-normalized tool selection so legacy `read` maps to `open`.
- Updated runtime, UI, and export layers to treat `open` as first-class while preserving `read` compatibility.
- Renamed read prompt docs to `open.md`/`open-chunk.md` and refreshed system guidance to recommend `open`.
- Updated tool-related tests and expectations from `read` to `open` (including test fixtures and aliases).
- Consolidated fetch tool into read tool with URL reading capability and caching support.
- Removed standalone fetch tool from all agent prompts and CLI documentation.
- Extended read tool schema with timeout and raw parameters for URL fetch control.
- Added URL caching mechanism to prevent redundant network requests during read operations.
- Refactored fetch module from class-based tool to standalone executeReadUrl function.
- Updated read tool documentation to describe multi-purpose capabilities including web pages, GitHub, Stack Overflow, Wikipedia, Reddit, NPM, arXiv, blogs, and feeds.
- Added retry mechanism for benchmark tasks with separate system and retry prompt templates to improve edit success rates.
- Introduced autocorrect tracking metrics including autocorrect-free success rate and edit autocorrect counts in task and benchmark summaries.
- Refactored prompt building into modular functions (buildBenchmarkSystemPrompt, buildInitialBenchmarkPrompt, buildRetryBenchmarkPrompt) with BenchmarkPromptDelivery type for distinguishing initial and follow-up messages.
- Added session management with cache-keyed provider session IDs using xxHash64 and centralized RPC argument building via prepareBenchmarkSessionSetup.