- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Export `directoryExists` in `utils` to safely validate working directories before traversal.
- Update `SessionManager` and startup logic to fallback to the launch directory if a session's recorded working directory no longer exists.
- Add regression tests to ensure sessions now correctly adopt the launch directory instead of crashing on missing paths.
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
- Refactor deep imports by targeting specific sub-modules in `@oh-my-pi/pi-ai` to reduce barrel file overhead.
- Utilize jitless ArkType scopes in schema definitions to reduce startup JIT codegen costs by approximately 65%.
- Reorganize internal `auth-storage` exports to maintain clean boundaries between core and broker-specific functionality.
- Switched to unique, timestamped backup file paths to avoid file lock collisions during binary replacement.
- Implemented a best-effort strategy for backup deletion, allowing successful updates even if the previous process image remains locked.
- Added a cleanup routine for sweeping stale backups, including legacy files, during each update attempt.
- Support single and double quoted file paths starting with @ in CLI arguments
- Update FILE_MENTION_REGEX and extraction logic to match and resolve quoted mentions with spaces
- Add regression tests covering unquoted, single-quoted, and double-quoted file mentions
- Added the `scan` action to `ttsr` CLI with support for gitignore-aware file globbing.
- Integrated AST pre-filtering and optimized AST/regex matching to evaluate scan rules.
- Implemented file size limits, binary file detection, and custom `--no-gitignore` and `--max-bytes` flags.
- Introduced comprehensive test suites validating directory mapping, exclusions, and size-limit enforcement.
- Added a top-level `ttsr` CLI command with `list` and `test` actions.
- Added snippet input handling for inline text, `--file` path, and stdin via `--file -`.
- Added test-mode context inference and result output for matched and unmatched rules.
- Added CLI tests for source inference, explicit source overrides, JSON output, and list mode.
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
- Added a new `--advisor` flag to argument parsing and launch flag definitions.
- Applied the parsed `advisor` flag as a runtime-only override of `advisor.enabled` during startup.
- Updated advisor-related docs and added parseArgs tests for `--advisor` behavior and position handling.
- Fixed OAuth credentials to keep unknown fields in schema while preserving existing shape checks.
- Fixed MCP OAuth IDs to be profile-scoped and avoid deleting credentials from non-active profiles.
- Fixed string-flag parsing so PROFILE_BOOTSTRAP_BOUNDARY tokens are not consumed as values.
- Fixed active-profile directory resolution to refresh after env updates so profile .env overrides apply.
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.
Without an after-separator state, the new unknown-flag guard rejected
flag-shaped prompts (`omp -p -- --explain-this`): the loop dropped the
`--` token and then re-validated `--explain-this`, recorded it in
`unrecognizedFlags`, and exited 2.
parseArgs now flips a `sawSeparator` latch on `--` and short-circuits
the remaining tokens straight into `messages` — no built-in dispatch,
no extension match, no `@file` expansion. Two regression tests cover
the flag-shaped and `@`-prefixed cases.
Refs #2459
- Added session-domain modules and exports for session-entries, context, listing, loader, and migrations.
- Changed persistence to async append writes plus writeTextAtomic, removing sync line APIs.
- Added compaction-aware session context rebuild with dangling tool-call cleanup.
- Added resumable session resolution with status inference, id/stem/suffix matching, and backup recovery.
Bare `omp --list-models` (or any other stale/typoed --flag) was silently
consumed by `parseArgs` and the agent went on to start a real session,
connect to the configured MCP servers, and hang waiting on the model.
Any positional after the unknown flag was reinterpreted as the initial
prompt, so a documentation drift turned into an unintended LLM invocation.
`parseArgs` now tracks flag-shaped tokens that did not match any built-in
or extension-registered flag in a new `unrecognizedFlags: string[]` field,
and `reportUnrecognizedFlags` prints a clean `Error: unknown flag(s): …`
line plus the `--help` hint. `runRootCommand` invokes the helper right
after the post-extension reparse and `process.exit(2)`s before any
session, MCP, or initial-message work runs.
The validation is gated on the extension-aware reparse, so extension
flags (`--spawn-peer`, `--headless`, `--plan`, …) still pass through
the same way `applyExtensionFlags` already handles them. `-` (stdin
marker) and `--` (POSIX separator) are deliberately allowed through.
Fixes#2459
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
- Added `omp models` command with `ls`, `find`, `canonical`, and `refresh` actions.
- Removed top-level `--list-models` parsing from CLI args, launch, and main command flow.
- Implemented action-driven model listing with provider filtering, extension loading, and `--json` output.
- Updated unknown provider/model errors and tests to direct users to `omp models` guidance.
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.
- Added usage snapshot persistence in sqlite with hour-bucket upsert behavior.
- Added listUsageHistory query support with optional provider and sinceMs filters.
- Added usage CLI history mode with `--history` and `--days` and trend rendering.
- Added changelog documentation for usage trend inspection and no-history exit behavior.
- Removed `resume` from task params and schema, requiring agent and assignment inputs.
- Dropped resume continuation paths in task execution and call rendering, always spawning a new agent.
- Removed the `irc.enabled` setting and computed IRC availability by task-depth rules.
- Updated task follow-up guidance to use IRC messaging/history links instead of `task(resume:)`.
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
Upstream force-rewrote history; this branch carried old-SHA twins of the
rewritten commits. All non-goal conflicts resolved to upstream (verified
ours == old upstream tip). Goal-side reconciliation:
- cli.ts: profile bootstrap woven into the new lazy-import/resolveCliArgv
structure; worker-host entry declaration deferred until after profile
selection (pi-utils/env eagerly snapshots the agent dir .env); the
floating runCli call guarded with import.meta.main || !Bun.isMainThread
so importing runCli stays side-effect free while Worker re-entry works.
- args.ts/flag-tables.ts: kept profile/alias branches; upstream's new
repeatable --config overlay flag moved into STRING_SETTERS.
- Changelogs: upstream-released bullets deduped out of Unreleased; profile
entries restored under Unreleased.
- task/index.ts: removed duplicated validateTaskIds block from auto-merge.
- Updated usage-window reporting to track remaining account quota instead of required-account counts.
- Replaced the computed "needed" metric with a non-negative remaining-quota value derived from each window's total minus used fraction.
- Changed usage output lines from "need" to "capacity" and now show used/total accounts with remaining quota multiplier.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Added per-account usage reporting in the `omp usage` command.
- Added `provider`, `json`, and `redact` options to customize usage output.
- Updated CLI wiring to route usage commands to the new per-account behavior.
- Added per-account usage reporting in the `omp usage` command.
- Added `provider`, `json`, and `redact` options to customize usage output.
- Updated CLI wiring to route usage commands to the new per-account behavior.
- Replaced canonical-row resolution with getCanonicalModelSelections in model lists and selector flow.
- Hydrated model selector state from registry on construction and kept cached selections during refresh.
- Preserved highlighted and cached model selection when offline refresh completed or reordered models.
- Added parity checks between getCanonicalModelSelections and resolveCanonicalModel via registry tests.