resolveModels("all") expanded the full TINY_LOCAL_MODELS registry, which now
includes the qwen3-1.7b entry marked unsupportedReason. loadPipeline() throws
for such specs, so the download worker reported it as failed and the bulk
command exited with "One or more tiny title models failed to download" even
when every usable model downloaded. Filter unsupported specs out of the `all`
prefetch path; explicit single-model requests are unchanged. Addresses the
unaddressed Codex P2 on PR #3133.
The session picker auto-switched into all-projects scope whenever the
current cwd had no sessions, so /resume from a fresh project silently
surfaced every other project's history. The empty-folder hint already
tells users to Tab into all-projects, but the auto-switch made it
unreachable. Both call sites (the /resume slash command and `omp
--resume` startup) now always open in folder scope; `omp --resume`
keeps the global probe only to early-exit with 'No sessions found' when
nothing exists anywhere. The component-level `startInAllScope` option
is deleted along with its callers.
Fixes#3099
- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Export `directoryExists` in `utils` to safely validate working directories before traversal.
- Update `SessionManager` and startup logic to fallback to the launch directory if a session's recorded working directory no longer exists.
- Add regression tests to ensure sessions now correctly adopt the launch directory instead of crashing on missing paths.
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
- Refactor deep imports by targeting specific sub-modules in `@oh-my-pi/pi-ai` to reduce barrel file overhead.
- Utilize jitless ArkType scopes in schema definitions to reduce startup JIT codegen costs by approximately 65%.
- Reorganize internal `auth-storage` exports to maintain clean boundaries between core and broker-specific functionality.
- Switched to unique, timestamped backup file paths to avoid file lock collisions during binary replacement.
- Implemented a best-effort strategy for backup deletion, allowing successful updates even if the previous process image remains locked.
- Added a cleanup routine for sweeping stale backups, including legacy files, during each update attempt.
- Support single and double quoted file paths starting with @ in CLI arguments
- Update FILE_MENTION_REGEX and extraction logic to match and resolve quoted mentions with spaces
- Add regression tests covering unquoted, single-quoted, and double-quoted file mentions
- Added the `scan` action to `ttsr` CLI with support for gitignore-aware file globbing.
- Integrated AST pre-filtering and optimized AST/regex matching to evaluate scan rules.
- Implemented file size limits, binary file detection, and custom `--no-gitignore` and `--max-bytes` flags.
- Introduced comprehensive test suites validating directory mapping, exclusions, and size-limit enforcement.
- Added a top-level `ttsr` CLI command with `list` and `test` actions.
- Added snippet input handling for inline text, `--file` path, and stdin via `--file -`.
- Added test-mode context inference and result output for matched and unmatched rules.
- Added CLI tests for source inference, explicit source overrides, JSON output, and list mode.
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
- Added a new `--advisor` flag to argument parsing and launch flag definitions.
- Applied the parsed `advisor` flag as a runtime-only override of `advisor.enabled` during startup.
- Updated advisor-related docs and added parseArgs tests for `--advisor` behavior and position handling.
- Fixed OAuth credentials to keep unknown fields in schema while preserving existing shape checks.
- Fixed MCP OAuth IDs to be profile-scoped and avoid deleting credentials from non-active profiles.
- Fixed string-flag parsing so PROFILE_BOOTSTRAP_BOUNDARY tokens are not consumed as values.
- Fixed active-profile directory resolution to refresh after env updates so profile .env overrides apply.
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.
Without an after-separator state, the new unknown-flag guard rejected
flag-shaped prompts (`omp -p -- --explain-this`): the loop dropped the
`--` token and then re-validated `--explain-this`, recorded it in
`unrecognizedFlags`, and exited 2.
parseArgs now flips a `sawSeparator` latch on `--` and short-circuits
the remaining tokens straight into `messages` — no built-in dispatch,
no extension match, no `@file` expansion. Two regression tests cover
the flag-shaped and `@`-prefixed cases.
Refs #2459
- Added session-domain modules and exports for session-entries, context, listing, loader, and migrations.
- Changed persistence to async append writes plus writeTextAtomic, removing sync line APIs.
- Added compaction-aware session context rebuild with dangling tool-call cleanup.
- Added resumable session resolution with status inference, id/stem/suffix matching, and backup recovery.
Bare `omp --list-models` (or any other stale/typoed --flag) was silently
consumed by `parseArgs` and the agent went on to start a real session,
connect to the configured MCP servers, and hang waiting on the model.
Any positional after the unknown flag was reinterpreted as the initial
prompt, so a documentation drift turned into an unintended LLM invocation.
`parseArgs` now tracks flag-shaped tokens that did not match any built-in
or extension-registered flag in a new `unrecognizedFlags: string[]` field,
and `reportUnrecognizedFlags` prints a clean `Error: unknown flag(s): …`
line plus the `--help` hint. `runRootCommand` invokes the helper right
after the post-extension reparse and `process.exit(2)`s before any
session, MCP, or initial-message work runs.
The validation is gated on the extension-aware reparse, so extension
flags (`--spawn-peer`, `--headless`, `--plan`, …) still pass through
the same way `applyExtensionFlags` already handles them. `-` (stdin
marker) and `--` (POSIX separator) are deliberately allowed through.
Fixes#2459
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
- Added `omp models` command with `ls`, `find`, `canonical`, and `refresh` actions.
- Removed top-level `--list-models` parsing from CLI args, launch, and main command flow.
- Implemented action-driven model listing with provider filtering, extension loading, and `--json` output.
- Updated unknown provider/model errors and tests to direct users to `omp models` guidance.
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.
- Added usage snapshot persistence in sqlite with hour-bucket upsert behavior.
- Added listUsageHistory query support with optional provider and sinceMs filters.
- Added usage CLI history mode with `--history` and `--days` and trend rendering.
- Added changelog documentation for usage trend inspection and no-history exit behavior.
- Removed `resume` from task params and schema, requiring agent and assignment inputs.
- Dropped resume continuation paths in task execution and call rendering, always spawning a new agent.
- Removed the `irc.enabled` setting and computed IRC availability by task-depth rules.
- Updated task follow-up guidance to use IRC messaging/history links instead of `task(resume:)`.
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
Upstream force-rewrote history; this branch carried old-SHA twins of the
rewritten commits. All non-goal conflicts resolved to upstream (verified
ours == old upstream tip). Goal-side reconciliation:
- cli.ts: profile bootstrap woven into the new lazy-import/resolveCliArgv
structure; worker-host entry declaration deferred until after profile
selection (pi-utils/env eagerly snapshots the agent dir .env); the
floating runCli call guarded with import.meta.main || !Bun.isMainThread
so importing runCli stays side-effect free while Worker re-entry works.
- args.ts/flag-tables.ts: kept profile/alias branches; upstream's new
repeatable --config overlay flag moved into STRING_SETTERS.
- Changelogs: upstream-released bullets deduped out of Unreleased; profile
entries restored under Unreleased.
- task/index.ts: removed duplicated validateTaskIds block from auto-merge.
- Updated usage-window reporting to track remaining account quota instead of required-account counts.
- Replaced the computed "needed" metric with a non-negative remaining-quota value derived from each window's total minus used fraction.
- Changed usage output lines from "need" to "capacity" and now show used/total accounts with remaining quota multiplier.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Added per-account usage reporting in the `omp usage` command.
- Added `provider`, `json`, and `redact` options to customize usage output.
- Updated CLI wiring to route usage commands to the new per-account behavior.