Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
- Added provider/model remoteCompaction metadata and models.yml propagation.\n- Routed configured OpenAI-compatible compaction endpoints for custom providers.\n- Added compactionModel as a summary-only model selector that leaves the active session model unchanged.\n\nFixes #3104
Added the missing models.yml schema field for compat.supportsImageDetailOriginal so custom Responses-compatible proxies can opt out of snapcompact's native-resolution image hint.
Covered the CC Switch-style custom provider override and the Codex Responses wire clamp from original to auto.
Fixes#3092
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
mergeDiscoveredModel re-applied baseUrl, headers, and compat from the
provider override when an existing bundled record matched, but dropped
`transport` because the raw /v1/models payload never carries one. With
`transport: pi-native` set for an auth-gateway-routed openrouter
provider in models.yml, the next /model switch after the background
catalog refresh picked the now-transport-less entry and routed through
the default openai-completions client to ${baseUrl}/chat/completions —
a path the auth-gateway never serves, so the gateway returned
"404 No route: POST /chat/completions".
Propagate transport with the same priority as baseUrl: provider
override wins, otherwise preserve the existing (boot-time override-
applied) value, otherwise the discovered value. Cover the regression
both at the merge unit level (xiaomi-tp-discovery-merge) and at the
ModelRegistry refresh level with a mock /models fetch.
Fixes#2555
- Applied a finite 64,000-token fallback for OpenAI completion requests when model maxTokens is unavailable.
- Refactored stream max-token computation to use a shared helper and preserve a finite cap when caller and model limits are null.
- Added regression coverage for null maxTokens flows and bumped the model-cache schema to invalidate entries using retired unknown-limit sentinels.
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
- Resolved retired effort-tier variant ids in `model-resolver.ts` through the hand-table aliases (`resolveVariantAlias`, `resolveBareVariantAlias`) plus the `X-thinking` → `X` grammar (`stripThinkingVariantToken`), with exact matches always winning while a raw id is live and explicit `:effort` suffixes transferring unchanged.
- Re-keyed models.yml `modelOverrides` and rate-limit selector suppressions from raw member ids onto the collapsed model in `model-registry.ts` (`normalizeSuppressedSelector`, lazy `hasLiveModel` checks so live raw ids keep their own overrides).
- Collapsed custom/config provider model lists at registry rebuild via `collapseBuiltModelVariants`, folding config-defined `X`/`X-thinking` twins into one entry.
- Extended `model-registry.test.ts` and `model-resolver.test.ts` with effort-tier variant collapsing and alias-resolution coverage.
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Replaced canonical-row resolution with getCanonicalModelSelections in model lists and selector flow.
- Hydrated model selector state from registry on construction and kept cached selections during refresh.
- Preserved highlighted and cached model selection when offline refresh completed or reordered models.
- Added parity checks between getCanonicalModelSelections and resolveCanonicalModel via registry tests.
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
Add `Model.omitMaxOutputTokens` (`models.yml` model definitions and
`modelOverrides` accept the same field). When set, the openai-responses
and openai-completions providers stop emitting `max_output_tokens` /
`max_tokens` / `max_completion_tokens` so the upstream API applies its
own default cap. Catalog `maxTokens` is still honoured for local
budgeting (compaction, context promotion); only the wire field is
suppressed.
Restores the pre-v15.8.3 behaviour for Ollama proxies fronting cloud
catalogs (GLM, Kimi, DeepSeek): OMP cannot discover their true output
limit, so users previously set `maxTokens` to the Ollama context window
to unlock the model. v15.8.3 began sending that value as
`max_output_tokens`, triggering HTTP 400 from the upstream provider.
`applyCommonResponsesSamplingParams` now takes the model object instead
of a bare provider string so the wire-suppression flag is available to
the sampling-params builder.
Fixes#1881
Allowed OpenAI-compatible providers to opt into Anthropic cache_control markers via compat.cacheControlFormat while preserving OpenRouter Anthropic defaults.
Fixes#1845
Loaded cached startup models for special built-in providers alongside standard provider descriptors so boot-time model resolution can see cached Google Antigravity, Gemini CLI, and OpenAI Codex discoveries before refresh.\n\nFixes #1721
- Filtered recall, fact, vector, temporal, and polyphonic voices to session-owned or explicitly global memories.
- Implemented graph_query and graph_link MCP tool handlers via EpisodicGraph; wired annotations and graph into Mnemosyne's external-db path.
- Fixed restore to stage and integrity-check before replacing the live database, rolling back on failure.
- Deduped generateId with a per-process nonce to prevent batch duplicate-content collisions.
Auto-discovered OpenAI-compatible / Ollama / llama.cpp / new-api proxy
models defaulted to maxTokens: 8192 across four discovery branches in
model-registry.ts. When models hit the 8K output cap mid-stream on
legitimate large tool calls (write/edit payloads >~5KB), providers
dropped the streaming connection and Bun surfaced it as the opaque
'socket connection was closed unexpectedly'.
Extracted DISCOVERY_DEFAULT_MAX_TOKENS = 32_768 and pointed all four
discovery sites at it. min(contextWindow, ...) still honors smaller
advertised context windows on local Ollama/llama.cpp servers.
Fixes#1528
- Mocked the Vertex stream E2E test to override the home directory and clear GOOGLE_APPLICATION_CREDENTIALS so token resolution uses metadata credentials instead of local ADC files.
- Updated wafer and model-registry test expectations to match current model metadata values (Qwen3.7 Max and claude-opus-4-8).
Treat Synthetic discovery as an authoritative catalog so deprecated bundled IDs are removed from resolved model lists and cache snapshots. Validate Synthetic API keys through the models endpoint instead of a model-specific chat request.\n\nFixes #1417
Threaded cache freshness/authoritativeness through #loadCachedStandardProviderModels so dropProviderModels only fires when the cached Vertex project-catalog row is both fresh and authoritative. A stale or non-authoritative snapshot (e.g. after ADC discovery failure rewrote the row with authoritative=0) now keeps the bundled Gemini fallback in place, which would otherwise be the last working catalog in API-key-only environments.
Refs #1412
Added Google Vertex OpenAI-compatible model discovery with ADC auth and treated authoritative Vertex project catalogs as replacements for bundled Gemini fallbacks in the model registry.
Fixes#1412
Prevent disabled providers from being registered for implicit local discovery and from creating built-in model discovery managers. Added regression coverage for disabled local providers during model registry refresh.
Fixes#1232
- Wrapped initial and subsequent print-mode prompts with `logger.time` for timing instrumentation.
- Printed collected timings after session run when `PI_TIMING` env var is set.
- Added `omp auth-gateway serve/token/status` — a forward-proxy injecting broker credentials for OpenAI Chat, Anthropic Messages, and OpenAI Responses wire formats.
- Added `GET /v1/usage` to auth-broker and auth-gateway; usage cache switched to 5-min per-credential TTL with jitter and last-good fallback on failure.
- Added `AuthStorage.setConfigApiKey/removeConfigApiKey/clearConfigApiKeys` so `models.yml` `apiKey` beats OAuth tokens without overriding `--api-key`.
- Added `omp auth-broker migrate --from-local` for idempotent upload of local SQLite/env credentials to the broker.
- Replaced fromTypeBox conversion with a JSON-schema validator flow in ai tool handling and execution paths.
- Added recursive schema validation and expanded TypeBox checks for refs, enums, uniqueItems, and constraint keywords.
- Sanitized Azure/CCA tool schemas by dropping unsupported fields and rewriting oneOf tool branches as anyOf.
- Tightened argument and model-config validation, preserving unknown tool fields and adding apiKey plus compatibility flags.
Loaded cached standard provider discovery models into ModelRegistry at startup so retry fallback validation can resolve Ollama Cloud models that are already visible through --list-models.
Added regression coverage for cached ollama-cloud fallback selectors and fixed a readonly notices type error exposed by the focused type check.
Fixes#1052
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
When cached or freshly-discovered provider models carry UNK_CONTEXT_WINDOW
(222222) / UNK_MAX_TOKENS (8888) sentinels, #mergeResolvedModels was
replacing the bundled model wholesale — wiping out the correct values.
Switch to a field-level merge that preserves the bundled model's
contextWindow and maxTokens when the replacement only has sentinel
fallbacks. Custom models (via #mergeCustomModels) already had this
protection via ?? fallback; provider discoveries didn't.
Fixes the TUI showing 222222/8888 instead of the real context/token
limits for discovered models.