- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).
Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.
Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.
Fixes#4893
Co-authored-by: roboomp <omp@can.ac>
- Honored per-model architecture.input_modalities from llama.cpp /v1/models during discovery and selected model refresh.
- Added regression coverage for full refresh and cached selected-model metadata refresh.
- Updated the coding-agent changelog.
Fixes#4719
- Propagated llama.cpp /props input modalities through selected-model runtime refresh.
- Added a regression test for cached text-only local vision models becoming image-capable after refresh.
- Updated the coding-agent changelog for the local vision detection fix.
Fixes#4654
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.
Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).
Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.
Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
llama-server in router/preset mode advertises each preset via /v1/models,
but meta.n_ctx / n_ctx_train are only merged in after the preset's child
instance loads. The router-level /props returns a dummy n_ctx: 0. As a
result every unloaded preset fell through to DISCOVERY_DEFAULT_CONTEXT_WINDOW
(128000), and picking a preset from /model kept surfacing 128k in the
status bar regardless of the configured --ctx-size — a restart didn't
help because discovery repopulated the cache from the same broken chain.
Parse each entry's status.args (rendered CLI vector) for --ctx-size or
-c, and fall back to ctx-size = N in status.preset (INI). Positive values
slot between runtimeContextWindow and serverMetadata in the resolution
chain so a running child's live n_ctx still wins; --ctx-size 0 ("loaded
from model") is correctly skipped so we don't publish 0.
The same fallback wires through discoverLlamaCppModelRuntimeMetadata so
the refresh triggered by /model uses the configured window even before
the child spawns.
Fixes#4190
Preserved the referenced model's OpenAI-compatible reasoning-effort support when openai-models-list discovery enriches a thin /v1/models payload. The discovered model still keeps conservative proxy-local store and developer-role defaults, but known reasoning models like gpt-5 no longer force supportsReasoningEffort false and trigger the omitReasoningEffort request path.
Added a regression assertion that a thin proxied gpt-5 keeps supportsReasoningEffort true and omitReasoningEffort false after reference enrichment.
Fixes#3983
Thin OpenAI-compatible proxies that omit context_length / max_model_len on
/v1/models made every discovered model fall back to
DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a
bundled model with a much larger intrinsic window. discoverProxyModels
and discoverLiteLLMModels already resolve ids against the bundled
reference index; discoverOpenAIModelsList (which also backs lm-studio
discovery) now does the same.
Behavior:
- Build the reference index once outside the loop and resolve each item
via resolveModelReference().
- contextWindow precedence keeps provider-reported values authoritative:
item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow
?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW.
- maxTokens uses reference?.maxTokens when available, otherwise the
api-specific discovery default, capped at contextWindow so a bundled
ref for a larger sibling can never over-request output tokens.
- name / reasoning / thinking / input inherit from the reference; native
lm-studio metadata still wins for input modality.
- Provider-specific baseUrl, headers, and local-unknown cost stay local.
- OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole
/ supportsReasoningEffort all false) to match the proxy sibling.
Also updated two pre-existing regression tests that used
deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional"
ids to exercise the default-fallback branch. Those model names have since
been added to the bundled catalog, so the tests were renamed to
vllm-lab-fork-* ids that unambiguously miss the reference index while
preserving each test's original default-fallback intent.
Fixes#3983
llama.cpp /props.default_generation_settings.params.{max_tokens,n_predict} are per-request defaults the server applies when a client omits the field, not a hard model cap. Only the -1 unlimited sentinel is promoted to the runtime context window now; positive values fall back to the discovery default so client-side per-request overrides remain unconstrained.
Fixes#3781
Resolved selected-model refresh maxTokens against the effective context window, including live contextWindow overrides, so unlimited llama.cpp caps cannot exceed the configured context.
Fixes#3781
Gated discoverLlamaCppModelRuntimeMetadata's /props context fallback on the selected entry being present in /models, so refreshSelectedModelMetadata never patches a stale cached id with a different model's runtime metadata.
Fixes#3781
Mapped llama.cpp -1 generation limits from /props to the discovered runtime context window instead of the generic discovery default, including selected-model metadata refresh.
Fixes#3781
Ollama and llama.cpp discovery now prefer runtime context settings over model training metadata, so compaction thresholds match the window local servers actually accept.
Fixes#3752
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
Use a non-resolving discovery context for selected-model llama.cpp metadata refresh so command-backed and OAuth credentials stay lazy during model switches.\n\nFixes #3310
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
Configured provider discovery now treats models.yml/models.json edits as a cache staleness boundary, forcing online-if-uncached refreshes instead of reusing fresh rows written before the config change.
Added regression coverage for Ollama metadata overrides with a pre-existing models.db row.
Fixes#3242
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
Add a proxy-discovery regression: a `context_length: 0` upstream value
must be rejected by `toPositiveNumberOrUndefined` and fall back to the
default window. Raw `??` would have pinned it at 0 (nullish coalescing
does not skip 0), so this guards the should-fix the helper was added for.
Also apply biome formatting to the discovery change (collapse the
multiline `contextWindow:` expression, drop a trailing space) and the new
tests so `biome check` passes — the PR as submitted failed the formatter.
Addresses review feedback on #2466.
Updated model discovery to use `context_length` reported by the API when available, falling back to bundled reference data and default. Added `context_length` field to parsed model response and modified context window assignment logic.
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.