Llama.cpp mirrors its /v1/ apis to / which omp currently uses, however
this change is somewhat recent of only a few months ago, so omp's
llama.cpp provider does not work with older versions of llama.cpp and
some forks.
to improve compatbility use /v1/ apis for requests. model discovery
still uses /models and /props directly without v1.
because modern versions of llama.cpp mirror these, people using recent
versions should see no impact from this, while people using older
version should see improved compatibility.
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
OpenAI-compatible endpoints such as Synthetic advertise vision support through a top-level input_modalities array. Include that response field alongside direct input and nested architecture.input_modalities, with regression coverage for a catalog-absent id.
Keep the richer LM Studio /api/v0/models input metadata ahead of the thin OpenAI-compatible row while retaining row-first behavior for generic openai-models-list discovery. Add a regression covering a text-only /v1/models row paired with a native VLM record.
discoverOpenAIModelsList never consulted the /v1/models row's input
field, so custom virtual tier ids absent from the bundled catalog fell
through to the ["text"] fallback and showed images: no even when the
server advertised input: ["text","image"]. Parse the direct input array
and OpenRouter-style architecture.input_modalities from the response row,
preferring server-reported modalities over native metadata and the
bundled reference.
Fixes#7583
- Rejected local model discovery at the configured deadline even when fetch ignored abort.
- Covered pending-transport behavior with deterministic fake timers.
Fixes#7482
- Added an optional `timeoutMs` property to the provider discovery configuration schema with validation.
- Passed custom discovery timeout values through model discovery and metadata probing functions.
Fixes#6952
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
Ollama's online-if-uncached path keyed every endpoint under the same
provider namespace. Changing OLLAMA_BASE_URL or OLLAMA_HOST therefore
reused fresh models routed to the previous endpoint until cache expiry.
Centralize an endpoint-normalized Ollama cache namespace and apply it to
both configured coding-agent discovery and the catalog model manager.
Add coverage proving a default refresh discovers the new endpoint even
while the previous endpoint has a fresh row.
Fixes#7087
llama.cpp and Ollama model discovery probed /models and /props with a
250ms timeout tuned for a loopback server. That cap also applied to a
host reached over the network, so a remote or LAN LLAMA_CPP_BASE_URL
(or OLLAMA_BASE_URL/OLLAMA_HOST) with normal round-trip latency timed
out, discovery returned no models, and the picker fell back to stale
127.0.0.1:8080 entries.
Select the probe timeout by host: strictly-loopback base URLs keep the
fast fail so a busy or foreign service on the default port never stalls
startup; every non-loopback host gets a generous discovery budget.
Fixes#7087
The previous commits patched each rebuild path individually to avoid feeding a
modifyModels hook its own output. That left the invariant implicit and the
provider-scoped path applying only a subset of hooks, which is wrong for a hook
that inspects or suppresses another provider's models.
Keep #unprojectedModels as the canonical pre-projection catalog and derive
#models from it at every mutation point, so projections are always a pure
function of the unprojected base:
- #composeUnprojectedStaticModels builds the catalog; #composeStaticModels
projects it. A scoped lookup with modifiers registered composes and projects
the whole catalog before narrowing, matching getAll() followed by a filter.
Providers without modifiers keep the cheap filtered path.
- Discovery completion, registerProvider, and runtime transport overrides
update the unprojected snapshot and reproject, instead of mutating an
already-projected array.
- Runtime metadata patches apply to the unprojected model, then reproject, so
a later registration cannot discard them.
- Provider lookup snapshots are invalidated wherever the projection changes.
Hooks no longer take a providerFilter: a modifier is a whole-catalog transform
and every rebuild now runs the full ordered set exactly once.
(cherry picked from commit e5d2e9eac7c371cc196e9b362f77d3a5d7bdf507)
Versioned request-header restoration metadata inside v10 cache rows so only markers written by the old id-only matcher can bypass an unrestorable marker through requestModelId. Current aliases whose live headers differ from their static base remain unresolved and are refetched or dropped.
Added catalog and startup-registry regressions for custom-header aliases while preserving legacy Copilot -1m cache recovery.
Fixes#6284
Copilot -1m long-context variants are synthesized with transport
headers and a requestModelId to a bundled base. The v10 cache omits
headers; the writer only matched a same-id static entry, so these
variants were flagged unrestorable and dropped on the next offline
read, vanishing from the picker with a "Could not restore model"
warning. The startup registry loader dropped them the same way.
Restore/match headers through requestModelId in the cache writer, the
model-manager restore path, and the coding-agent startup loader, and
bypass a stale unrestorable marker written by the old id-only writer.
Fixes#6284
resolveCodexDiscoveryAccounts now returns null when any stored Codex OAuth
account fails to resolve (e.g. a transient refresh failure), and the manager's
resolveAccounts callback propagates null to skip discovery. Previously a failed
account was silently dropped before unionCodexModels saw it, so the remaining
accounts were unioned and cached as the authoritative catalog, hiding the
failed account's models for the cache TTL. Aborting keeps the previous/bundled
catalog.
Add a ModelRegistry regression test with one refreshable and one failing Codex
account asserting discovery makes no /models call and bundled models survive.
Fixes#6265
Keep the bearer resolved by the discovery preflight when it is not among the
stored OAuth account resolutions. This preserves Codex discovery for env,
runtime/config override, and stored non-OAuth credential sources while still
unioning all configured OAuth account catalogs.
Add a ModelRegistry regression test that drives runtime-key discovery and
asserts the resolved bearer reaches the Codex models endpoint.
Fixes#6265
A Qwen-family model served through llama.cpp ships a jinja chat template that
defaults `enable_thinking: true`, but `discoverLlamaCppModels` stamped every
local model with `reasoning: false` and an empty compat, so `--thinking off`
never reached the wire and the model kept emitting a reasoning block.
Route Qwen-family ids (plus the Qwen3.6-derived PrismLM Ternary Bonsai GGUFs,
matched with a scoped pattern rather than broadening the global
`isQwenModelId`) through a shared `applyLlamaCppQwenThinking` upgrade. It gives
them `reasoning: true` with the `qwen-template-false` disable dialect and
`qwenPreserveThinking`, since omp emits `preserve_thinking` inside
`chat_template_kwargs` for Qwen, and switches them to the chat-completions API
because the implicit llama.cpp provider defaults to `openai-responses`, whose
disable path has no Qwen encoding. The runtime base URL gains a `/v1` suffix so
the completions request does not POST to the native root, which serves
`/models` and `/props` but not `/chat/completions`; a model kept on a custom
transport (e.g. `pi-native`, whose client appends `/v1/pi/stream`) retains its
base URL so the suffix is not doubled. Non-Qwen local models keep the
configured api, base URL, and minimal compat.
The upgrade is idempotent and re-applied as the outermost transform after
discovery merges, provider/transport overrides, and cache fallbacks, so a
configured native-root `baseUrl` (which wins in `mergeDiscoveredModel`) or a
fallback to a pre-fix cached row cannot leave the routed model on the old
`openai-responses` / `reasoning: false` spec. Because routed models carry a
`/v1` base URL, the runtime metadata refresh probes the native `/models`
endpoint (stripping `/v1`, matching the existing `/props` probe) so a model's
`meta`, `status.args`, and `architecture.input_modalities` fields are not lost.
Adds discovery tests pinning the resolved reasoning/api/base URL/compat for
Qwen and non-Qwen local ids, that a configured native-root provider keeps the
`/v1` runtime URL, that a pi-native-transport model keeps its gateway URL, and
that the runtime metadata refresh for a routed model stays on native `/models`.
Signed-off-by: Christian Stewart <christian@aperture.us>
Built-in discovery skipped the OAuth refresh whenever a fresh authoritative cache existed, so an openai-codex user with an expired access token never got the model manager constructed and stale bundled models (e.g. gpt-5.4-nano) stayed selectable for the full cache TTL. Force the refresh for authoritative providers and forward the registry fetch through the Codex manager so discovery honors the configured transport.
Fixes#5364
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).
Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.
Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.
Fixes#4893
Co-authored-by: roboomp <omp@can.ac>
- Honored per-model architecture.input_modalities from llama.cpp /v1/models during discovery and selected model refresh.
- Added regression coverage for full refresh and cached selected-model metadata refresh.
- Updated the coding-agent changelog.
Fixes#4719
- Propagated llama.cpp /props input modalities through selected-model runtime refresh.
- Added a regression test for cached text-only local vision models becoming image-capable after refresh.
- Updated the coding-agent changelog for the local vision detection fix.
Fixes#4654
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.
Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).
Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.
Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
llama-server in router/preset mode advertises each preset via /v1/models,
but meta.n_ctx / n_ctx_train are only merged in after the preset's child
instance loads. The router-level /props returns a dummy n_ctx: 0. As a
result every unloaded preset fell through to DISCOVERY_DEFAULT_CONTEXT_WINDOW
(128000), and picking a preset from /model kept surfacing 128k in the
status bar regardless of the configured --ctx-size — a restart didn't
help because discovery repopulated the cache from the same broken chain.
Parse each entry's status.args (rendered CLI vector) for --ctx-size or
-c, and fall back to ctx-size = N in status.preset (INI). Positive values
slot between runtimeContextWindow and serverMetadata in the resolution
chain so a running child's live n_ctx still wins; --ctx-size 0 ("loaded
from model") is correctly skipped so we don't publish 0.
The same fallback wires through discoverLlamaCppModelRuntimeMetadata so
the refresh triggered by /model uses the configured window even before
the child spawns.
Fixes#4190
Preserved the referenced model's OpenAI-compatible reasoning-effort support when openai-models-list discovery enriches a thin /v1/models payload. The discovered model still keeps conservative proxy-local store and developer-role defaults, but known reasoning models like gpt-5 no longer force supportsReasoningEffort false and trigger the omitReasoningEffort request path.
Added a regression assertion that a thin proxied gpt-5 keeps supportsReasoningEffort true and omitReasoningEffort false after reference enrichment.
Fixes#3983
Thin OpenAI-compatible proxies that omit context_length / max_model_len on
/v1/models made every discovered model fall back to
DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a
bundled model with a much larger intrinsic window. discoverProxyModels
and discoverLiteLLMModels already resolve ids against the bundled
reference index; discoverOpenAIModelsList (which also backs lm-studio
discovery) now does the same.
Behavior:
- Build the reference index once outside the loop and resolve each item
via resolveModelReference().
- contextWindow precedence keeps provider-reported values authoritative:
item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow
?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW.
- maxTokens uses reference?.maxTokens when available, otherwise the
api-specific discovery default, capped at contextWindow so a bundled
ref for a larger sibling can never over-request output tokens.
- name / reasoning / thinking / input inherit from the reference; native
lm-studio metadata still wins for input modality.
- Provider-specific baseUrl, headers, and local-unknown cost stay local.
- OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole
/ supportsReasoningEffort all false) to match the proxy sibling.
Also updated two pre-existing regression tests that used
deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional"
ids to exercise the default-fallback branch. Those model names have since
been added to the bundled catalog, so the tests were renamed to
vllm-lab-fork-* ids that unambiguously miss the reference index while
preserving each test's original default-fallback intent.
Fixes#3983
llama.cpp /props.default_generation_settings.params.{max_tokens,n_predict} are per-request defaults the server applies when a client omits the field, not a hard model cap. Only the -1 unlimited sentinel is promoted to the runtime context window now; positive values fall back to the discovery default so client-side per-request overrides remain unconstrained.
Fixes#3781
Resolved selected-model refresh maxTokens against the effective context window, including live contextWindow overrides, so unlimited llama.cpp caps cannot exceed the configured context.
Fixes#3781
Gated discoverLlamaCppModelRuntimeMetadata's /props context fallback on the selected entry being present in /models, so refreshSelectedModelMetadata never patches a stale cached id with a different model's runtime metadata.
Fixes#3781
Mapped llama.cpp -1 generation limits from /props to the discovered runtime context window instead of the generic discovery default, including selected-model metadata refresh.
Fixes#3781
Ollama and llama.cpp discovery now prefer runtime context settings over model training metadata, so compaction thresholds match the window local servers actually accept.
Fixes#3752
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
Use a non-resolving discovery context for selected-model llama.cpp metadata refresh so command-backed and OAuth credentials stay lazy during model switches.\n\nFixes #3310
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
Configured provider discovery now treats models.yml/models.json edits as a cache staleness boundary, forcing online-if-uncached refreshes instead of reusing fresh rows written before the config change.
Added regression coverage for Ollama metadata overrides with a pre-existing models.db row.
Fixes#3242
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
Add a proxy-discovery regression: a `context_length: 0` upstream value
must be rejected by `toPositiveNumberOrUndefined` and fall back to the
default window. Raw `??` would have pinned it at 0 (nullish coalescing
does not skip 0), so this guards the should-fix the helper was added for.
Also apply biome formatting to the discovery change (collapse the
multiline `contextWindow:` expression, drop a trailing space) and the new
tests so `biome check` passes — the PR as submitted failed the formatter.
Addresses review feedback on #2466.
Updated model discovery to use `context_length` reported by the API when available, falling back to bundled reference data and default. Added `context_length` field to parsed model response and modified context window assignment logic.
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.