52 Commits

Author SHA1 Message Date
can1357 0d929bf586 Merge PR #8459: fix(coding-agent): use /v1/ APIs for llama.cpp for better compatibility (@cphlipot) 2026-08-16 02:13:37 +02:00
can1357 1e1ee2c330 test(coding-agent): updated model cache tests to resolve provider ids
- Import and use `resolveModelCacheProviderId` for github-copilot in model discovery and cache header tests.
2026-08-14 07:31:12 +02:00
Chris Phlipot 9b49684723 use /v1/ APIs for llama.cpp for better compatibility
Llama.cpp mirrors its /v1/ apis to / which omp currently uses, however
this change is somewhat recent of only a few months ago, so omp's
llama.cpp provider does not work with older versions of llama.cpp and
some forks.

to improve compatbility use /v1/ apis for requests. model discovery
still uses /models and /props directly without v1.

because modern versions of llama.cpp mirror these, people using recent
versions should see no impact from this, while people using older
version should see improved compatibility.
2026-08-13 20:21:15 -07:00
can1357 6b4823181b test: cleaned test suites and documented filtering guidelines
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
2026-08-13 08:28:42 +02:00
Vinh Nguyen ed4cfeec99 fix(coding-agent): prefer proxy-reported model name over bundled catalog name 2026-08-06 21:57:03 +07:00
roboomp 9011e8ed44 fix(config): parsed top-level input modalities
OpenAI-compatible endpoints such as Synthetic advertise vision support through a top-level input_modalities array. Include that response field alongside direct input and nested architecture.input_modalities, with regression coverage for a catalog-absent id.
2026-08-04 03:37:38 +00:00
roboomp 0c27e521c4 fix(config): preserved LM Studio native modalities
Keep the richer LM Studio /api/v0/models input metadata ahead of the thin OpenAI-compatible row while retaining row-first behavior for generic openai-models-list discovery. Add a regression covering a text-only /v1/models row paired with a native VLM record.
2026-08-04 03:25:03 +00:00
roboomp a29acbfa40 fix(config): read server input modalities in openai-models-list discovery
discoverOpenAIModelsList never consulted the /v1/models row's input
field, so custom virtual tier ids absent from the bundled catalog fell
through to the ["text"] fallback and showed images: no even when the
server advertised input: ["text","image"]. Parse the direct input array
and OpenRouter-style architecture.input_modalities from the response row,
preferring server-reported modalities over native metadata and the
bundled reference.

Fixes #7583
2026-08-04 03:19:07 +00:00
roboomp f948c61256 fix(discovery): enforced hard model probe timeouts
- Rejected local model discovery at the configured deadline even when fetch ignored abort.
- Covered pending-transport behavior with deterministic fake timers.

Fixes #7482
2026-08-03 16:43:52 +02:00
can1357 a872d77068 chore: cleanup dumb tests 2026-08-02 20:39:23 +02:00
can1357 212f0b1e00 feat(coding-agent/config): supported custom timeout configuration for model discovery
- Added an optional `timeoutMs` property to the provider discovery configuration schema with validation.
- Passed custom discovery timeout values through model discovery and metadata probing functions.

Fixes #6952
2026-08-02 20:32:31 +02:00
can1357 c565e6fb44 feat(catalog): migrated model catalog endpoints to stencil.so with ETag and ZStd
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
2026-07-31 20:38:49 +02:00
roboomp fad7e97d53 fix(catalog): scoped Ollama model caches by endpoint
Ollama's online-if-uncached path keyed every endpoint under the same
provider namespace. Changing OLLAMA_BASE_URL or OLLAMA_HOST therefore
reused fresh models routed to the previous endpoint until cache expiry.

Centralize an endpoint-normalized Ollama cache namespace and apply it to
both configured coding-agent discovery and the catalog model manager.
Add coverage proving a default refresh discovers the new endpoint even
while the previous endpoint has a fresh row.

Fixes #7087
2026-07-30 13:31:08 +00:00
roboomp ec2caea840 fix(coding-agent): honor remote llama.cpp/ollama discovery base urls
llama.cpp and Ollama model discovery probed /models and /props with a
250ms timeout tuned for a loopback server. That cap also applied to a
host reached over the network, so a remote or LAN LLAMA_CPP_BASE_URL
(or OLLAMA_BASE_URL/OLLAMA_HOST) with normal round-trip latency timed
out, discovery returned no models, and the picker fell back to stale
127.0.0.1:8080 entries.

Select the probe timeout by host: strictly-loopback base URLs keep the
fast fail so a busy or foreign service on the default port never stalls
startup; every non-loopback host gets a generous discovery budget.

Fixes #7087
2026-07-30 13:19:41 +00:00
Abhishek Sharma c7a113c9cb refactor(model-registry): derive projected catalog from an unprojected snapshot
The previous commits patched each rebuild path individually to avoid feeding a
modifyModels hook its own output. That left the invariant implicit and the
provider-scoped path applying only a subset of hooks, which is wrong for a hook
that inspects or suppresses another provider's models.

Keep #unprojectedModels as the canonical pre-projection catalog and derive
#models from it at every mutation point, so projections are always a pure
function of the unprojected base:

- #composeUnprojectedStaticModels builds the catalog; #composeStaticModels
  projects it. A scoped lookup with modifiers registered composes and projects
  the whole catalog before narrowing, matching getAll() followed by a filter.
  Providers without modifiers keep the cheap filtered path.
- Discovery completion, registerProvider, and runtime transport overrides
  update the unprojected snapshot and reproject, instead of mutating an
  already-projected array.
- Runtime metadata patches apply to the unprojected model, then reproject, so
  a later registration cannot discard them.
- Provider lookup snapshots are invalidated wherever the projection changes.

Hooks no longer take a providerFilter: a modifier is a whole-catalog transform
and every rebuild now runs the full ordered set exactly once.

(cherry picked from commit e5d2e9eac7c371cc196e9b362f77d3a5d7bdf507)
2026-07-29 23:08:37 +02:00
can1357 45df1a0fd5 Merge PR #6286: fix(catalog): restore cached Copilot 1M models via requestModelId (@roboomp) 2026-07-22 21:13:12 +02:00
roboomp b823c18837 fix(catalog): guarded request-model header recovery
Versioned request-header restoration metadata inside v10 cache rows so only markers written by the old id-only matcher can bypass an unrestorable marker through requestModelId. Current aliases whose live headers differ from their static base remain unresolved and are refetched or dropped.

Added catalog and startup-registry regressions for custom-header aliases while preserving legacy Copilot -1m cache recovery.

Fixes #6284
2026-07-22 11:10:23 +00:00
roboomp d9bab7ce8b fix(catalog): restore cached request-model variants via requestModelId
Copilot -1m long-context variants are synthesized with transport
headers and a requestModelId to a bundled base. The v10 cache omits
headers; the writer only matched a same-id static entry, so these
variants were flagged unrestorable and dropped on the next offline
read, vanishing from the picker with a "Could not restore model"
warning. The startup registry loader dropped them the same way.

Restore/match headers through requestModelId in the cache writer, the
model-manager restore path, and the coding-agent startup loader, and
bypass a stale unrestorable marker written by the old id-only writer.

Fixes #6284
2026-07-22 10:59:10 +00:00
roboomp dfeaa7aed1 fix(catalog): aborted codex discovery on failed account refresh
resolveCodexDiscoveryAccounts now returns null when any stored Codex OAuth
account fails to resolve (e.g. a transient refresh failure), and the manager's
resolveAccounts callback propagates null to skip discovery. Previously a failed
account was silently dropped before unionCodexModels saw it, so the remaining
accounts were unioned and cached as the authoritative catalog, hiding the
failed account's models for the cache TTL. Aborting keeps the previous/bundled
catalog.

Add a ModelRegistry regression test with one refreshable and one failing Codex
account asserting discovery makes no /models call and bundled models survive.

Fixes #6265
2026-07-22 07:19:11 +00:00
roboomp 2eff59e060 fix(catalog): preserved codex discovery token fallback
Keep the bearer resolved by the discovery preflight when it is not among the
stored OAuth account resolutions. This preserves Codex discovery for env,
runtime/config override, and stored non-OAuth credential sources while still
unioning all configured OAuth account catalogs.

Add a ModelRegistry regression test that drives runtime-key discovery and
asserts the resolved bearer reaches the Codex models endpoint.

Fixes #6265
2026-07-22 07:08:22 +00:00
can1357 ffab5689a0 Merge PR #5497: fix(catalog): prune unsupported Codex account models (@roboomp)
# Conflicts:
#	packages/catalog/src/discovery/codex.ts
#	packages/catalog/src/models.json
#	packages/catalog/test/codex-discovery.test.ts
2026-07-18 21:02:46 +02:00
Christian Stewart e3aa6594e5 fix(coding-agent): disable thinking on local llama.cpp Qwen models
A Qwen-family model served through llama.cpp ships a jinja chat template that
defaults `enable_thinking: true`, but `discoverLlamaCppModels` stamped every
local model with `reasoning: false` and an empty compat, so `--thinking off`
never reached the wire and the model kept emitting a reasoning block.

Route Qwen-family ids (plus the Qwen3.6-derived PrismLM Ternary Bonsai GGUFs,
matched with a scoped pattern rather than broadening the global
`isQwenModelId`) through a shared `applyLlamaCppQwenThinking` upgrade. It gives
them `reasoning: true` with the `qwen-template-false` disable dialect and
`qwenPreserveThinking`, since omp emits `preserve_thinking` inside
`chat_template_kwargs` for Qwen, and switches them to the chat-completions API
because the implicit llama.cpp provider defaults to `openai-responses`, whose
disable path has no Qwen encoding. The runtime base URL gains a `/v1` suffix so
the completions request does not POST to the native root, which serves
`/models` and `/props` but not `/chat/completions`; a model kept on a custom
transport (e.g. `pi-native`, whose client appends `/v1/pi/stream`) retains its
base URL so the suffix is not doubled. Non-Qwen local models keep the
configured api, base URL, and minimal compat.

The upgrade is idempotent and re-applied as the outermost transform after
discovery merges, provider/transport overrides, and cache fallbacks, so a
configured native-root `baseUrl` (which wins in `mergeDiscoveredModel`) or a
fallback to a pre-fix cached row cannot leave the routed model on the old
`openai-responses` / `reasoning: false` spec. Because routed models carry a
`/v1` base URL, the runtime metadata refresh probes the native `/models`
endpoint (stripping `/v1`, matching the existing `/props` probe) so a model's
`meta`, `status.args`, and `architecture.input_modalities` fields are not lost.

Adds discovery tests pinning the resolved reasoning/api/base URL/compat for
Qwen and non-Qwen local ids, that a configured native-root provider keeps the
`/v1` runtime URL, that a pi-native-transport model keeps its gateway URL, and
that the runtime metadata refresh for a routed model stays on native `/models`.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-15 21:47:48 -07:00
roboomp 5ec402fc10 fix(catalog): force codex refresh for authoritative pruning
Built-in discovery skipped the OAuth refresh whenever a fresh authoritative cache existed, so an openai-codex user with an expired access token never got the model manager constructed and stale bundled models (e.g. gpt-5.4-nano) stayed selectable for the full cache TTL. Force the refresh for authoritative providers and forward the registry fetch through the Codex manager so discovery honors the configured transport.

Fixes #5364
2026-07-14 20:47:46 +00:00
can1357 68bc6ba605 test(coding-agent): updated thinking model metadata expectations
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
2026-07-10 14:10:41 +02:00
can1357 898643f9a9 fix(coding-agent): refreshed expired OAuth in built-in discovery
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).

Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.

Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.

Fixes #4893

Co-authored-by: roboomp <omp@can.ac>
2026-07-09 18:27:23 +02:00
roboomp 90aabbac9c fix(coding-agent): detected llama.cpp model vision metadata
- Honored per-model architecture.input_modalities from llama.cpp /v1/models during discovery and selected model refresh.
- Added regression coverage for full refresh and cached selected-model metadata refresh.
- Updated the coding-agent changelog.

Fixes #4719
2026-07-06 14:25:59 +00:00
roboomp 306bde3b7e fix(coding-agent): refreshed llama vision input metadata
- Propagated llama.cpp /props input modalities through selected-model runtime refresh.
- Added a regression test for cached text-only local vision models becoming image-capable after refresh.
- Updated the coding-agent changelog for the local vision detection fix.

Fixes #4654
2026-07-06 02:00:20 +00:00
metaphorics 39668f36f7 fix(model-discovery): auto-update ZenMux models into models.db without a key
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.

Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).

Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.

Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
2026-07-02 13:50:06 +09:00
roboomp e35d772030 style: bun run fix 2026-07-02 01:16:36 +00:00
roboomp e3e79127ba fix(model-discovery): honor --ctx-size for llama.cpp router presets
llama-server in router/preset mode advertises each preset via /v1/models,
but meta.n_ctx / n_ctx_train are only merged in after the preset's child
instance loads. The router-level /props returns a dummy n_ctx: 0. As a
result every unloaded preset fell through to DISCOVERY_DEFAULT_CONTEXT_WINDOW
(128000), and picking a preset from /model kept surfacing 128k in the
status bar regardless of the configured --ctx-size — a restart didn't
help because discovery repopulated the cache from the same broken chain.

Parse each entry's status.args (rendered CLI vector) for --ctx-size or
-c, and fall back to ctx-size = N in status.preset (INI). Positive values
slot between runtimeContextWindow and serverMetadata in the resolution
chain so a running child's live n_ctx still wins; --ctx-size 0 ("loaded
from model") is correctly skipped so we don't publish 0.

The same fallback wires through discoverLlamaCppModelRuntimeMetadata so
the refresh triggered by /model uses the configured window even before
the child spawns.

Fixes #4190
2026-07-02 01:16:23 +00:00
roboomp 57a891d014 fix(coding-agent): preserved referenced reasoning effort support
Preserved the referenced model's OpenAI-compatible reasoning-effort support when openai-models-list discovery enriches a thin /v1/models payload. The discovered model still keeps conservative proxy-local store and developer-role defaults, but known reasoning models like gpt-5 no longer force supportsReasoningEffort false and trigger the omitReasoningEffort request path.

Added a regression assertion that a thin proxied gpt-5 keeps supportsReasoningEffort true and omitReasoningEffort false after reference enrichment.

Fixes #3983
2026-07-01 03:16:56 +00:00
roboomp bf75d80836 fix(coding-agent): resolve bundled reference in discoverOpenAIModelsList
Thin OpenAI-compatible proxies that omit context_length / max_model_len on
/v1/models made every discovered model fall back to
DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a
bundled model with a much larger intrinsic window. discoverProxyModels
and discoverLiteLLMModels already resolve ids against the bundled
reference index; discoverOpenAIModelsList (which also backs lm-studio
discovery) now does the same.

Behavior:
- Build the reference index once outside the loop and resolve each item
  via resolveModelReference().
- contextWindow precedence keeps provider-reported values authoritative:
  item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow
  ?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW.
- maxTokens uses reference?.maxTokens when available, otherwise the
  api-specific discovery default, capped at contextWindow so a bundled
  ref for a larger sibling can never over-request output tokens.
- name / reasoning / thinking / input inherit from the reference; native
  lm-studio metadata still wins for input modality.
- Provider-specific baseUrl, headers, and local-unknown cost stay local.
- OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole
  / supportsReasoningEffort all false) to match the proxy sibling.

Also updated two pre-existing regression tests that used
deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional"
ids to exercise the default-fallback branch. Those model names have since
been added to the bundled catalog, so the tests were renamed to
vllm-lab-fork-* ids that unambiguously miss the reference index while
preserving each test's original default-fallback intent.

Fixes #3983
2026-07-01 03:07:14 +00:00
roboomp a2f8b3915e fix(providers): treated positive llama.cpp props defaults as per-request
llama.cpp /props.default_generation_settings.params.{max_tokens,n_predict} are per-request defaults the server applies when a client omits the field, not a hard model cap. Only the -1 unlimited sentinel is promoted to the runtime context window now; positive values fall back to the discovery default so client-side per-request overrides remain unconstrained.

Fixes #3781
2026-06-29 04:10:54 +00:00
roboomp 00a41749e4 fix(providers): clamped llama.cpp refreshed output cap
Resolved selected-model refresh maxTokens against the effective context window, including live contextWindow overrides, so unlimited llama.cpp caps cannot exceed the configured context.

Fixes #3781
2026-06-29 04:03:56 +00:00
roboomp bc7e5b31a1 fix(providers): skipped llama.cpp /props fallback when model absent
Gated discoverLlamaCppModelRuntimeMetadata's /props context fallback on the selected entry being present in /models, so refreshSelectedModelMetadata never patches a stale cached id with a different model's runtime metadata.

Fixes #3781
2026-06-29 03:54:38 +00:00
roboomp dca8a7afba fix(providers): honored llama.cpp unlimited output cap
Mapped llama.cpp -1 generation limits from /props to the discovered runtime context window instead of the generic discovery default, including selected-model metadata refresh.

Fixes #3781
2026-06-29 03:47:11 +00:00
roboomp 89305f724f fix(providers): preserved llama train context fallback
Kept llama.cpp n_ctx_train as a bounded fallback only after runtime n_ctx and server props are unavailable.
2026-06-28 21:10:35 +00:00
roboomp 3600dc0cc4 fix(providers): honored local runtime context limits
Ollama and llama.cpp discovery now prefer runtime context settings over model training metadata, so compaction thresholds match the window local servers actually accept.

Fixes #3752
2026-06-28 21:03:34 +00:00
can1357 74d7dfd2fa Merge PR #3354: fix: migrate coding-agent tests to removeWithRetries (@oldschoola) 2026-06-27 02:06:38 +02:00
Jean-Luc Davern 3dc9099058 feat(coding-agent): discover rich LiteLLM metadata 2026-06-26 11:26:27 -05:00
oldschoola a2854ba768 fix: migrate coding-agent tests from fs.rm to removeWithRetries
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.

The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
2026-06-23 15:28:05 -07:00
roboomp f552f260f4 fix(providers): avoided llama cpp key resolution on switch
Use a non-resolving discovery context for selected-model llama.cpp metadata refresh so command-backed and OAuth credentials stay lazy during model switches.\n\nFixes #3310
2026-06-23 12:45:35 +00:00
roboomp 15d0e95d6f fix(providers): preserved custom llama cpp limits
Treat same-id custom llama.cpp model contextWindow and maxTokens fields as pinned limits during selected-model metadata refresh.\n\nFixes #3310
2026-06-23 12:32:55 +00:00
roboomp 24adad2890 fix(providers): honored llama cpp model context
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
2026-06-23 12:23:59 +00:00
roboomp c2339cd477 fix(providers): refreshed discovery cache after config edits
Configured provider discovery now treats models.yml/models.json edits as a cache staleness boundary, forcing online-if-uncached refreshes instead of reusing fresh rows written before the config change.

Added regression coverage for Ollama metadata overrides with a pre-existing models.db row.

Fixes #3242
2026-06-22 08:04:17 +00:00
can1357 291b3c74c2 feat: enhanced model reasoning, schema normalization, and loop guarding
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
2026-06-18 04:51:43 +02:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
can1357 bcde055998 test(config): cover non-positive context_length fallback and satisfy formatter
Add a proxy-discovery regression: a `context_length: 0` upstream value
must be rejected by `toPositiveNumberOrUndefined` and fall back to the
default window. Raw `??` would have pinned it at 0 (nullish coalescing
does not skip 0), so this guards the should-fix the helper was added for.

Also apply biome formatting to the discovery change (collapse the
multiline `contextWindow:` expression, drop a trailing space) and the new
tests so `biome check` passes — the PR as submitted failed the formatter.

Addresses review feedback on #2466.
2026-06-14 17:30:20 +02:00
Nicolas Lorin 87c9e77dbe fix(config): prefer context_length from API in model discovery
Updated model discovery to use `context_length` reported by the API when available, falling back to bundled reference data and default. Added `context_length` field to parsed model response and modified context window assignment logic.
2026-06-14 17:30:20 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00