Commit Graph
29 Commits
Author SHA1 Message Date
can1357 68bc6ba605 test(coding-agent): updated thinking model metadata expectations
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
2026-07-10 14:10:41 +02:00
can1357androboomp 898643f9a9 fix(coding-agent): refreshed expired OAuth in built-in discovery
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).

Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.

Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.

Fixes #4893

Co-authored-by: roboomp <omp@can.ac>
2026-07-09 18:27:23 +02:00
roboomp 90aabbac9c fix(coding-agent): detected llama.cpp model vision metadata
- Honored per-model architecture.input_modalities from llama.cpp /v1/models during discovery and selected model refresh.
- Added regression coverage for full refresh and cached selected-model metadata refresh.
- Updated the coding-agent changelog.

Fixes #4719
2026-07-06 14:25:59 +00:00
roboomp 306bde3b7e fix(coding-agent): refreshed llama vision input metadata
- Propagated llama.cpp /props input modalities through selected-model runtime refresh.
- Added a regression test for cached text-only local vision models becoming image-capable after refresh.
- Updated the coding-agent changelog for the local vision detection fix.

Fixes #4654
2026-07-06 02:00:20 +00:00
metaphorics 39668f36f7 fix(model-discovery): auto-update ZenMux models into models.db without a key
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.

Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).

Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.

Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
2026-07-02 13:50:06 +09:00
roboomp e35d772030 style: bun run fix 2026-07-02 01:16:36 +00:00
roboomp e3e79127ba fix(model-discovery): honor --ctx-size for llama.cpp router presets
llama-server in router/preset mode advertises each preset via /v1/models,
but meta.n_ctx / n_ctx_train are only merged in after the preset's child
instance loads. The router-level /props returns a dummy n_ctx: 0. As a
result every unloaded preset fell through to DISCOVERY_DEFAULT_CONTEXT_WINDOW
(128000), and picking a preset from /model kept surfacing 128k in the
status bar regardless of the configured --ctx-size — a restart didn't
help because discovery repopulated the cache from the same broken chain.

Parse each entry's status.args (rendered CLI vector) for --ctx-size or
-c, and fall back to ctx-size = N in status.preset (INI). Positive values
slot between runtimeContextWindow and serverMetadata in the resolution
chain so a running child's live n_ctx still wins; --ctx-size 0 ("loaded
from model") is correctly skipped so we don't publish 0.

The same fallback wires through discoverLlamaCppModelRuntimeMetadata so
the refresh triggered by /model uses the configured window even before
the child spawns.

Fixes #4190
2026-07-02 01:16:23 +00:00
roboomp 57a891d014 fix(coding-agent): preserved referenced reasoning effort support
Preserved the referenced model's OpenAI-compatible reasoning-effort support when openai-models-list discovery enriches a thin /v1/models payload. The discovered model still keeps conservative proxy-local store and developer-role defaults, but known reasoning models like gpt-5 no longer force supportsReasoningEffort false and trigger the omitReasoningEffort request path.

Added a regression assertion that a thin proxied gpt-5 keeps supportsReasoningEffort true and omitReasoningEffort false after reference enrichment.

Fixes #3983
2026-07-01 03:16:56 +00:00
roboomp bf75d80836 fix(coding-agent): resolve bundled reference in discoverOpenAIModelsList
Thin OpenAI-compatible proxies that omit context_length / max_model_len on
/v1/models made every discovered model fall back to
DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a
bundled model with a much larger intrinsic window. discoverProxyModels
and discoverLiteLLMModels already resolve ids against the bundled
reference index; discoverOpenAIModelsList (which also backs lm-studio
discovery) now does the same.

Behavior:
- Build the reference index once outside the loop and resolve each item
  via resolveModelReference().
- contextWindow precedence keeps provider-reported values authoritative:
  item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow
  ?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW.
- maxTokens uses reference?.maxTokens when available, otherwise the
  api-specific discovery default, capped at contextWindow so a bundled
  ref for a larger sibling can never over-request output tokens.
- name / reasoning / thinking / input inherit from the reference; native
  lm-studio metadata still wins for input modality.
- Provider-specific baseUrl, headers, and local-unknown cost stay local.
- OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole
  / supportsReasoningEffort all false) to match the proxy sibling.

Also updated two pre-existing regression tests that used
deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional"
ids to exercise the default-fallback branch. Those model names have since
been added to the bundled catalog, so the tests were renamed to
vllm-lab-fork-* ids that unambiguously miss the reference index while
preserving each test's original default-fallback intent.

Fixes #3983
2026-07-01 03:07:14 +00:00
roboomp a2f8b3915e fix(providers): treated positive llama.cpp props defaults as per-request
llama.cpp /props.default_generation_settings.params.{max_tokens,n_predict} are per-request defaults the server applies when a client omits the field, not a hard model cap. Only the -1 unlimited sentinel is promoted to the runtime context window now; positive values fall back to the discovery default so client-side per-request overrides remain unconstrained.

Fixes #3781
2026-06-29 04:10:54 +00:00
roboomp 00a41749e4 fix(providers): clamped llama.cpp refreshed output cap
Resolved selected-model refresh maxTokens against the effective context window, including live contextWindow overrides, so unlimited llama.cpp caps cannot exceed the configured context.

Fixes #3781
2026-06-29 04:03:56 +00:00
roboomp bc7e5b31a1 fix(providers): skipped llama.cpp /props fallback when model absent
Gated discoverLlamaCppModelRuntimeMetadata's /props context fallback on the selected entry being present in /models, so refreshSelectedModelMetadata never patches a stale cached id with a different model's runtime metadata.

Fixes #3781
2026-06-29 03:54:38 +00:00
roboomp dca8a7afba fix(providers): honored llama.cpp unlimited output cap
Mapped llama.cpp -1 generation limits from /props to the discovered runtime context window instead of the generic discovery default, including selected-model metadata refresh.

Fixes #3781
2026-06-29 03:47:11 +00:00
roboomp 89305f724f fix(providers): preserved llama train context fallback
Kept llama.cpp n_ctx_train as a bounded fallback only after runtime n_ctx and server props are unavailable.
2026-06-28 21:10:35 +00:00
roboomp 3600dc0cc4 fix(providers): honored local runtime context limits
Ollama and llama.cpp discovery now prefer runtime context settings over model training metadata, so compaction thresholds match the window local servers actually accept.

Fixes #3752
2026-06-28 21:03:34 +00:00
can1357 74d7dfd2fa Merge PR #3354: fix: migrate coding-agent tests to removeWithRetries (@oldschoola) 2026-06-27 02:06:38 +02:00
Jean-Luc Davern 3dc9099058 feat(coding-agent): discover rich LiteLLM metadata 2026-06-26 11:26:27 -05:00
oldschoola a2854ba768 fix: migrate coding-agent tests from fs.rm to removeWithRetries
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.

The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
2026-06-23 15:28:05 -07:00
roboomp f552f260f4 fix(providers): avoided llama cpp key resolution on switch
Use a non-resolving discovery context for selected-model llama.cpp metadata refresh so command-backed and OAuth credentials stay lazy during model switches.\n\nFixes #3310
2026-06-23 12:45:35 +00:00
roboomp 15d0e95d6f fix(providers): preserved custom llama cpp limits
Treat same-id custom llama.cpp model contextWindow and maxTokens fields as pinned limits during selected-model metadata refresh.\n\nFixes #3310
2026-06-23 12:32:55 +00:00
roboomp 24adad2890 fix(providers): honored llama cpp model context
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
2026-06-23 12:23:59 +00:00
roboomp c2339cd477 fix(providers): refreshed discovery cache after config edits
Configured provider discovery now treats models.yml/models.json edits as a cache staleness boundary, forcing online-if-uncached refreshes instead of reusing fresh rows written before the config change.

Added regression coverage for Ollama metadata overrides with a pre-existing models.db row.

Fixes #3242
2026-06-22 08:04:17 +00:00
can1357 291b3c74c2 feat: enhanced model reasoning, schema normalization, and loop guarding
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
2026-06-18 04:51:43 +02:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
can1357 bcde055998 test(config): cover non-positive context_length fallback and satisfy formatter
Add a proxy-discovery regression: a `context_length: 0` upstream value
must be rejected by `toPositiveNumberOrUndefined` and fall back to the
default window. Raw `??` would have pinned it at 0 (nullish coalescing
does not skip 0), so this guards the should-fix the helper was added for.

Also apply biome formatting to the discovery change (collapse the
multiline `contextWindow:` expression, drop a trailing space) and the new
tests so `biome check` passes — the PR as submitted failed the formatter.

Addresses review feedback on #2466.
2026-06-14 17:30:20 +02:00
Nicolas Lorinandcan1357 87c9e77dbe fix(config): prefer context_length from API in model discovery
Updated model discovery to use `context_length` reported by the API when available, falling back to bundled reference data and default. Added `context_length` field to parsed model response and modified context window assignment logic.
2026-06-14 17:30:20 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 ae415199dc feat: added build-time compatibility in ModelSpec/buildModel pipeline
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
2026-06-10 06:20:51 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00