Commit Graph
375 Commits
Author SHA1 Message Date
can1357 9ab6ea6c8f test(catalog): cover discovered Token Plan limits 2026-08-07 13:37:55 +02:00
can1357 ded3bbfce0 Merge PR #7849: fix(catalog): enrich Alibaba Token Plan discovered model limits (@Mustaqeem66) 2026-08-07 13:37:55 +02:00
can1357 0394bf4a29 Merge PR #7865: fix(catalog): update Devin reasoning family routing (@will-bogusz) 2026-08-07 13:37:55 +02:00
Voon Foo 2f24d4457e fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).

- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
  a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
  Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
  tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
  and lazy terminal errors carry the structural errorId classification
  so session auto-retry classifies stalls without text matching.
2026-08-07 14:18:46 +08:00
Will 1ad85b5584 fix(catalog): collapse current Devin reasoning families 2026-08-06 21:36:11 -04:00
Will 7e95b61ffb fix(catalog): route Devin GPT-5.6 fast max effort 2026-08-06 21:35:56 -04:00
Mustaqeem66 b80a888dc8 fix(catalog): enrich Alibaba Token Plan discovered model limits 2026-08-06 17:12:21 +00:00
roboomp e97d1fd21a test(catalog): removed tautological effort assertions 2026-08-05 02:34:04 +00:00
roboomp 736b496cc6 fix(catalog): exposed low effort tier for deepseek-v4-flash
DeepSeek's API accepts reasoning_effort low/high/max and only
deepseek-v4-flash supports all three tiers (V4 Pro is high/max). The
identity-derived effort fallback blanket-applied high/max to every
direct-DeepSeek reasoning model, hiding the low tier flash accepts.

Added isDeepseekV4FlashModelId and route flash to the low/high/max ladder
on every host; non-flash DeepSeek keeps high/max (high-only on OpenRouter).

Fixes #7668
2026-08-05 02:22:27 +00:00
Pete Samwel fce059930e fix(catalog): emit AWS GovCloud us-gov Bedrock Claude inference profiles
Bare anthropic.claude-* Bedrock rows already derived eu.* selectors; also
emit us-gov.* so GovCloud accounts can resolve system inference profiles
without requiring a full partition ARN.
2026-08-04 13:13:18 -05:00
can1357 35ed5db9e1 Merge PR #7473: fix(catalog): use live Copilot default-tier prices (@roboomp) 2026-08-03 14:46:23 +02:00
can1357 11559395c2 Merge PR #7468: fix(catalog): add reasoning config for deepseek-v4 family in alibaba-token-plan (@21307369) 2026-08-03 14:46:23 +02:00
can1357 c48376d8f0 fix(catalog): made bedrock-mantle dynamic discovery authoritative
- Account-scoped bearer /v1/models responses now replace the static seed
  instead of merging, so models disabled for the account are not selectable.
- Extended the catalog regression to run a real online refresh and assert
  the static seeds are pruned to the fetched IDs.
2026-08-03 14:36:57 +02:00
can1357 e06ccbd907 Merge PR #7080: fix(ai): add authenticated Bedrock Mantle routing (@anatoli-tsinovoy)
# Conflicts:
#	packages/ai/src/registry/registry.ts
#	packages/catalog/scripts/generated-policies.ts
#	packages/catalog/src/models.json
2026-08-03 14:36:52 +02:00
roboomp 27c1d6e8d4 fix(catalog): used live copilot default-tier prices
Applied GitHub Copilot's discovered default token-price tier to base models while preserving the provider fallback for unreported cache-write costs. Added regression coverage for GPT-5.6 Luna base and long-context pricing.

Fixes #7471
2026-08-03 08:38:14 +00:00
lsmir2 7beee683b9 fix(catalog): add reasoning config for deepseek-v4 family in alibaba-token-plan
Dynamically discovered deepseek-v4* models (e.g. deepseek-v4-flash-0731)
now receive reasoning: true and [high, max] thinking efforts via prefix
matching in the discovery mapper.
2026-08-03 15:55:14 +08:00
can1357 3177f6bdf7 test(catalog): cover DeepSeek policy without bundled models 2026-08-02 20:55:16 +02:00
can1357 990ec2b0e8 fix(catalog): exclude Token Plan speech recognition models 2026-08-02 20:53:35 +02:00
roboomp ac403dbb71 fix(catalog): excluded token plan embedding models
Filtered text-embedding model IDs from authoritative Alibaba Token Plan chat discovery.

Covered text-embedding-v4 alongside the existing media-only discovery fixtures.

Fixes #7391
2026-08-02 16:16:24 +00:00
roboomp 3224245e98 fix(catalog): admitted discovered token plan chat models
Replaced the stale static discovery allowlist with explicit filters for media-only model families.

Covered newly advertised DeepSeek, Kimi, and MiniMax chat models while retaining image, audio, and video exclusions.

Fixes #7391
2026-08-02 16:09:07 +00:00
can1357 8fe2b8f9ba Merge PR #7311: fix(catalog): honor openrouter deepseek effort metadata (@roboomp) 2026-08-01 21:29:35 +02:00
can1357 e154c8ed0b fix(catalog): sourced antigravity claude pricing from google vertex
- Claude ids now alias to google-vertex suffixed entries (claude-opus-4-6@default etc.) so Antigravity follows Google's price if it diverges from Anthropic's list price; plain-id anthropic lookup remains as dangling-alias fallback.
- Regenerated models.json via gen:models.
2026-08-01 21:27:51 +02:00
can1357 d0d15f1a55 fix(catalog): implemented fallback pricing for unpriced antigravity models
- Added `applyAntigravityPricingFallback` to back-fill unpriced `google-antigravity` models using first-party list prices from `google` and `anthropic`.
- Mapped Antigravity model IDs to preview-id aliases (e.g., `gemini-3.1-pro` to `gemini-3.1-pro-preview`) for correct pricing resolution.
- Added test coverage in `packages/catalog/test/generated-policies.test.ts` verifying fallback pricing and preservation of billable costs.
2026-08-01 21:23:17 +02:00
roboomp 2cfaeb6116 fix(catalog): honored openrouter deepseek effort metadata
- Parsed OpenRouter reasoning effort ladders and defaults during discovery.

- Preserved explicit thinking metadata from models.yml patches.

- Regenerated the catalog and covered both regression paths.

Fixes #7307
2026-08-01 19:06:44 +00:00
roboomp c00790fa3a fix(catalog): cap ollama cloud deepseek-v4 output at 65536
Ollama Cloud's deepseek-v4-pro and deepseek-v4-flash deployments reject any
output budget above 65536 with HTTP 400, despite advertising a 1M context /
384K output (ollama/ollama#16890). Ollama's /api/show never reports this cap,
so the catalog left the base models at the full context window and the dated
tag deepseek-v4-flash:0731 at a stale 8192 fallback. Pin these ids (base plus
tag variants) to min(contextWindow, 65536) at both runtime discovery and
generation; other cloud models keep their discovered limits.

Fixes #7266
2026-08-01 15:30:17 +00:00
Anatoli Tsinovoy e5541a577a fix(ai): address Bedrock Mantle review feedback 2026-08-01 12:29:07 +03:00
can1357 71c755d5f6 feat: introduced aiand provider registry entry and model catalog
- Implement the ai& provider registry entry with API-key authentication and login support.
- Add model descriptors, static model seeding, and openai-compatible model discovery for the ai& provider.
- Update the model catalog with ai& provider models, pricing, and updated provider model names.
- Add unit tests for the ai& provider environment resolution, metadata, and dynamic model mapping.
2026-08-01 08:30:37 +02:00
can1357 c565e6fb44 feat(catalog): migrated model catalog endpoints to stencil.so with ETag and ZStd
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
2026-07-31 20:38:49 +02:00
can1357 160907233b Merge PR #7097: fix(catalog): route Copilot Grok 4.5 to Responses (@roboomp) 2026-07-31 19:28:42 +02:00
can1357 fa84e3924d test(catalog): cover Cursor v2 cache invalidation 2026-07-31 19:28:10 +02:00
can1357 7bcd1aff41 Merge PR #7072: feat(cursor): expose 1M context windows in model discovery (@mmmeff)
# Conflicts:
#	packages/catalog/src/discovery/cursor.ts
2026-07-31 19:28:10 +02:00
can1357 8cb10f3f5b Merge PR #7159: fix(catalog): read Synthetic's advertised model capabilities (@Gareth-Rouse) 2026-07-31 19:08:31 +02:00
can1357 79fde0a8b7 Merge PR #7191: fix(catalog): disable store field for Google AI Studio openai-compat host (@yevman) 2026-07-31 19:07:27 +02:00
can1357 ac736d3ee1 Merge PR #7186: fix(cursor): preserve structured K3 history replay (@roboomp) 2026-07-31 19:07:19 +02:00
can1357 685311807c fix(catalog,ai): seeded GMI Cloud default model into bundled catalog
- Added GMI_CLOUD_STATIC_MODELS seed wired into gen:models so a fresh
  install resolves deepseek-ai/DeepSeek-V4-Flash synchronously at boot,
  before async /v1/models discovery fires; live discovery stays
  authoritative and replaces the seed.
- Regenerated models.json with the seeded gmi-cloud slice.
- Dropped the GeneratedProvider cast now that gmi-cloud is bundled.
- Moved gmi-cloud to the end of the /login provider order.
- Moved changelog entries from released 17.0.3 sections to Unreleased.
- Added a regression test asserting the seed covers the descriptor's
  defaultModel.
2026-07-31 18:40:09 +02:00
yevman 2429efab86 fix(catalog): disable store field for Google AI Studio openai-compat host
Google AI Studio's OpenAI-compatible endpoint
(generativelanguage.googleapis.com/v1beta/openai/chat/completions)
implements a subset of the chat-completions schema and rejects the
`store` field with HTTP 400:

  Invalid JSON payload received. Unknown name "store": Cannot find field.

OMP unconditionally emits `store: false` for any openai-completions
model whose resolved compat has supportsStore: true, so every request
to a custom provider pointed at this host (a common wiring for the
`vision` role, e.g. gemini-2.5-flash) failed before the first token.

Add a googleAistudio host class (URL-marker only, like chutes, since
users wire this host under arbitrary provider ids in models.yml) and
include it in the isNonStandard set so supportsStore resolves false.
Same failure shape as openclaw/openclaw#22704.

Detection is URL-based: verified against a captured 400 request dump
from generativelanguage.googleapis.com/v1beta/openai/chat/completions.
2026-07-31 12:29:21 -04:00
roboomp 53eae701bd fix(cursor): preserved structured K3 history replay
- Rebuilt assistant thinking, tool calls, and paired tool results in Cursor-native history shapes.
- Rejected K3 continuation when prior same-model thinking cannot be replayed safely.
- Marked dynamically discovered Cursor K3 variants as reasoning models.

Fixes #7184
2026-07-31 15:23:52 +00:00
Gareth-Rouse a2b0e2347a fix(catalog): keep Synthetic's wire-off reasoning through the manager merge
Codex P2 on b88869dc6 (reproduced live): the mapper's `reasoning: false`
for a `none`-only route only held for the raw fetcher. The production path
normalizes through `createModelManager`, where `mergeDynamicModel` merged
the dynamic row over the bundled reference with
`existingModel.reasoning || dynamicModel.reasoning` — so a stale bundled
`reasoning: true` (e.g. `hf:zai-org/GLM-5.2`) won again and `buildModel`
fabricated a `[minimal,low,medium,high,xhigh]` ladder for a route that
advertised only the `none` off-state.

`mergeDynamicModel` now treats Synthetic's discovered `reasoning` as
authoritative (both-sides-synthetic guard inside the per-id merge),
mirroring the existing Copilot `dynamicInputAuthoritative` precedent in
the same function. This loses nothing: the Synthetic mapper already folds
the reference's reasoning vote into the dynamic row whenever the wire is
silent on reasoning, so the reference still gets its say — at the mapper,
where the wire vocabulary can be consulted, instead of via a blind OR at
the merge.

The new regression test drives the full production path
(`createModelManager.refresh("online")`) rather than calling the fetcher
and `buildModel` directly, closing the gap Codex called out.
2026-07-31 11:52:23 +01:00
Gareth-Rouse 850ecb0d5f fix(catalog): keep single-tier Synthetic reasoning routes reasoning
Codex P2 on b88869dc6: the reasoning gate required more than one ladder
entry, so a route advertising exactly one named tier (e.g.
`reasoning_parameters.efforts: ["high"]`) was marked non-reasoning and its
only wire-accepted `reasoning_effort` was hidden and dropped. The
off-switch is a vocabulary with no named tiers (`["none"]` or unrecognized
values alone), not a one-tier ladder — gate on the named-tier count
instead. Adds a single-tier regression test asserted through buildModel.
2026-07-31 11:34:50 +01:00
Gareth-Rouse b88869dc64 fix(catalog): let Synthetic's wire vocabulary override reference metadata
Addresses two review items on the Synthetic capability fix:

- Codex P2: an advertised `reasoning_parameters.efforts` list is now
  authoritative over the bundled reference's `reasoning` flag. Previously a
  `hf:*` reference with `reasoning: true` re-armed the effort dial even when
  the wire advertised only the `none` off-state, so the raw spec carried a
  misleading `reasoning: true` plus a one-stop minimal ladder. When the wire
  names tiers, the reference no longer gets a vote.
- roboomp should-fix: a present-but-empty `supported_features: []` was
  treated like absent metadata and left `supportsTools` unset, which the
  request layer reads as tool-capable. A present array (empty included) with
  no `tools` entry now marks the route tool-less, while a reference that
  already vouched for tools still wins over a merely-incomplete wire list.

Tests: stale-reference override is asserted through the built-model path
(`buildModel`), and a new fixture covers the empty-features contract.
7 pass; full packages/catalog suite 525 pass, 0 fail.
2026-07-31 11:19:38 +01:00
Gareth-Rouse 306245cb78 fix(catalog): keep Synthetic effort ladders inside the wire vocabulary
Code review on b2880217e found three edge-shape leaks in the new mapper:

- A route advertising efforts: ["none"] (or only unrecognized tiers) set
  reasoning: true with thinking: undefined, so resolveModelThinking fell
  through to identity inference and fabricated an unadvertised
  low/medium/high/xhigh ladder for the request layer. None-only routes
  now map onto minimal -> none and report non-reasoning; unrecognized-only
  vocabularies keep the reference reasoning flag but never grow tiers the
  wire did not name.
- The same undefined-thinking path let a stale bundled reference ladder
  (e.g. hf:zai-org/GLM-5.2's baked tiers) survive past the wire
  vocabulary. Wire efforts now always override the reference ladder.
- supportsTools: false was written from a populated but tool-less
  supported_features list with no reference OR-fallback, unlike
  reasoning/input, so an incomplete wire list could hard-disable tools
  on a route the reference marked tool-capable. The strip now also
  honors an explicit reference supportsTools: false and only downgrades
  when neither source vouches for tools.

Tests: none-only vocabulary, stale-reference override, and the discovery
request's Authorization header. 6 pass; full packages/catalog suite 524
pass, 0 fail.
2026-07-31 10:31:35 +01:00
Gareth-Rouse b2880217e9 fix(catalog): read Synthetic's advertised model capabilities
The Synthetic discovery mapper looked for `supports_reasoning`,
`supports_vision` and `max_tokens`. Synthetic's /openai/v1/models sends
none of those: capabilities arrive in `supported_features`, the accepted
`reasoning_effort` vocabulary in `reasoning_parameters.efforts`,
modalities in `input_modalities`, the output cap in `max_output_length`,
and per-token prices in `pricing` as `$`-prefixed strings.

Every capability therefore fell back to the bundled reference, so any
route without one — the `syn:*` router aliases, newly added routes such
as `hf:moonshotai/Kimi-K3` — resolved to `reasoning: false`, text-only,
zero cost, `maxTokens: 8192`. `reasoning: false` makes
`getSupportedEfforts` return `[]`, which hides the thinking selector and
silently drops a `provider/model:high` suffix, so those models never
receive `reasoning_effort` at all; the 8k cap is low enough that verbose
models stop on `length` every turn, which reads as an incomplete
response and triggers recovery compaction at ~9% context.

Read the fields Synthetic actually sends and derive the effort ladder
from the per-model wire vocabulary. `none` is the router's thinking-off
state rather than a user tier, so it backs `minimal` through the wire
map, mirroring the existing Fireworks `minimal -> none` mapping.

Verified live against api.synthetic.new: `reasoning_effort` is accepted
on every route (`none` yields no reasoning, `high` < `max`), and the
`syn:*:vision` aliases do accept image input.
2026-07-31 10:05:52 +01:00
Matt 8dca9772b9 fix(cursor): recognized Kimi's official bare k3 id as 1M
Cursor custom ids can arrive as k3 (or kimi/k3), which
isKimiK3ModelId does not match; the bundled kimi-code catalog defines
k3 as the 1,048,576-token Kimi K3 model while k3-256k is the 256k
SKU, so gate the bare form with an exact-match pattern that keeps
k3-256k at the default window.
2026-07-30 14:50:36 -06:00
Matt 6409cb21e9 fix(cursor): broadened native 1M family matching, restored changelog
Compose the native-1M gate from the shared family parsers
(isKimiK3ModelId, parseGlmModel + semverGte 5.2 floor) so namespaced
ids (moonshotai/kimi-k3, z-ai/glm-5.2) and future GLM versions
(glm-5.10, glm-6) are covered, addressing review feedback. Restore the
Unreleased changelog entry dropped by the applied empty suggestion.
2026-07-30 11:20:20 -06:00
roboomp b7c2026980 fix(catalog): invalidated cached grok context variants
Drop the synthesized grok-4.5-1m completion route alongside its base model during the Copilot Responses migration.

Cover the failed-refresh fallback so neither stale selectable route survives when discovery is unavailable.

Fixes #7096
2026-07-30 15:49:01 +00:00
roboomp cafbe6dec0 fix(catalog): refreshed endpoint-migration caches
Fold provider cache-drop policies into the catalog fingerprint and force online-if-uncached discovery when an affected cached model is present.

Seed the Copilot migration regression with the real bundled fingerprint and verify stale Grok and MAI completion routes are rewritten through Responses.

Fixes #7096
2026-07-30 15:40:34 +00:00
roboomp 1e3ac8f69d fix(catalog): invalidated cached copilot grok route
Drop cached grok-4.5 Chat Completions rows when the bundled Copilot catalog fingerprint changes, matching the existing MAI endpoint migration.

Cover both cached endpoint migrations through the model manager's default online-if-uncached path.

Fixes #7096
2026-07-30 15:32:16 +00:00
roboomp 4819341053 fix(catalog): routed copilot grok 4.5 to responses
Route the Copilot-discovered grok-4.5 model through the Responses API so requests no longer hit the unsupported Chat Completions endpoint.

Add focused discovery coverage for the endpoint contract.

Fixes #7096
2026-07-30 15:25:46 +00:00
roboomp 6fada0c3ba fix(catalog): preserved Ollama cache path prefixes
The endpoint-scoped Ollama cache key collapsed every base URL on the
same origin. Catalog discovery preserves reverse-proxy path prefixes,
so different tenants could still reuse one fresh cache row.

Include the normalized native path in the cache fingerprint while
removing only a terminal /v1 suffix and trailing slashes. Cover tenant
isolation and equivalent native/OpenAI-compatible URL spellings.

Fixes #7087
2026-07-30 13:39:17 +00:00
roboomp fad7e97d53 fix(catalog): scoped Ollama model caches by endpoint
Ollama's online-if-uncached path keyed every endpoint under the same
provider namespace. Changing OLLAMA_BASE_URL or OLLAMA_HOST therefore
reused fresh models routed to the previous endpoint until cache expiry.

Centralize an endpoint-normalized Ollama cache namespace and apply it to
both configured coding-agent discovery and the catalog model manager.
Add coverage proving a default refresh discovers the new endpoint even
while the previous endpoint has a fresh row.

Fixes #7087
2026-07-30 13:31:08 +00:00