The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).
- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
and lazy terminal errors carry the structural errorId classification
so session auto-retry classifies stalls without text matching.
DeepSeek's API accepts reasoning_effort low/high/max and only
deepseek-v4-flash supports all three tiers (V4 Pro is high/max). The
identity-derived effort fallback blanket-applied high/max to every
direct-DeepSeek reasoning model, hiding the low tier flash accepts.
Added isDeepseekV4FlashModelId and route flash to the low/high/max ladder
on every host; non-flash DeepSeek keeps high/max (high-only on OpenRouter).
Fixes#7668
Bare anthropic.claude-* Bedrock rows already derived eu.* selectors; also
emit us-gov.* so GovCloud accounts can resolve system inference profiles
without requiring a full partition ARN.
- Account-scoped bearer /v1/models responses now replace the static seed
instead of merging, so models disabled for the account are not selectable.
- Extended the catalog regression to run a real online refresh and assert
the static seeds are pruned to the fetched IDs.
Applied GitHub Copilot's discovered default token-price tier to base models while preserving the provider fallback for unreported cache-write costs. Added regression coverage for GPT-5.6 Luna base and long-context pricing.
Fixes#7471
Dynamically discovered deepseek-v4* models (e.g. deepseek-v4-flash-0731)
now receive reasoning: true and [high, max] thinking efforts via prefix
matching in the discovery mapper.
Filtered text-embedding model IDs from authoritative Alibaba Token Plan chat discovery.
Covered text-embedding-v4 alongside the existing media-only discovery fixtures.
Fixes#7391
Replaced the stale static discovery allowlist with explicit filters for media-only model families.
Covered newly advertised DeepSeek, Kimi, and MiniMax chat models while retaining image, audio, and video exclusions.
Fixes#7391
- Claude ids now alias to google-vertex suffixed entries (claude-opus-4-6@default etc.) so Antigravity follows Google's price if it diverges from Anthropic's list price; plain-id anthropic lookup remains as dangling-alias fallback.
- Regenerated models.json via gen:models.
- Added `applyAntigravityPricingFallback` to back-fill unpriced `google-antigravity` models using first-party list prices from `google` and `anthropic`.
- Mapped Antigravity model IDs to preview-id aliases (e.g., `gemini-3.1-pro` to `gemini-3.1-pro-preview`) for correct pricing resolution.
- Added test coverage in `packages/catalog/test/generated-policies.test.ts` verifying fallback pricing and preservation of billable costs.
- Parsed OpenRouter reasoning effort ladders and defaults during discovery.
- Preserved explicit thinking metadata from models.yml patches.
- Regenerated the catalog and covered both regression paths.
Fixes#7307
Ollama Cloud's deepseek-v4-pro and deepseek-v4-flash deployments reject any
output budget above 65536 with HTTP 400, despite advertising a 1M context /
384K output (ollama/ollama#16890). Ollama's /api/show never reports this cap,
so the catalog left the base models at the full context window and the dated
tag deepseek-v4-flash:0731 at a stale 8192 fallback. Pin these ids (base plus
tag variants) to min(contextWindow, 65536) at both runtime discovery and
generation; other cloud models keep their discovered limits.
Fixes#7266
- Implement the ai& provider registry entry with API-key authentication and login support.
- Add model descriptors, static model seeding, and openai-compatible model discovery for the ai& provider.
- Update the model catalog with ai& provider models, pricing, and updated provider model names.
- Add unit tests for the ai& provider environment resolution, metadata, and dynamic model mapping.
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
- Added GMI_CLOUD_STATIC_MODELS seed wired into gen:models so a fresh
install resolves deepseek-ai/DeepSeek-V4-Flash synchronously at boot,
before async /v1/models discovery fires; live discovery stays
authoritative and replaces the seed.
- Regenerated models.json with the seeded gmi-cloud slice.
- Dropped the GeneratedProvider cast now that gmi-cloud is bundled.
- Moved gmi-cloud to the end of the /login provider order.
- Moved changelog entries from released 17.0.3 sections to Unreleased.
- Added a regression test asserting the seed covers the descriptor's
defaultModel.
Google AI Studio's OpenAI-compatible endpoint
(generativelanguage.googleapis.com/v1beta/openai/chat/completions)
implements a subset of the chat-completions schema and rejects the
`store` field with HTTP 400:
Invalid JSON payload received. Unknown name "store": Cannot find field.
OMP unconditionally emits `store: false` for any openai-completions
model whose resolved compat has supportsStore: true, so every request
to a custom provider pointed at this host (a common wiring for the
`vision` role, e.g. gemini-2.5-flash) failed before the first token.
Add a googleAistudio host class (URL-marker only, like chutes, since
users wire this host under arbitrary provider ids in models.yml) and
include it in the isNonStandard set so supportsStore resolves false.
Same failure shape as openclaw/openclaw#22704.
Detection is URL-based: verified against a captured 400 request dump
from generativelanguage.googleapis.com/v1beta/openai/chat/completions.
Codex P2 on b88869dc6 (reproduced live): the mapper's `reasoning: false`
for a `none`-only route only held for the raw fetcher. The production path
normalizes through `createModelManager`, where `mergeDynamicModel` merged
the dynamic row over the bundled reference with
`existingModel.reasoning || dynamicModel.reasoning` — so a stale bundled
`reasoning: true` (e.g. `hf:zai-org/GLM-5.2`) won again and `buildModel`
fabricated a `[minimal,low,medium,high,xhigh]` ladder for a route that
advertised only the `none` off-state.
`mergeDynamicModel` now treats Synthetic's discovered `reasoning` as
authoritative (both-sides-synthetic guard inside the per-id merge),
mirroring the existing Copilot `dynamicInputAuthoritative` precedent in
the same function. This loses nothing: the Synthetic mapper already folds
the reference's reasoning vote into the dynamic row whenever the wire is
silent on reasoning, so the reference still gets its say — at the mapper,
where the wire vocabulary can be consulted, instead of via a blind OR at
the merge.
The new regression test drives the full production path
(`createModelManager.refresh("online")`) rather than calling the fetcher
and `buildModel` directly, closing the gap Codex called out.
Codex P2 on b88869dc6: the reasoning gate required more than one ladder
entry, so a route advertising exactly one named tier (e.g.
`reasoning_parameters.efforts: ["high"]`) was marked non-reasoning and its
only wire-accepted `reasoning_effort` was hidden and dropped. The
off-switch is a vocabulary with no named tiers (`["none"]` or unrecognized
values alone), not a one-tier ladder — gate on the named-tier count
instead. Adds a single-tier regression test asserted through buildModel.
Addresses two review items on the Synthetic capability fix:
- Codex P2: an advertised `reasoning_parameters.efforts` list is now
authoritative over the bundled reference's `reasoning` flag. Previously a
`hf:*` reference with `reasoning: true` re-armed the effort dial even when
the wire advertised only the `none` off-state, so the raw spec carried a
misleading `reasoning: true` plus a one-stop minimal ladder. When the wire
names tiers, the reference no longer gets a vote.
- roboomp should-fix: a present-but-empty `supported_features: []` was
treated like absent metadata and left `supportsTools` unset, which the
request layer reads as tool-capable. A present array (empty included) with
no `tools` entry now marks the route tool-less, while a reference that
already vouched for tools still wins over a merely-incomplete wire list.
Tests: stale-reference override is asserted through the built-model path
(`buildModel`), and a new fixture covers the empty-features contract.
7 pass; full packages/catalog suite 525 pass, 0 fail.
Code review on b2880217e found three edge-shape leaks in the new mapper:
- A route advertising efforts: ["none"] (or only unrecognized tiers) set
reasoning: true with thinking: undefined, so resolveModelThinking fell
through to identity inference and fabricated an unadvertised
low/medium/high/xhigh ladder for the request layer. None-only routes
now map onto minimal -> none and report non-reasoning; unrecognized-only
vocabularies keep the reference reasoning flag but never grow tiers the
wire did not name.
- The same undefined-thinking path let a stale bundled reference ladder
(e.g. hf:zai-org/GLM-5.2's baked tiers) survive past the wire
vocabulary. Wire efforts now always override the reference ladder.
- supportsTools: false was written from a populated but tool-less
supported_features list with no reference OR-fallback, unlike
reasoning/input, so an incomplete wire list could hard-disable tools
on a route the reference marked tool-capable. The strip now also
honors an explicit reference supportsTools: false and only downgrades
when neither source vouches for tools.
Tests: none-only vocabulary, stale-reference override, and the discovery
request's Authorization header. 6 pass; full packages/catalog suite 524
pass, 0 fail.
The Synthetic discovery mapper looked for `supports_reasoning`,
`supports_vision` and `max_tokens`. Synthetic's /openai/v1/models sends
none of those: capabilities arrive in `supported_features`, the accepted
`reasoning_effort` vocabulary in `reasoning_parameters.efforts`,
modalities in `input_modalities`, the output cap in `max_output_length`,
and per-token prices in `pricing` as `$`-prefixed strings.
Every capability therefore fell back to the bundled reference, so any
route without one — the `syn:*` router aliases, newly added routes such
as `hf:moonshotai/Kimi-K3` — resolved to `reasoning: false`, text-only,
zero cost, `maxTokens: 8192`. `reasoning: false` makes
`getSupportedEfforts` return `[]`, which hides the thinking selector and
silently drops a `provider/model:high` suffix, so those models never
receive `reasoning_effort` at all; the 8k cap is low enough that verbose
models stop on `length` every turn, which reads as an incomplete
response and triggers recovery compaction at ~9% context.
Read the fields Synthetic actually sends and derive the effort ladder
from the per-model wire vocabulary. `none` is the router's thinking-off
state rather than a user tier, so it backs `minimal` through the wire
map, mirroring the existing Fireworks `minimal -> none` mapping.
Verified live against api.synthetic.new: `reasoning_effort` is accepted
on every route (`none` yields no reasoning, `high` < `max`), and the
`syn:*:vision` aliases do accept image input.
Cursor custom ids can arrive as k3 (or kimi/k3), which
isKimiK3ModelId does not match; the bundled kimi-code catalog defines
k3 as the 1,048,576-token Kimi K3 model while k3-256k is the 256k
SKU, so gate the bare form with an exact-match pattern that keeps
k3-256k at the default window.
Compose the native-1M gate from the shared family parsers
(isKimiK3ModelId, parseGlmModel + semverGte 5.2 floor) so namespaced
ids (moonshotai/kimi-k3, z-ai/glm-5.2) and future GLM versions
(glm-5.10, glm-6) are covered, addressing review feedback. Restore the
Unreleased changelog entry dropped by the applied empty suggestion.
Drop the synthesized grok-4.5-1m completion route alongside its base model during the Copilot Responses migration.
Cover the failed-refresh fallback so neither stale selectable route survives when discovery is unavailable.
Fixes#7096
Fold provider cache-drop policies into the catalog fingerprint and force online-if-uncached discovery when an affected cached model is present.
Seed the Copilot migration regression with the real bundled fingerprint and verify stale Grok and MAI completion routes are rewritten through Responses.
Fixes#7096
Drop cached grok-4.5 Chat Completions rows when the bundled Copilot catalog fingerprint changes, matching the existing MAI endpoint migration.
Cover both cached endpoint migrations through the model manager's default online-if-uncached path.
Fixes#7096
Route the Copilot-discovered grok-4.5 model through the Responses API so requests no longer hit the unsupported Chat Completions endpoint.
Add focused discovery coverage for the endpoint contract.
Fixes#7096
The endpoint-scoped Ollama cache key collapsed every base URL on the
same origin. Catalog discovery preserves reverse-proxy path prefixes,
so different tenants could still reuse one fresh cache row.
Include the normalized native path in the cache fingerprint while
removing only a terminal /v1 suffix and trailing slashes. Cover tenant
isolation and equivalent native/OpenAI-compatible URL spellings.
Fixes#7087
Ollama's online-if-uncached path keyed every endpoint under the same
provider namespace. Changing OLLAMA_BASE_URL or OLLAMA_HOST therefore
reused fresh models routed to the previous endpoint until cache expiry.
Centralize an endpoint-normalized Ollama cache namespace and apply it to
both configured coding-agent discovery and the catalog model manager.
Add coverage proving a default refresh discovers the new endpoint even
while the previous endpoint has a fresh row.
Fixes#7087