fix(catalog): scoped CoreWeave catalog metadata
This commit is contained in:
@@ -7,6 +7,7 @@
|
||||
- Added a third streaming thinking-loop detection heuristic to catch "progress-lexicon stalls" where models endlessly reshuffle motivational filler without introducing new vocabulary or concrete technical references
|
||||
- Added branded wordmark and logo animation to authentication flow
|
||||
- Added a third streaming thinking-loop detection shape — a *progress-lexicon stall* — alongside verbatim tail repetition and near-duplicate (trigram) segments. It catches reasoning-summarizer loops that reshuffle the same motivational filler ("just doing it, pushing ahead, maintaining momentum") into fresh word order every paragraph: word-trigrams never cluster, but a run of substantial segments that recycle the recent vocabulary and introduce no *new* concrete reference (path / identifier / code-span) trips the guard. Summarizer title/heading lines (`**Bold Title**`, `## Heading`) are stripped before analysis so their ever-changing wording cannot mask the stall by inflating novelty. Calibrated against 537k real non-Gemini reasoning blocks (zero false positives at novelty floor 0.2 / run length 8; the real loop sustains runs of 10+).
|
||||
- Added CoreWeave Serverless Inference provider login support via `COREWEAVE_API_KEY` and `WANDB_API_KEY` fallback.
|
||||
|
||||
### Changed
|
||||
|
||||
@@ -79,7 +80,6 @@
|
||||
- Fixed OpenRouter Anthropic models on the Responses path omitting `cache_control`, so prompt caching engages without forcing Chat Completions. ([#3397](https://github.com/can1357/oh-my-pi/issues/3397))
|
||||
- Fixed OpenRouter Anthropic Responses follow-up requests replaying prior reasoning items with stale signatures, which caused HTTP 400 `Invalid signature in thinking block` errors after a thinking turn. ([#3399](https://github.com/can1357/oh-my-pi/issues/3399))
|
||||
- Fixed OpenRouter Anthropic models on the Responses path omitting `cache_control`, so prompt caching engages without forcing Chat Completions. `cacheRetention: "long"` now upgrades the breakpoint to `ttl: "1h"`. ([#3397](https://github.com/can1357/oh-my-pi/issues/3397))
|
||||
- Added CoreWeave Serverless Inference provider login support via `COREWEAVE_API_KEY` and `WANDB_API_KEY` fallback.
|
||||
|
||||
## [16.1.16] - 2026-06-23
|
||||
|
||||
|
||||
@@ -5,6 +5,7 @@
|
||||
### Added
|
||||
|
||||
- Added `OpenAICompat.qwenPreserveThinking` — auto-enabled when the resolved `thinkingFormat` is `"qwen"` or `"qwen-chat-template"` AND `replayReasoningContent` is on (i.e. the four built-in local OpenAI-compatible providers, or a custom provider pointed at a loopback / RFC1918 / `*.local` baseUrl). Pairs with the chat-completions encoder change so the request body carries `preserve_thinking: true` (twin top-level + `chat_template_kwargs` emission), keeping Qwen3.6+ from stripping `<think>...</think>` off older assistant turns and breaking the local slot's KV cache between user messages. Non-Qwen chat templates ignore the parameter, so the flag stays a no-op outside the Qwen path; users on a cloud Qwen host (Alibaba Dashscope / Qwen Portal) can opt in with `compat.qwenPreserveThinking: true`. ([#3541](https://github.com/can1357/oh-my-pi/issues/3541))
|
||||
- Added CoreWeave Serverless Inference as an OpenAI-compatible provider with models.dev-backed bundled catalog metadata.
|
||||
|
||||
## [16.1.22] - 2026-06-26
|
||||
|
||||
@@ -25,7 +26,6 @@
|
||||
- Fixed the Umans GLM-5.2 thinking-level picker collapsing to a single `high` tier after dynamic discovery: the `max` upstream level now resolves to the internal `xhigh` effort, the picker shows both `high` and `xhigh`, and the metadata maps `xhigh` back to Umans's native `max` wire tier. ([#3192](https://github.com/can1357/oh-my-pi/issues/3192))
|
||||
- Fixed GitHub Copilot business and enterprise endpoints accepting image inputs that they reject with `400 vision is not supported`. The Copilot `/models` response advertises `capabilities.supports.vision = true` for Claude/GPT chat models on every host, but only the canonical personal endpoint (`https://api.githubcopilot.com`) actually serves them; `githubCopilotModelManagerOptions` now forces `input: ["text"]` whenever discovery resolves to a non-personal base URL, and `mergeDynamicModel` honours the dynamic value (instead of OR-upgrading) when the merged endpoint differs from the bundled reference. ([#3387](https://github.com/can1357/oh-my-pi/issues/3387))
|
||||
- Fixed OpenRouter Anthropic compat to strip Responses reasoning history during replay so signed thinking blocks are not sent back to routed Anthropic providers. ([#3399](https://github.com/can1357/oh-my-pi/issues/3399))
|
||||
- Added CoreWeave Serverless Inference as an OpenAI-compatible provider with models.dev-backed bundled catalog metadata.
|
||||
|
||||
## [16.1.14] - 2026-06-22
|
||||
|
||||
|
||||
@@ -13936,7 +13936,7 @@
|
||||
"api": "openai-completions",
|
||||
"provider": "coreweave",
|
||||
"baseUrl": "https://api.inference.wandb.ai/v1",
|
||||
"reasoning": false,
|
||||
"reasoning": true,
|
||||
"input": [
|
||||
"text"
|
||||
],
|
||||
@@ -13947,7 +13947,15 @@
|
||||
"cacheWrite": 0
|
||||
},
|
||||
"contextWindow": 131072,
|
||||
"maxTokens": 131072
|
||||
"maxTokens": 131072,
|
||||
"thinking": {
|
||||
"mode": "effort",
|
||||
"efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
]
|
||||
}
|
||||
},
|
||||
"openai/gpt-oss-20b": {
|
||||
"id": "openai/gpt-oss-20b",
|
||||
@@ -13955,7 +13963,7 @@
|
||||
"api": "openai-completions",
|
||||
"provider": "coreweave",
|
||||
"baseUrl": "https://api.inference.wandb.ai/v1",
|
||||
"reasoning": false,
|
||||
"reasoning": true,
|
||||
"input": [
|
||||
"text"
|
||||
],
|
||||
@@ -13966,7 +13974,15 @@
|
||||
"cacheWrite": 0
|
||||
},
|
||||
"contextWindow": 131072,
|
||||
"maxTokens": 131072
|
||||
"maxTokens": 131072,
|
||||
"thinking": {
|
||||
"mode": "effort",
|
||||
"efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
]
|
||||
}
|
||||
},
|
||||
"OpenPipe/Qwen3-14B-Instruct": {
|
||||
"id": "OpenPipe/Qwen3-14B-Instruct",
|
||||
@@ -88438,4 +88454,4 @@
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3688,7 +3688,18 @@ const MODELS_DEV_PROVIDER_DESCRIPTORS_CORE: readonly ModelsDevProviderDescriptor
|
||||
// --- Together ---
|
||||
openAiCompletionsDescriptor("togetherai", "together", "https://api.together.xyz/v1"),
|
||||
// --- CoreWeave Serverless Inference ---
|
||||
openAiCompletionsDescriptor("wandb", "coreweave", "https://api.inference.wandb.ai/v1"),
|
||||
openAiCompletionsDescriptor("wandb", "coreweave", "https://api.inference.wandb.ai/v1", {
|
||||
transformModel: model => {
|
||||
if (!model.id.startsWith("openai/gpt-oss-")) {
|
||||
return model;
|
||||
}
|
||||
return {
|
||||
...model,
|
||||
reasoning: true,
|
||||
thinking: { mode: "effort", efforts: [Effort.Low, Effort.Medium, Effort.High] },
|
||||
};
|
||||
},
|
||||
}),
|
||||
// --- NVIDIA ---
|
||||
openAiCompletionsDescriptor("nvidia", "nvidia", "https://integrate.api.nvidia.com/v1", {
|
||||
defaultContextWindow: 131072,
|
||||
|
||||
@@ -98,7 +98,7 @@ describe("CoreWeave Serverless Inference provider support", () => {
|
||||
id: "openai/gpt-oss-120b",
|
||||
name: "GPT OSS 120B",
|
||||
tool_call: true,
|
||||
reasoning: true,
|
||||
reasoning: false,
|
||||
modalities: { input: ["text"] },
|
||||
limit: { context: 131072, output: 32768 },
|
||||
cost: { input: 0.15, output: 0.6 },
|
||||
@@ -116,6 +116,7 @@ describe("CoreWeave Serverless Inference provider support", () => {
|
||||
provider: "coreweave",
|
||||
baseUrl: "https://api.inference.wandb.ai/v1",
|
||||
reasoning: true,
|
||||
thinking: { mode: "effort", efforts: ["low", "medium", "high"] },
|
||||
contextWindow: 131072,
|
||||
maxTokens: 32768,
|
||||
cost: { input: 0.15, output: 0.6, cacheRead: 0, cacheWrite: 0 },
|
||||
|
||||
@@ -10,9 +10,8 @@
|
||||
- Added `tui.renderMermaid` to control Mermaid fenced-block ASCII rendering; disabling it also removes the Mermaid diagram hint from the generated system prompt so Mermaid blocks fall back to ordinary highlighted code fences.
|
||||
- Added `/resume <session-id>` in the interactive command system, reusing the existing session-id/prefix resolver while bare `/resume` still opens the selector.
|
||||
|
||||
### Added
|
||||
|
||||
- Added manual `omp gc` storage maintenance with `gc.*` defaults for blob sweeping, cold-session archiving, and SQLite WAL checkpointing.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed `/resume <session-id>` in the interactive TUI only searching the active cwd's session directory; id-prefix lookup now falls back to sessions from other cwd buckets like CLI `--resume <session-id>`.
|
||||
|
||||
Reference in New Issue
Block a user