diff --git a/packages/ai/CHANGELOG.md b/packages/ai/CHANGELOG.md index 65c2349f2..aa78228f3 100644 --- a/packages/ai/CHANGELOG.md +++ b/packages/ai/CHANGELOG.md @@ -7,6 +7,7 @@ - Added a third streaming thinking-loop detection heuristic to catch "progress-lexicon stalls" where models endlessly reshuffle motivational filler without introducing new vocabulary or concrete technical references - Added branded wordmark and logo animation to authentication flow - Added a third streaming thinking-loop detection shape — a *progress-lexicon stall* — alongside verbatim tail repetition and near-duplicate (trigram) segments. It catches reasoning-summarizer loops that reshuffle the same motivational filler ("just doing it, pushing ahead, maintaining momentum") into fresh word order every paragraph: word-trigrams never cluster, but a run of substantial segments that recycle the recent vocabulary and introduce no *new* concrete reference (path / identifier / code-span) trips the guard. Summarizer title/heading lines (`**Bold Title**`, `## Heading`) are stripped before analysis so their ever-changing wording cannot mask the stall by inflating novelty. Calibrated against 537k real non-Gemini reasoning blocks (zero false positives at novelty floor 0.2 / run length 8; the real loop sustains runs of 10+). +- Added CoreWeave Serverless Inference provider login support via `COREWEAVE_API_KEY` and `WANDB_API_KEY` fallback. ### Changed @@ -79,7 +80,6 @@ - Fixed OpenRouter Anthropic models on the Responses path omitting `cache_control`, so prompt caching engages without forcing Chat Completions. ([#3397](https://github.com/can1357/oh-my-pi/issues/3397)) - Fixed OpenRouter Anthropic Responses follow-up requests replaying prior reasoning items with stale signatures, which caused HTTP 400 `Invalid signature in thinking block` errors after a thinking turn. ([#3399](https://github.com/can1357/oh-my-pi/issues/3399)) - Fixed OpenRouter Anthropic models on the Responses path omitting `cache_control`, so prompt caching engages without forcing Chat Completions. `cacheRetention: "long"` now upgrades the breakpoint to `ttl: "1h"`. ([#3397](https://github.com/can1357/oh-my-pi/issues/3397)) -- Added CoreWeave Serverless Inference provider login support via `COREWEAVE_API_KEY` and `WANDB_API_KEY` fallback. ## [16.1.16] - 2026-06-23 diff --git a/packages/catalog/CHANGELOG.md b/packages/catalog/CHANGELOG.md index c379b435d..634c672a4 100644 --- a/packages/catalog/CHANGELOG.md +++ b/packages/catalog/CHANGELOG.md @@ -5,6 +5,7 @@ ### Added - Added `OpenAICompat.qwenPreserveThinking` — auto-enabled when the resolved `thinkingFormat` is `"qwen"` or `"qwen-chat-template"` AND `replayReasoningContent` is on (i.e. the four built-in local OpenAI-compatible providers, or a custom provider pointed at a loopback / RFC1918 / `*.local` baseUrl). Pairs with the chat-completions encoder change so the request body carries `preserve_thinking: true` (twin top-level + `chat_template_kwargs` emission), keeping Qwen3.6+ from stripping `...` off older assistant turns and breaking the local slot's KV cache between user messages. Non-Qwen chat templates ignore the parameter, so the flag stays a no-op outside the Qwen path; users on a cloud Qwen host (Alibaba Dashscope / Qwen Portal) can opt in with `compat.qwenPreserveThinking: true`. ([#3541](https://github.com/can1357/oh-my-pi/issues/3541)) +- Added CoreWeave Serverless Inference as an OpenAI-compatible provider with models.dev-backed bundled catalog metadata. ## [16.1.22] - 2026-06-26 @@ -25,7 +26,6 @@ - Fixed the Umans GLM-5.2 thinking-level picker collapsing to a single `high` tier after dynamic discovery: the `max` upstream level now resolves to the internal `xhigh` effort, the picker shows both `high` and `xhigh`, and the metadata maps `xhigh` back to Umans's native `max` wire tier. ([#3192](https://github.com/can1357/oh-my-pi/issues/3192)) - Fixed GitHub Copilot business and enterprise endpoints accepting image inputs that they reject with `400 vision is not supported`. The Copilot `/models` response advertises `capabilities.supports.vision = true` for Claude/GPT chat models on every host, but only the canonical personal endpoint (`https://api.githubcopilot.com`) actually serves them; `githubCopilotModelManagerOptions` now forces `input: ["text"]` whenever discovery resolves to a non-personal base URL, and `mergeDynamicModel` honours the dynamic value (instead of OR-upgrading) when the merged endpoint differs from the bundled reference. ([#3387](https://github.com/can1357/oh-my-pi/issues/3387)) - Fixed OpenRouter Anthropic compat to strip Responses reasoning history during replay so signed thinking blocks are not sent back to routed Anthropic providers. ([#3399](https://github.com/can1357/oh-my-pi/issues/3399)) -- Added CoreWeave Serverless Inference as an OpenAI-compatible provider with models.dev-backed bundled catalog metadata. ## [16.1.14] - 2026-06-22 diff --git a/packages/catalog/src/models.json b/packages/catalog/src/models.json index de16152ba..75c72f1fc 100644 --- a/packages/catalog/src/models.json +++ b/packages/catalog/src/models.json @@ -13936,7 +13936,7 @@ "api": "openai-completions", "provider": "coreweave", "baseUrl": "https://api.inference.wandb.ai/v1", - "reasoning": false, + "reasoning": true, "input": [ "text" ], @@ -13947,7 +13947,15 @@ "cacheWrite": 0 }, "contextWindow": 131072, - "maxTokens": 131072 + "maxTokens": 131072, + "thinking": { + "mode": "effort", + "efforts": [ + "low", + "medium", + "high" + ] + } }, "openai/gpt-oss-20b": { "id": "openai/gpt-oss-20b", @@ -13955,7 +13963,7 @@ "api": "openai-completions", "provider": "coreweave", "baseUrl": "https://api.inference.wandb.ai/v1", - "reasoning": false, + "reasoning": true, "input": [ "text" ], @@ -13966,7 +13974,15 @@ "cacheWrite": 0 }, "contextWindow": 131072, - "maxTokens": 131072 + "maxTokens": 131072, + "thinking": { + "mode": "effort", + "efforts": [ + "low", + "medium", + "high" + ] + } }, "OpenPipe/Qwen3-14B-Instruct": { "id": "OpenPipe/Qwen3-14B-Instruct", @@ -88438,4 +88454,4 @@ } } } -} \ No newline at end of file +} diff --git a/packages/catalog/src/provider-models/openai-compat.ts b/packages/catalog/src/provider-models/openai-compat.ts index 0233cb99c..2e8e96122 100644 --- a/packages/catalog/src/provider-models/openai-compat.ts +++ b/packages/catalog/src/provider-models/openai-compat.ts @@ -3688,7 +3688,18 @@ const MODELS_DEV_PROVIDER_DESCRIPTORS_CORE: readonly ModelsDevProviderDescriptor // --- Together --- openAiCompletionsDescriptor("togetherai", "together", "https://api.together.xyz/v1"), // --- CoreWeave Serverless Inference --- - openAiCompletionsDescriptor("wandb", "coreweave", "https://api.inference.wandb.ai/v1"), + openAiCompletionsDescriptor("wandb", "coreweave", "https://api.inference.wandb.ai/v1", { + transformModel: model => { + if (!model.id.startsWith("openai/gpt-oss-")) { + return model; + } + return { + ...model, + reasoning: true, + thinking: { mode: "effort", efforts: [Effort.Low, Effort.Medium, Effort.High] }, + }; + }, + }), // --- NVIDIA --- openAiCompletionsDescriptor("nvidia", "nvidia", "https://integrate.api.nvidia.com/v1", { defaultContextWindow: 131072, diff --git a/packages/catalog/test/coreweave-provider.test.ts b/packages/catalog/test/coreweave-provider.test.ts index dfcd06109..d695eaeea 100644 --- a/packages/catalog/test/coreweave-provider.test.ts +++ b/packages/catalog/test/coreweave-provider.test.ts @@ -98,7 +98,7 @@ describe("CoreWeave Serverless Inference provider support", () => { id: "openai/gpt-oss-120b", name: "GPT OSS 120B", tool_call: true, - reasoning: true, + reasoning: false, modalities: { input: ["text"] }, limit: { context: 131072, output: 32768 }, cost: { input: 0.15, output: 0.6 }, @@ -116,6 +116,7 @@ describe("CoreWeave Serverless Inference provider support", () => { provider: "coreweave", baseUrl: "https://api.inference.wandb.ai/v1", reasoning: true, + thinking: { mode: "effort", efforts: ["low", "medium", "high"] }, contextWindow: 131072, maxTokens: 32768, cost: { input: 0.15, output: 0.6, cacheRead: 0, cacheWrite: 0 }, diff --git a/packages/coding-agent/CHANGELOG.md b/packages/coding-agent/CHANGELOG.md index 8849edb97..ce03aa589 100644 --- a/packages/coding-agent/CHANGELOG.md +++ b/packages/coding-agent/CHANGELOG.md @@ -10,9 +10,8 @@ - Added `tui.renderMermaid` to control Mermaid fenced-block ASCII rendering; disabling it also removes the Mermaid diagram hint from the generated system prompt so Mermaid blocks fall back to ordinary highlighted code fences. - Added `/resume ` in the interactive command system, reusing the existing session-id/prefix resolver while bare `/resume` still opens the selector. -### Added - - Added manual `omp gc` storage maintenance with `gc.*` defaults for blob sweeping, cold-session archiving, and SQLite WAL checkpointing. + ### Fixed - Fixed `/resume ` in the interactive TUI only searching the active cwd's session directory; id-prefix lookup now falls back to sessions from other cwd buckets like CLI `--resume `.