diff --git a/packages/ai/CHANGELOG.md b/packages/ai/CHANGELOG.md
index 65c2349f2..aa78228f3 100644
--- a/packages/ai/CHANGELOG.md
+++ b/packages/ai/CHANGELOG.md
@@ -7,6 +7,7 @@
- Added a third streaming thinking-loop detection heuristic to catch "progress-lexicon stalls" where models endlessly reshuffle motivational filler without introducing new vocabulary or concrete technical references
- Added branded wordmark and logo animation to authentication flow
- Added a third streaming thinking-loop detection shape — a *progress-lexicon stall* — alongside verbatim tail repetition and near-duplicate (trigram) segments. It catches reasoning-summarizer loops that reshuffle the same motivational filler ("just doing it, pushing ahead, maintaining momentum") into fresh word order every paragraph: word-trigrams never cluster, but a run of substantial segments that recycle the recent vocabulary and introduce no *new* concrete reference (path / identifier / code-span) trips the guard. Summarizer title/heading lines (`**Bold Title**`, `## Heading`) are stripped before analysis so their ever-changing wording cannot mask the stall by inflating novelty. Calibrated against 537k real non-Gemini reasoning blocks (zero false positives at novelty floor 0.2 / run length 8; the real loop sustains runs of 10+).
+- Added CoreWeave Serverless Inference provider login support via `COREWEAVE_API_KEY` and `WANDB_API_KEY` fallback.
### Changed
@@ -79,7 +80,6 @@
- Fixed OpenRouter Anthropic models on the Responses path omitting `cache_control`, so prompt caching engages without forcing Chat Completions. ([#3397](https://github.com/can1357/oh-my-pi/issues/3397))
- Fixed OpenRouter Anthropic Responses follow-up requests replaying prior reasoning items with stale signatures, which caused HTTP 400 `Invalid signature in thinking block` errors after a thinking turn. ([#3399](https://github.com/can1357/oh-my-pi/issues/3399))
- Fixed OpenRouter Anthropic models on the Responses path omitting `cache_control`, so prompt caching engages without forcing Chat Completions. `cacheRetention: "long"` now upgrades the breakpoint to `ttl: "1h"`. ([#3397](https://github.com/can1357/oh-my-pi/issues/3397))
-- Added CoreWeave Serverless Inference provider login support via `COREWEAVE_API_KEY` and `WANDB_API_KEY` fallback.
## [16.1.16] - 2026-06-23
diff --git a/packages/catalog/CHANGELOG.md b/packages/catalog/CHANGELOG.md
index c379b435d..634c672a4 100644
--- a/packages/catalog/CHANGELOG.md
+++ b/packages/catalog/CHANGELOG.md
@@ -5,6 +5,7 @@
### Added
- Added `OpenAICompat.qwenPreserveThinking` — auto-enabled when the resolved `thinkingFormat` is `"qwen"` or `"qwen-chat-template"` AND `replayReasoningContent` is on (i.e. the four built-in local OpenAI-compatible providers, or a custom provider pointed at a loopback / RFC1918 / `*.local` baseUrl). Pairs with the chat-completions encoder change so the request body carries `preserve_thinking: true` (twin top-level + `chat_template_kwargs` emission), keeping Qwen3.6+ from stripping `...` off older assistant turns and breaking the local slot's KV cache between user messages. Non-Qwen chat templates ignore the parameter, so the flag stays a no-op outside the Qwen path; users on a cloud Qwen host (Alibaba Dashscope / Qwen Portal) can opt in with `compat.qwenPreserveThinking: true`. ([#3541](https://github.com/can1357/oh-my-pi/issues/3541))
+- Added CoreWeave Serverless Inference as an OpenAI-compatible provider with models.dev-backed bundled catalog metadata.
## [16.1.22] - 2026-06-26
@@ -25,7 +26,6 @@
- Fixed the Umans GLM-5.2 thinking-level picker collapsing to a single `high` tier after dynamic discovery: the `max` upstream level now resolves to the internal `xhigh` effort, the picker shows both `high` and `xhigh`, and the metadata maps `xhigh` back to Umans's native `max` wire tier. ([#3192](https://github.com/can1357/oh-my-pi/issues/3192))
- Fixed GitHub Copilot business and enterprise endpoints accepting image inputs that they reject with `400 vision is not supported`. The Copilot `/models` response advertises `capabilities.supports.vision = true` for Claude/GPT chat models on every host, but only the canonical personal endpoint (`https://api.githubcopilot.com`) actually serves them; `githubCopilotModelManagerOptions` now forces `input: ["text"]` whenever discovery resolves to a non-personal base URL, and `mergeDynamicModel` honours the dynamic value (instead of OR-upgrading) when the merged endpoint differs from the bundled reference. ([#3387](https://github.com/can1357/oh-my-pi/issues/3387))
- Fixed OpenRouter Anthropic compat to strip Responses reasoning history during replay so signed thinking blocks are not sent back to routed Anthropic providers. ([#3399](https://github.com/can1357/oh-my-pi/issues/3399))
-- Added CoreWeave Serverless Inference as an OpenAI-compatible provider with models.dev-backed bundled catalog metadata.
## [16.1.14] - 2026-06-22
diff --git a/packages/catalog/src/models.json b/packages/catalog/src/models.json
index de16152ba..75c72f1fc 100644
--- a/packages/catalog/src/models.json
+++ b/packages/catalog/src/models.json
@@ -13936,7 +13936,7 @@
"api": "openai-completions",
"provider": "coreweave",
"baseUrl": "https://api.inference.wandb.ai/v1",
- "reasoning": false,
+ "reasoning": true,
"input": [
"text"
],
@@ -13947,7 +13947,15 @@
"cacheWrite": 0
},
"contextWindow": 131072,
- "maxTokens": 131072
+ "maxTokens": 131072,
+ "thinking": {
+ "mode": "effort",
+ "efforts": [
+ "low",
+ "medium",
+ "high"
+ ]
+ }
},
"openai/gpt-oss-20b": {
"id": "openai/gpt-oss-20b",
@@ -13955,7 +13963,7 @@
"api": "openai-completions",
"provider": "coreweave",
"baseUrl": "https://api.inference.wandb.ai/v1",
- "reasoning": false,
+ "reasoning": true,
"input": [
"text"
],
@@ -13966,7 +13974,15 @@
"cacheWrite": 0
},
"contextWindow": 131072,
- "maxTokens": 131072
+ "maxTokens": 131072,
+ "thinking": {
+ "mode": "effort",
+ "efforts": [
+ "low",
+ "medium",
+ "high"
+ ]
+ }
},
"OpenPipe/Qwen3-14B-Instruct": {
"id": "OpenPipe/Qwen3-14B-Instruct",
@@ -88438,4 +88454,4 @@
}
}
}
-}
\ No newline at end of file
+}
diff --git a/packages/catalog/src/provider-models/openai-compat.ts b/packages/catalog/src/provider-models/openai-compat.ts
index 0233cb99c..2e8e96122 100644
--- a/packages/catalog/src/provider-models/openai-compat.ts
+++ b/packages/catalog/src/provider-models/openai-compat.ts
@@ -3688,7 +3688,18 @@ const MODELS_DEV_PROVIDER_DESCRIPTORS_CORE: readonly ModelsDevProviderDescriptor
// --- Together ---
openAiCompletionsDescriptor("togetherai", "together", "https://api.together.xyz/v1"),
// --- CoreWeave Serverless Inference ---
- openAiCompletionsDescriptor("wandb", "coreweave", "https://api.inference.wandb.ai/v1"),
+ openAiCompletionsDescriptor("wandb", "coreweave", "https://api.inference.wandb.ai/v1", {
+ transformModel: model => {
+ if (!model.id.startsWith("openai/gpt-oss-")) {
+ return model;
+ }
+ return {
+ ...model,
+ reasoning: true,
+ thinking: { mode: "effort", efforts: [Effort.Low, Effort.Medium, Effort.High] },
+ };
+ },
+ }),
// --- NVIDIA ---
openAiCompletionsDescriptor("nvidia", "nvidia", "https://integrate.api.nvidia.com/v1", {
defaultContextWindow: 131072,
diff --git a/packages/catalog/test/coreweave-provider.test.ts b/packages/catalog/test/coreweave-provider.test.ts
index dfcd06109..d695eaeea 100644
--- a/packages/catalog/test/coreweave-provider.test.ts
+++ b/packages/catalog/test/coreweave-provider.test.ts
@@ -98,7 +98,7 @@ describe("CoreWeave Serverless Inference provider support", () => {
id: "openai/gpt-oss-120b",
name: "GPT OSS 120B",
tool_call: true,
- reasoning: true,
+ reasoning: false,
modalities: { input: ["text"] },
limit: { context: 131072, output: 32768 },
cost: { input: 0.15, output: 0.6 },
@@ -116,6 +116,7 @@ describe("CoreWeave Serverless Inference provider support", () => {
provider: "coreweave",
baseUrl: "https://api.inference.wandb.ai/v1",
reasoning: true,
+ thinking: { mode: "effort", efforts: ["low", "medium", "high"] },
contextWindow: 131072,
maxTokens: 32768,
cost: { input: 0.15, output: 0.6, cacheRead: 0, cacheWrite: 0 },
diff --git a/packages/coding-agent/CHANGELOG.md b/packages/coding-agent/CHANGELOG.md
index 8849edb97..ce03aa589 100644
--- a/packages/coding-agent/CHANGELOG.md
+++ b/packages/coding-agent/CHANGELOG.md
@@ -10,9 +10,8 @@
- Added `tui.renderMermaid` to control Mermaid fenced-block ASCII rendering; disabling it also removes the Mermaid diagram hint from the generated system prompt so Mermaid blocks fall back to ordinary highlighted code fences.
- Added `/resume ` in the interactive command system, reusing the existing session-id/prefix resolver while bare `/resume` still opens the selector.
-### Added
-
- Added manual `omp gc` storage maintenance with `gc.*` defaults for blob sweeping, cold-session archiving, and SQLite WAL checkpointing.
+
### Fixed
- Fixed `/resume ` in the interactive TUI only searching the active cwd's session directory; id-prefix lookup now falls back to sessions from other cwd buckets like CLI `--resume `.