xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
The OpenRouter non-Flash DeepSeek V4 effort override forced HIGH_ONLY for every id, so getModelDefinedEfforts clamped deepseek-v4-pro-0813 to high even though OpenRouter's /models advertises reasoning.supported_efforts [low,high,max] and the route accepts them. Carve out the dated SKU to the wire-exact low/high/max ladder while keeping the undated deepseek-v4-pro route high-only.
Fixes#8517
Copilot discovery writes an authoritative cache, so online-if-uncached served the prior endpoint for the full TTL after COPILOT_GITHUB_TOKEN switched accounts. Keying the cache namespace on the credential forces fresh discovery for a new token instead of reusing a stale personal-endpoint cache.
Fixes#8507
Threaded the shared 10s discovery AbortSignal into the copilot_internal/user probe so a stalled endpoint falls back to the personal host instead of hanging startup or refresh.
Fixes#8507
Shared the plan-endpoint probe between OAuth login and raw token model discovery so Business credentials route to their advertised API host.
Added regression coverage for the raw environment-token path.
Fixes#8507
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
DeepSeek's Chat Completions API now advertises reasoning_effort
low/high/max for both deepseek-v4-flash and deepseek-v4-pro, but the
catalog gated the low tier behind isDeepseekV4FlashModelId, so V4 Pro
(and every non-Flash reasoner) fell through to high/max. Broaden the
DeepSeek effort ladder to give any V4 SKU the low/high/max scale on the
direct API and faithful aggregator routes, keeping OpenRouter's non-Flash
route at high-only and the older V3.x/R1 reasoners at high/max.
Fixes#8405
- Deduplicated the re-imported changelog bullets from the PR #8403 merge,
keeping the condensed register with only the new #8402 entry.
- getUserHomeCandidates memoizes the WSL home candidate keyed by
platform + WSL markers + USERPROFILE, so a wedged interop pipe costs
one bounded probe per process instead of one 500ms stall per
discovery loader, while env changes (tests, SDK embeddings) still
recompute.
- Generalized thinking loop guard and helper functions to support Gemini, DeepSeek, and Grok model families.
- Replaced `withGeminiThinkingLoopGuard` and related Gemini-specific symbols with generalized counterparts.
- Removed deprecated `enableGeminiThinkingLoopGuard` options and associated tests.
- Updated test suites and agent session logic to use the generalized thinking loop guard and model family tokens.
- Added `isGrok46ModelId` boundary-aware model predicate to the catalog package.
- Included Grok 4.6 models in the thinking-loop guard to prevent runaway reasoning streams.
- The mux pane-growth oracle treated every physical scroll as a logical append, but immutable-history recovery can recommit a corrected suffix after an off-screen mutation without advancing the shadow tape; exempt only changed shared history prefixes.
- The frame-neutral oracle compared prepared rows only; at narrow widths distinct raw rows prepare identically, so the renderer's raw-prefix divergence recovery recommits legitimately. Snapshot raw frames and allow declared transient growth.
- OSC66 spacer preservation intentionally composes six bounded context rows above the resize viewport; assert that exact bound instead of zero above-fold rendering.
Both oracle false positives reproduce identically at the PR head that introduced the harness (4cc9725037); three full randomized stress passes green after the fix.
Ollama Cloud serves the DeepSeek V4 family over the ollama-chat API,
which bypassed the deepseek effort-ladder branch in model-thinking.ts.
deepseek-v4-flash exposed the generic minimal/low/medium/high/xhigh
ladder with no max tier instead of the low/high/max contract the
catalog encodes on every other host.
Broadened the branch to also cover the ollama-cloud ollama-chat surface:
Flash keeps low/high/max, V4 Pro and older reasoners top out at
high/max. Regenerated models.json accordingly.
Fixes#8334
Built-in OpenAI-compatible provider managers call fetchOpenAICompatibleModels with neither a signal nor a timeoutMs, and the no-timeout branch issued the request with signal: undefined. A stalled /models endpoint left the fetch pending forever, so createAgentSession's awaited resolveModelDiscoveryFallback discovery pass never returned and startup hung.
Apply a default 10s deadline (DEFAULT_OPENAI_COMPATIBLE_DISCOVERY_TIMEOUT_MS) when the caller supplies neither signal nor timeoutMs, matching the coding-agent remote-discovery budget. Callers passing their own signal keep owning its lifecycle.
Fixes#8315
- Bound PAX sparse record memory overhead by caching sparse markers and specific keys.
- Update system prompt phrasing and tests for tool inventory and date displays.
- Added Google provider thinking configuration parameters and force-reasoning-off controls.
- Implemented MCP SSE stream resumption using Last-Event-ID and `SSEResumeError`.
- Added support for TAR old-GNU sparse extension blocks, path length checks, and archive entry overrides.
- Restricted external thinking support to specific models and added semver fallback parsing.
Adds zai-org/GLM-5.2-Fast to the Baseten reasoning allowlists, matching
the sibling zai-org/GLM-5.2. Also fixes parseGlmModel to handle uppercase
GLM model IDs (used by Baseten, CoreWeave, HuggingFace, etc.), so the
identity deriver correctly classifies them as GLM-5.2 reasoning models
during catalog generation — previously the case-sensitive regex caused
rebakeModelThinking to fall back to the generic effort ladder.
Regenerated models.json with a live Baseten API key: GLM-5.2 and
GLM-5.2-Fast now bundle reasoning:true with the correct high/max effort
ladder, and other uppercase GLM-5.2 resellers (CoreWeave, HuggingFace,
Synthetic, Together, Wafer) get the corrected minimal..max ladder.
- Update packages/catalog/test/meta-provider.test.ts to expect
three META_MUSE_STATIC_MODELS entries (1.1, 1.2, contributor)
with input ["text","image"] — fixes blocking failure.
- Add ## [Unreleased] entry in packages/catalog/CHANGELOG.md per
AGENTS.md.
- Remove trailing newline from models.json to match
generate-models.ts (Bun.write without \n).
Co-authored-by: roboomp <roboomp@users.noreply.github.com>
- The regenerated bundle (merged with PR #8021) now ships a
responses-route github-copilot/grok-4.5, so the id legitimately
resurfaces from the bundle when the migration refresh fails; the
contract worth defending is that the stale cached completions route
never returns and the unbundled long-context variant stays dropped.
- The OpenCode Go gateway does not serve DSV4-Flash at
/zen/go/v1/chat/completions; /zen/go/v1/responses works (user-verified
against the live gateway). Added a per-id override in
OPENCODE_GO_API_RESOLUTION so both bundled generation and the runtime
/v1/models refresh route it to openai-responses; deepseek-v4-pro keeps
chat completions.
- Regenerated models.json from the resolver source.
Applied curated Alibaba Token Plan seeds after generic models.dev fallback so bundled capabilities cannot be overwritten by incomplete upstream metadata.
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).
- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
and lazy terminal errors carry the structural errorId classification
so session auto-retry classifies stalls without text matching.
DeepSeek's API accepts reasoning_effort low/high/max and only
deepseek-v4-flash supports all three tiers (V4 Pro is high/max). The
identity-derived effort fallback blanket-applied high/max to every
direct-DeepSeek reasoning model, hiding the low tier flash accepts.
Added isDeepseekV4FlashModelId and route flash to the low/high/max ladder
on every host; non-flash DeepSeek keeps high/max (high-only on OpenRouter).
Fixes#7668