Commit Graph

416 Commits

Author SHA1 Message Date
Yang Yang 86866b8bbf fix(catalog): omit Responses penalties on all first-party xAI models
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
2026-08-14 22:02:52 -07:00
Yang Yang 76faddf886 fix(catalog): omit reasoningEffortMap on no-dial xAI rows
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
2026-08-14 22:02:52 -07:00
Yang Yang 09830d2bd6 fix(catalog): drop unsupported xhigh effort from first-party Grok
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
2026-08-14 22:02:52 -07:00
Yang Yang 01db5b04ee fix(catalog): omit unsupported reasoning.summary on paid xAI Responses
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
2026-08-14 22:02:07 -07:00
Yang Yang b49b5b88d2 fix(catalog): strip stale xAI Responses effort dials from generated rows
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
2026-08-14 22:02:07 -07:00
Yang Yang 3fb57803f3 fix(catalog): omit penalty and stop params on xAI reasoning models
Grok 4.5 rejects presencePenalty, frequencyPenalty, and stop. After the
paid default moved off a non-reasoning model, configured penalties 400ed.
2026-08-14 22:00:51 -07:00
Yang Yang 7a3a558895 fix(catalog): clamp paid xAI Responses minimal effort to low
Grok 4.5 on XAI_API_KEY kept a minimal dial without SuperGrok's
minimal→low wire map, which can 400 on /v1/responses.
2026-08-14 22:00:50 -07:00
Yang Yang ef7759782d fix(catalog): drop stale xAI Chat Completions model-cache rows
Invalidate cached paid-xAI ids on static fingerprint mismatch so the
Responses migration is not stuck behind a fresh completions cache overlay.
2026-08-14 22:00:50 -07:00
Yang Yang 651f20957b feat(catalog): replay xAI encrypted reasoning on later turns
Stop stripping type=reasoning history for xai and xai-oauth so
encrypted_content from include is sent back on the next Responses request.
2026-08-14 22:00:50 -07:00
Yang Yang c228dea58b feat(catalog): route paid xAI through Responses like SuperGrok
Switch XAI_API_KEY models from Chat Completions to /v1/responses, default
both xai and xai-oauth to grok-4.5, and include reasoning.encrypted_content.
2026-08-14 22:00:50 -07:00
can1357 69a856ea3e Merge PR #8519: fix(catalog): grant low/high/max to OpenRouter deepseek-v4-pro-0813 (@roboomp) 2026-08-14 14:11:38 +02:00
roboomp 2d0eb6c41e fix(catalog): grant low/high/max to OpenRouter deepseek-v4-pro-0813
The OpenRouter non-Flash DeepSeek V4 effort override forced HIGH_ONLY for every id, so getModelDefinedEfforts clamped deepseek-v4-pro-0813 to high even though OpenRouter's /models advertises reasoning.supported_efforts [low,high,max] and the route accepts them. Carve out the dated SKU to the wire-exact low/high/max ladder while keeping the undated deepseek-v4-pro route high-only.

Fixes #8517
2026-08-14 06:18:34 +00:00
roboomp c0394ba53d fix(catalog): scope Copilot model cache by credential
Copilot discovery writes an authoritative cache, so online-if-uncached served the prior endpoint for the full TTL after COPILOT_GITHUB_TOKEN switched accounts. Keying the cache namespace on the credential forces fresh discovery for a new token instead of reusing a stale personal-endpoint cache.

Fixes #8507
2026-08-14 05:12:06 +00:00
roboomp ce65a40539 fix(catalog): bound Copilot endpoint probe with discovery timeout
Threaded the shared 10s discovery AbortSignal into the copilot_internal/user probe so a stalled endpoint falls back to the personal host instead of hanging startup or refresh.

Fixes #8507
2026-08-14 04:47:42 +00:00
roboomp c92ba97538 fix(catalog): discovered Copilot endpoint for env tokens
Shared the plan-endpoint probe between OAuth login and raw token model discovery so Business credentials route to their advertised API host.

Added regression coverage for the raw environment-token path.

Fixes #8507
2026-08-14 04:41:27 +00:00
can1357 0d6a7146a3 feat(coding-agent): supported allSettled and any promise combinators in browser scope
- Added `allSettled` and `any` to tracked promise combinators in run scope.
- Updated browser cancellation tests to cover new promise combinators.
2026-08-14 00:11:04 +02:00
can1357 3ce33d436b feat(catalog): added Gemini Flash model definitions and routing
- Added Gemini 3.7 Flash model definitions and reasoning effort routing configurations.
- Implemented geminiLevelFlashFamily in variant-collapse.ts to manage 3.6+ flash thinking levels.
- Added test coverage for Gemini 3.7 Flash variant collapse and discovery routing.
2026-08-13 19:36:51 +02:00
can1357 b279db1790 test: refactored test suites to eliminate time-based sleeps and polling loops
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
2026-08-13 19:32:22 +02:00
can1357 6b4823181b test: cleaned test suites and documented filtering guidelines
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
2026-08-13 08:28:42 +02:00
roboomp 93795acdd4 fix(catalog): expose low reasoning effort for deepseek-v4-pro
DeepSeek's Chat Completions API now advertises reasoning_effort
low/high/max for both deepseek-v4-flash and deepseek-v4-pro, but the
catalog gated the low tier behind isDeepseekV4FlashModelId, so V4 Pro
(and every non-Flash reasoner) fell through to high/max. Broaden the
DeepSeek effort ladder to give any V4 SKU the low/high/max scale on the
direct API and faithful aggregator routes, keeping OpenRouter's non-Flash
route at high-only and the older V3.x/R1 reasoners at high/max.

Fixes #8405
2026-08-13 07:18:46 +02:00
can1357 a69d166e00 perf(discovery): memoized WSL host-home probe per environment
- Deduplicated the re-imported changelog bullets from the PR #8403 merge,
  keeping the condensed register with only the new #8402 entry.
- getUserHomeCandidates memoizes the WSL home candidate keyed by
  platform + WSL markers + USERPROFILE, so a wedged interop pipe costs
  one bounded probe per process instead of one 500ms stall per
  discovery loader, while env changes (tests, SDK embeddings) still
  recompute.
2026-08-13 06:51:35 +02:00
can1357 2d512cb253 refactor: generalized thinking loop guard for multiple model families
- Generalized thinking loop guard and helper functions to support Gemini, DeepSeek, and Grok model families.
- Replaced `withGeminiThinkingLoopGuard` and related Gemini-specific symbols with generalized counterparts.
- Removed deprecated `enableGeminiThinkingLoopGuard` options and associated tests.
- Updated test suites and agent session logic to use the generalized thinking loop guard and model family tokens.
2026-08-13 06:15:00 +02:00
can1357 021fdc5ba7 feat: added Grok 4.6 thinking-loop guard and model predicate
- Added `isGrok46ModelId` boundary-aware model predicate to the catalog package.
- Included Grok 4.6 models in the thinking-loop guard to prevent runaway reasoning streams.
2026-08-13 05:39:28 +02:00
can1357 2e492a3076 test(tui): corrected stress oracles and bounded-context resize expectation
- The mux pane-growth oracle treated every physical scroll as a logical append, but immutable-history recovery can recommit a corrected suffix after an off-screen mutation without advancing the shadow tape; exempt only changed shared history prefixes.
- The frame-neutral oracle compared prepared rows only; at narrow widths distinct raw rows prepare identically, so the renderer's raw-prefix divergence recovery recommits legitimately. Snapshot raw frames and allow declared transient growth.
- OSC66 spacer preservation intentionally composes six bounded context rows above the resize viewport; assert that exact bound instead of zero above-fold rendering.
Both oracle false positives reproduce identically at the PR head that introduced the harness (4cc9725037); three full randomized stress passes green after the fix.
2026-08-13 02:19:28 +02:00
pickpocket e7e280a6fb fix(catalog): complete GPT-5.6 off and pricing support
(cherry picked from commit fc034d61ff69e8ad1870f674c72ff3f761741862)
2026-08-13 02:00:59 +02:00
pickpocket 37af70b086 fix(catalog): classify Codex Daybreak aliases
(cherry picked from commit eba697926e9fc030d357d170e47856f9e3141dcc)
2026-08-13 02:00:59 +02:00
pickpocket 685055356a feat: add OpenAI Daybreak model support
(cherry picked from commit 889b55bbca0309e9eee6ce0c4287659bfc4fccb3)
2026-08-13 02:00:59 +02:00
can1357 584447f280 Merge PR #8317: fix(catalog): bound openai-compatible model discovery with default timeout (@roboomp) 2026-08-13 02:00:49 +02:00
roboomp 28115cfc1f fix(catalog): applied deepseek effort contract on ollama cloud
Ollama Cloud serves the DeepSeek V4 family over the ollama-chat API,
which bypassed the deepseek effort-ladder branch in model-thinking.ts.
deepseek-v4-flash exposed the generic minimal/low/medium/high/xhigh
ladder with no max tier instead of the low/high/max contract the
catalog encodes on every other host.

Broadened the branch to also cover the ollama-cloud ollama-chat surface:
Flash keeps low/high/max, V4 Pro and older reasoners top out at
high/max. Regenerated models.json accordingly.

Fixes #8334
2026-08-12 09:46:03 +00:00
roboomp 100d2ef547 fix(catalog): bound openai-compatible model discovery with default timeout
Built-in OpenAI-compatible provider managers call fetchOpenAICompatibleModels with neither a signal nor a timeoutMs, and the no-timeout branch issued the request with signal: undefined. A stalled /models endpoint left the fetch pending forever, so createAgentSession's awaited resolveModelDiscoveryFallback discovery pass never returned and startup hung.

Apply a default 10s deadline (DEFAULT_OPENAI_COMPATIBLE_DISCOVERY_TIMEOUT_MS) when the caller supplies neither signal nor timeoutMs, matching the coding-agent remote-discovery budget. Callers passing their own signal keep owning its lifecycle.

Fixes #8315
2026-08-12 05:26:49 +00:00
can1357 94a76a8f27 feat: hardened tar parser and optimize prompt handling
- Bound PAX sparse record memory overhead by caching sparse markers and specific keys.
- Update system prompt phrasing and tests for tool inventory and date displays.
2026-08-12 03:04:07 +02:00
can1357 a4d8860a6c feat: added google reasoning controls mcp stream resumption and tar support
- Added Google provider thinking configuration parameters and force-reasoning-off controls.
- Implemented MCP SSE stream resumption using Last-Event-ID and `SSEResumeError`.
- Added support for TAR old-GNU sparse extension blocks, path length checks, and archive entry overrides.
- Restricted external thinking support to specific models and added semver fallback parsing.
2026-08-12 02:32:45 +02:00
can1357 64baa7c1bd chore(format): applied biome formatting and removed dead code from merged prs 2026-08-11 15:14:15 +02:00
can1357 1340ce6d18 Merge PR #8200: fix(catalog): Fix reasoning levels of GLM-5.2 models; add Baseten GLM 5.2 Fast (@jcfrancisco) 2026-08-11 15:09:34 +02:00
Carlo Francisco aecd4c76e5 fix(catalog): mark Baseten GLM-5.2-Fast as reasoning with high/max effort
Adds zai-org/GLM-5.2-Fast to the Baseten reasoning allowlists, matching
the sibling zai-org/GLM-5.2. Also fixes parseGlmModel to handle uppercase
GLM model IDs (used by Baseten, CoreWeave, HuggingFace, etc.), so the
identity deriver correctly classifies them as GLM-5.2 reasoning models
during catalog generation — previously the case-sensitive regex caused
rebakeModelThinking to fall back to the generic effort ladder.

Regenerated models.json with a live Baseten API key: GLM-5.2 and
GLM-5.2-Fast now bundle reasoning:true with the correct high/max effort
ladder, and other uppercase GLM-5.2 resellers (CoreWeave, HuggingFace,
Synthetic, Together, Wafer) get the corrected minimal..max ladder.
2026-08-10 22:30:11 -04:00
Chen Buskilla 7d7964c14b address review: fix test, changelog and models.json newline
- Update packages/catalog/test/meta-provider.test.ts to expect
  three META_MUSE_STATIC_MODELS entries (1.1, 1.2, contributor)
  with input ["text","image"] — fixes blocking failure.

- Add ## [Unreleased] entry in packages/catalog/CHANGELOG.md per
  AGENTS.md.

- Remove trailing newline from models.json to match
  generate-models.ts (Bun.write without \n).

Co-authored-by: roboomp <roboomp@users.noreply.github.com>
2026-08-09 18:50:43 +03:00
can1357 60d4cb997e test(catalog): pinned copilot grok-4.5 migration to the responses route
- The regenerated bundle (merged with PR #8021) now ships a
  responses-route github-copilot/grok-4.5, so the id legitimately
  resurfaces from the bundle when the migration refresh fails; the
  contract worth defending is that the stale cached completions route
  never returns and the unbundled long-context variant stays dropped.
2026-08-08 20:57:10 +02:00
can1357 bf04fbfc8d fix(catalog): routed opencode-go deepseek-v4-flash through the responses api
- The OpenCode Go gateway does not serve DSV4-Flash at
  /zen/go/v1/chat/completions; /zen/go/v1/responses works (user-verified
  against the live gateway). Added a per-id override in
  OPENCODE_GO_API_RESOLUTION so both bundled generation and the runtime
  /v1/models refresh route it to openai-responses; deepseek-v4-pro keeps
  chat completions.
- Regenerated models.json from the resolver source.
2026-08-08 20:57:10 +02:00
can1357 ac55ea7697 fix(catalog): toggle qwen3.8 max thinking on wire 2026-08-08 19:38:31 +02:00
roboomp 4d2c6e37f1 fix(catalog): preserved token plan preview vision
Applied curated Alibaba Token Plan seeds after generic models.dev fallback so bundled capabilities cannot be overwritten by incomplete upstream metadata.
2026-08-08 14:45:44 +00:00
roboomp 155fdaedba fix(catalog): corrected qwen3.8 max discovery metadata
Curated reasoning, multimodal input, context limits, and the provider-specific effort ladder for the discovered Alibaba Token Plan model.

Fixes #8019
2026-08-08 14:37:24 +00:00
can1357 9ab6ea6c8f test(catalog): cover discovered Token Plan limits 2026-08-07 13:37:55 +02:00
can1357 ded3bbfce0 Merge PR #7849: fix(catalog): enrich Alibaba Token Plan discovered model limits (@Mustaqeem66) 2026-08-07 13:37:55 +02:00
can1357 0394bf4a29 Merge PR #7865: fix(catalog): update Devin reasoning family routing (@will-bogusz) 2026-08-07 13:37:55 +02:00
Voon Foo 2f24d4457e fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).

- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
  a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
  Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
  tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
  and lazy terminal errors carry the structural errorId classification
  so session auto-retry classifies stalls without text matching.
2026-08-07 14:18:46 +08:00
Will 1ad85b5584 fix(catalog): collapse current Devin reasoning families 2026-08-06 21:36:11 -04:00
Will 7e95b61ffb fix(catalog): route Devin GPT-5.6 fast max effort 2026-08-06 21:35:56 -04:00
Mustaqeem66 b80a888dc8 fix(catalog): enrich Alibaba Token Plan discovered model limits 2026-08-06 17:12:21 +00:00
roboomp e97d1fd21a test(catalog): removed tautological effort assertions 2026-08-05 02:34:04 +00:00
roboomp 736b496cc6 fix(catalog): exposed low effort tier for deepseek-v4-flash
DeepSeek's API accepts reasoning_effort low/high/max and only
deepseek-v4-flash supports all three tiers (V4 Pro is high/max). The
identity-derived effort fallback blanket-applied high/max to every
direct-DeepSeek reasoning model, hiding the low tier flash accepts.

Added isDeepseekV4FlashModelId and route flash to the low/high/max ladder
on every host; non-flash DeepSeek keeps high/max (high-only on OpenRouter).

Fixes #7668
2026-08-05 02:22:27 +00:00