Commit Graph

36 Commits

Author SHA1 Message Date
Yang Yang 86866b8bbf fix(catalog): omit Responses penalties on all first-party xAI models
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
2026-08-14 22:02:52 -07:00
Yang Yang 09830d2bd6 fix(catalog): drop unsupported xhigh effort from first-party Grok
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
2026-08-14 22:02:52 -07:00
Yang Yang 01db5b04ee fix(catalog): omit unsupported reasoning.summary on paid xAI Responses
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
2026-08-14 22:02:07 -07:00
Yang Yang 3fb57803f3 fix(catalog): omit penalty and stop params on xAI reasoning models
Grok 4.5 rejects presencePenalty, frequencyPenalty, and stop. After the
paid default moved off a non-reasoning model, configured penalties 400ed.
2026-08-14 22:00:51 -07:00
Yang Yang 7a3a558895 fix(catalog): clamp paid xAI Responses minimal effort to low
Grok 4.5 on XAI_API_KEY kept a minimal dial without SuperGrok's
minimal→low wire map, which can 400 on /v1/responses.
2026-08-14 22:00:50 -07:00
Yang Yang 651f20957b feat(catalog): replay xAI encrypted reasoning on later turns
Stop stripping type=reasoning history for xai and xai-oauth so
encrypted_content from include is sent back on the next Responses request.
2026-08-14 22:00:50 -07:00
Yang Yang c228dea58b feat(catalog): route paid xAI through Responses like SuperGrok
Switch XAI_API_KEY models from Chat Completions to /v1/responses, default
both xai and xai-oauth to grok-4.5, and include reasoning.encrypted_content.
2026-08-14 22:00:50 -07:00
pickpocket e7e280a6fb fix(catalog): complete GPT-5.6 off and pricing support
(cherry picked from commit fc034d61ff69e8ad1870f674c72ff3f761741862)
2026-08-13 02:00:59 +02:00
can1357 3177f6bdf7 test(catalog): cover DeepSeek policy without bundled models 2026-08-02 20:55:16 +02:00
roboomp 2cfaeb6116 fix(catalog): honored openrouter deepseek effort metadata
- Parsed OpenRouter reasoning effort ladders and defaults during discovery.

- Preserved explicit thinking metadata from models.yml patches.

- Regenerated the catalog and covered both regression paths.

Fixes #7307
2026-08-01 19:06:44 +00:00
can1357 93ddcccc47 Merge PR #6622: fix(catalog): retry empty model discovery (@roboomp) 2026-07-26 15:41:26 +02:00
usr-bin-roygbiv ab35562418 fix(catalog): infer runtime Azure computer use 2026-07-25 21:21:59 +00:00
roboomp 95bd9923c6 fix(catalog): keep empty discovery authoritative for pruning
Decoupled cache-retry authority from result authority: an empty successful fetch stays authoritative for the cycle so ModelRegistry prunes removed models, while the cache row is written non-authoritative so the short retry still recovers transient empties.

Fixes #6620
2026-07-25 14:23:56 +00:00
roboomp 423fdd6c67 fix(catalog): retried empty model discovery
Marked empty successful dynamic catalogs non-authoritative so the existing short retry interval can recover transient provider startup and ACL states.

Fixes #6620
2026-07-25 14:14:51 +00:00
usr-bin-roygbiv 91c31c18aa fix(catalog): invalidate stale capability caches 2026-07-25 03:04:20 +00:00
usr-bin-roygbiv e33f665e4c fix(catalog): retain computer capability provenance 2026-07-25 02:50:10 +00:00
usr-bin-roygbiv 9bab62ee55 fix(computer-use): address routing review gaps 2026-07-25 02:14:36 +00:00
usr-bin-roygbiv c12c282aa5 fix(computer): gate native Responses transport 2026-07-25 00:42:23 +00:00
Alexander Kirilin 944ff76b77 test(openai): cover future cache boundaries 2026-07-23 16:48:40 -04:00
Alexander Kirilin 8906595c49 fix(openai): preserve explicit prompt-cache boundaries 2026-07-23 16:39:10 -04:00
Alexander Kirilin 7ae6ebfa8b feat(catalog): gate OpenAI prompt cache breakpoints 2026-07-23 15:18:37 -04:00
roboomp b823c18837 fix(catalog): guarded request-model header recovery
Versioned request-header restoration metadata inside v10 cache rows so only markers written by the old id-only matcher can bypass an unrestorable marker through requestModelId. Current aliases whose live headers differ from their static base remain unresolved and are refetched or dropped.

Added catalog and startup-registry regressions for custom-header aliases while preserving legacy Copilot -1m cache recovery.

Fixes #6284
2026-07-22 11:10:23 +00:00
roboomp d9bab7ce8b fix(catalog): restore cached request-model variants via requestModelId
Copilot -1m long-context variants are synthesized with transport
headers and a requestModelId to a bundled base. The v10 cache omits
headers; the writer only matched a same-id static entry, so these
variants were flagged unrestorable and dropped on the next offline
read, vanishing from the picker with a "Could not restore model"
warning. The startup registry loader dropped them the same way.

Restore/match headers through requestModelId in the cache writer, the
model-manager restore path, and the coding-agent startup loader, and
bypass a stale unrestorable marker written by the old id-only writer.

Fixes #6284
2026-07-22 10:59:10 +00:00
can1357 401fca701b Merge PR #5448: fix(coding-agent): rescue snapcompact dead-end when nothing is summarizable (@roboomp) 2026-07-18 21:01:41 +02:00
roboomp cc688c6b5c fix(catalog): restored headers omitted from model cache
Recorded which cached model ids had headers omitted and which header sets cannot be reconstructed from current static inputs.

Fresh/offline cache reads now restore exact static headers before returning models. Dynamic-only or dynamically-augmented header models bypass fresh-cache reuse and refetch online; offline and failure fallbacks omit them instead of returning unusable models without required transport headers.

Bumped the cache schema to v10 and added static, dynamic, offline, and raw-persistence regressions.

Fixes #5780
2026-07-17 03:27:12 +00:00
can1357 b865e6a4d7 fix(catalog): scope Kimi K2.7 timeout to Moonshot 2026-07-16 03:32:02 +02:00
can1357 9c78b7e7b0 merged PR #5428: fix(catalog): extend streamIdleTimeoutMs floor to Kimi K2.7 Code 2026-07-16 03:32:02 +02:00
roboomp 0069a84c9f fix(catalog): apply local stream-timeout floor to loopback proxies
A litellm proxy on a loopback baseUrl fronting a local llama.cpp/vLLM server
was excluded from isLocalOpenAICompatBackend (PROXY_OPENAI_COMPAT_PROVIDERS)
so replayReasoningContent stays off proxies that may forward to an unrelated
cloud upstream. That exclusion also stripped the 300s LOCAL_OPENAI_COMPAT
stream-timeout floor, so the first-event budget fell back to the 100s default.
A slow prefill on a large prompt (llama.cpp "non-consecutive token position"
KV thrash + reprocess) then exceeded 100s, aborting with "OpenAI completions
stream timed out while waiting for the first event" and retry-looping.

Decouple the stream-timeout floor from the replay gate: the floor now applies
to any loopback/RFC1918 backend (including proxies) because widening the abort
ceiling only helps a slow local upstream and never forwards an extra wire
field. Reasoning replay stays gated to first-party local providers. Applied to
both the completions and responses compat builders.

Fixes #4786
2026-07-15 14:32:39 +00:00
roboomp 29c8cae9ba fix(catalog): extend stream idle floor to kimi k2.7 code
Native Kimi K2.7 Code (kimi-k2.7-code / kimi-k2.7-code-highspeed) reasons
for minutes before the first stream event like K2.6, but the
streamIdleTimeoutMs branch in buildOpenAICompat gated only on
isKimiK26ModelId, so K2.7 Code fell through to the 120s default and
aborted on long reasoning turns. Match matchesKimiK27CodeFamily in the
same branch and rename the constant to KIMI_REASONING_STREAM_IDLE_TIMEOUT_MS.

Fixes #4836
2026-07-14 16:39:54 +00:00
roboomp b7aa046ed6 fix(catalog): preserved static limits on cache mismatch
- Sanitized same-id cached model limits when the cache fingerprint no longer matches the current static catalog.

- Added regression coverage for offline resolution with stale cache rows and updated static limits.

Fixes #4956
2026-07-09 17:46:27 +00:00
can1357 c4c0331345 fix(coding-agent/session): prevented data loss in session serialization
- Persist signed message blocks (`text`, `thinking`, `toolCall`) and encrypted reasoning payloads verbatim during session serialization instead of clearing or truncating them.
- Preserve signature keys instead of replacing them with empty strings when they exceed persistence size limits.
- Exempt official first-party OpenAI and Anthropic API endpoints from the leaked-thinking stream healing wrapper to prevent misfires on legitimate visible text fences.
2026-07-02 03:58:10 +02:00
can1357 221f4102fb fix: resolved Fireworks Qwen models to openai thinking format
- Updated `buildOpenAICompat` to override the `qwen` thinking format for Fireworks-hosted models, ensuring they use `openai` thinking parameters instead.
- Prevented invalid `enable_thinking` payload errors by ensuring Fireworks-hosted Qwen requests conform to their strict schema.
- Updated `AgentSession` to allow Fireworks fast-fallback logic to execute even when standard retries are disabled.
2026-06-20 09:25:36 +02:00
can1357 291b3c74c2 feat: enhanced model reasoning, schema normalization, and loop guarding
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
2026-06-18 04:51:43 +02:00
can1357 d4317d3d20 feat(ai): consolidated OpenAI-family streaming and add OpenRouter API support
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.

Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
    - Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
    - Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
    - Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
2026-06-17 21:36:48 +02:00
can1357 f30ec6e089 feat(catalog): stripped gateway prefixes and promo tags from model display names
- Added `cleanModelName` to `utils.ts`, dropping gateway author prefixes (`OpenAI: …`), `(latest)` alias markers, `(Antigravity)` attribution, price tiers (`($$$$)`), and promo/lifecycle tags (`(20% off)`, `(retires …)`) while preserving variant tags that map to distinct wire ids (`(Thinking)`, `(free)`, `(Fast)`, dates, regions).
- Applied it in `buildModel` (covers live discovery and stale caches) and as a display-name normalization pass in `generate-models.ts`; Antigravity discovery no longer appends `(Antigravity)` to display names.
- Added name-cleaning coverage to `build.test.ts`.
- Changelog entry for this change landed with the variant-collapse commit (same contiguous `CHANGELOG.md` run).
2026-06-12 07:37:38 +02:00
can1357 ae415199dc feat: added build-time compatibility in ModelSpec/buildModel pipeline
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
2026-06-10 06:20:51 +02:00