Commit Graph

135 Commits

Author SHA1 Message Date
can1357 5a044dc0da feat: consolidated and automate changelog management
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
2026-06-27 03:08:52 +02:00
can1357 f6e7d8ebc7 Merge PR #3574: feat(coding-agent): discover rich LiteLLM proxy metadata (@jdavv) 2026-06-27 01:39:32 +02:00
can1357 3e5360ca4c Merge PR #3060: feat(provider): add GitLab Duo Agent provider (@jiwangyihao) 2026-06-27 01:39:31 +02:00
Jean-Luc Davern 3dc9099058 feat(coding-agent): discover rich LiteLLM metadata 2026-06-26 11:26:27 -05:00
can1357 d7f05ac640 fix(catalog): scoped CoreWeave catalog metadata 2026-06-26 17:12:01 +02:00
Lance Tuller efdcadf0d5 feat(ai): add CoreWeave Serverless Inference provider 2026-06-26 09:13:41 -04:00
jiwangyihao 6d280d4349 fix(gitlab-duo): 命名空间发现翻页遍历顶层 group 并放宽 SSH 远程端口比较
处理两条 Codex 审阅意见:(1) fetchTopLevelGroupNamespaceCandidates 只取第一页,token 属于 >100 个顶层 group 时后续页的可用 Duo namespace 不可见,现在跟随 GitLab x-next-page 分页(受 GITLAB_DUO_WORKFLOW_MAX_GROUP_PAGES 限制)遍历所有页再校验候选;(2) 自管 GitLab 常把 SSH 暴露在独立端口(ssh://git@host:2222/...)而 web 是 https://host,原 host:port 严格比较会误拒该远程,现在 SSH/scp 远程按裸 hostname 比较,HTTP(S) 远程仍严格比 host:port 以区分同主机不同服务。各加 1 个发现测试。
2026-06-26 16:13:27 +08:00
jiwangyihao bf1b692df9 fix(gitlab-duo): direct_access 错误保留 HTTP 状态码以触发凭据轮换
机器人指出两处问题,本提交处理:

1) direct_access 带 body message 的错误丢弃了 response.status,导致
   streaming auth-retry 路径(extractStatusFromAssistantError ->
   extractHttpStatusFromError)无法恢复状态、无法刷新/轮换过期 OAuth 或
   配额受限的 broker 凭据。现在即便有 body message 也内嵌 HTTP <status>,
   与 create 路径既有约定一致。新增 401 Unauthorized 回归测试断言状态可恢复,
   并强化既有 403 配额测试。

2) generate-models 种子注释过度暗示生成期会跑 namespace-scoped 发现。
   实际上 gitlab-duo-agent descriptor 故意不带 catalogDiscovery,被
   isCatalogDescriptor 过滤排除在生成发现循环之外,因此生成期绝不会拉取
   某账号的 aiChatAvailableModels,只播种通用 namespace-free fallback。
   修正注释表述,新增两条针对 descriptor 的回归测试:断言 descriptor 无
   catalogDiscovery(永不参与生成发现),以及 fallback 模型不带
   gitlabDuoWorkflowRootNamespaceId。
2026-06-26 16:13:24 +08:00
jiwangyihao 0e78a24443 fix(catalog): 将 GitLab Duo Agent fallback 模型纳入 bundled models.json
机器人指出 gitlab-duo-agent 不在 models.json,fresh 安装(尚无带凭据的动态
发现/缓存)时内置 catalog 看不到默认模型。generator 现按 Sakana 同样方式播种
gitlab-duo-agent 的 fallback 模型(claude_sonnet_4_6_vertex):live aiChatAvailableModels
发现成功时其条目按 id 去重胜出,仅在无凭据/失败 regen 时落种子。

- generate-models.ts:导入 buildGitLabDuoWorkflowFallbackModel,在 authoritative
  发现未命中时 push 种子。
- 重新生成 models.json,仅保留 gitlab-duo-agent provider 块的新增,其余 provider
  数据维持基线不变(regen 在无凭据环境下未触碰其它 provider)。
- 新增针对 descriptor 的回归测试(非 bundled JSON):断言 manager options 暴露
  fallback 静态模型,符合 AGENTS.md 要求。
2026-06-26 16:13:24 +08:00
jiwangyihao 20e4c3bd5b fix(ai): 处理 Codex 对项目自动发现、设置启用与 changelog 归属的三项审阅意见
- 自动项目发现优先复用命名空间解析所依据的项目(工作区 git remote 或显式
  项目),而不是从 group 列表泛选;多项目 group 下不再把工作流 scope 到无关项目。
  命名空间选择沿用 remote/project 来源的 projectPath,运行时据此 scope。
- Duo 设置启用仅在确定性尝试(任意 HTTP 响应,含 4xx)后才标记 ensured;
  瞬时网络错误/5xx 返回 false,使后续 turn 可重试,避免 fresh 命名空间因一次
  瞬时失败而永久跳过 PUT。
- agent CHANGELOG 把 cwd 转发条目从已发布 [16.1.8] 段移至 [Unreleased];
  移动后 [16.1.8] 与上游 main 一致。
2026-06-26 16:13:24 +08:00
jiwangyihao 3298c35eec fix(catalog): 处理 Codex 对 Duo 模型缓存与远程端口比较的两项审阅意见
- 动态模型缓存键改为镜像 discoverGitLabDuoWorkflowNamespace 真实解析输入:
  内置发现只传 apiKey/baseUrl/fetch,原键里的 namespaceId/projectId/cwd 恒为
  空,退化为 apiKey+baseUrl,导致同 token 下两个不同 group 的工作区互相复用
  权威模型缓存。改为按 (凭据, baseUrl, 命名空间/项目 config 或同名环境变量,
  有效 cwd) 指纹分区。
- 远程 host 比较改用 host(含端口)而非 hostname:同主机不同端口的自托管 GitLab
  不再被误判为同实例,避免从错误远程推导项目路径。SCP 远程无端口概念,仍按裸
  host 比较。新增跨端口远程不被当作工作区项目的回归测试。
2026-06-26 16:13:23 +08:00
jiwangyihao 65108346e9 feat(catalog): add GitLab Duo Agent model discovery 2026-06-26 16:13:20 +08:00
can1357 3188fc6bb9 Merge remote-tracking branch 'origin/farm/371a49e5/anthropic-claude-4-5-budget-effort' 2026-06-25 22:35:22 +02:00
roboomp 7a4cf0e30c fix(catalog): classified direct Anthropic Sonnet/Haiku 4.5 as plain budget thinking
Direct Anthropic Claude Sonnet 4.5 and Haiku 4.5 (plus their Cloudflare,
Vertex, GitLab-Duo, Copilot, OpenCode-Zen, and Bedrock cross-region passes)
were classified as anthropic-budget-effort, which made the Anthropic provider
serialize output_config.effort alongside the thinking.budget_tokens block.
Anthropic only honors output_config.effort on Opus 4.5 and adaptive (4.6+)
Messages-API models — Sonnet 4.5 and Haiku 4.5 reject every request with
HTTP 400 'This model does not support the effort parameter.', so the advisor
(and any agent on a Sonnet/Haiku 4.5 SKU) failed every turn.

inferThinkingControlMode now gates anthropic-budget-effort to
parsedModel.kind === 'opus' && semverGte(version, '4.5') on both
anthropic-messages and bedrock-converse-stream. Sonnet/Haiku 4.5 fall through
to mode: 'budget' (effort still scales the per-tier thinking budget via
ANTHROPIC_THINKING[reasoning]); Opus 4.5 keeps anthropic-budget-effort and
continues to emit output_config.effort. anthropic-budget-effort also remains
in use for Anthropic-compatible third-party backends that natively support
the field (Umans GLM 5.2).

Regenerated models.json so the 19 Sonnet/Haiku 4.5 first-party entries flip
to mode: 'budget' and the 12 Opus 4.5 entries stay on anthropic-budget-effort.

Regression tests cover both sides: anthropic-alignment.test.ts asserts the
Sonnet 4.5 wire body omits output_config and the Opus 4.5 wire body emits
output_config.effort: 'medium'.

Fixes #3497
2026-06-25 20:16:34 +00:00
roboomp 80862b79da fix(agent): handled ollama-cloud task backoff
Added ollama-cloud subagent concurrency limiting, role fallback-chain inheritance, and visible empty length errors for native Ollama responses.

Fixes #3464
2026-06-25 11:57:51 +00:00
can1357 0eb21efa1a Merge remote-tracking branch 'origin/farm/7ef98714/snapcompact-copilot-vision-gate' 2026-06-24 21:00:37 +02:00
can1357 4c4fb21caa Merge remote-tracking branch 'origin/farm/cf20a2c7/clamp-ollama-cloud-max-output-tokens' 2026-06-24 21:00:03 +02:00
roboomp 0edb3d32b8 fix(ai/ollama): clamped num_predict at the 65536 Ollama Cloud cap
Ollama Cloud rejects any chat request whose options.num_predict exceeds
65536 with HTTP 400, but the existing safety relied on the load-time
omitMaxOutputTokens policy in model-registry.ts. Stale models.db rows
predating that policy (and custom modelOverrides re-enabling output
caps) carried maxTokens: 1048576 forward to the wire layer untouched,
so every request to deepseek-v4-pro / deepseek-v4-flash 400'd with:

  max_tokens (1048576) exceeds model's maximum output tokens (65536)
  for model deepseek-v4-pro

createChatBody now resolves num_predict through resolveNumPredict,
which clamps every ollama-cloud request at the documented cap before
serialization — independent of the cached spec, on top of the existing
omitMaxOutputTokens path. Self-hosted ollama traffic is unaffected.

Fixes #3392
2026-06-24 17:29:28 +00:00
can1357 a4dbe6396c style: biome organize-imports on merged test files (#3193, #3312) 2026-06-24 19:03:32 +02:00
roboomp 714051d795 fix(catalog,coding-agent): disable vision on non-personal copilot endpoints
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.

- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
  forces input=['text'] whenever the resolved baseUrl is not the
  canonical personal-Copilot host, so the upstream's vision flag is
  honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
  OR-upgrading with the bundled reference) when the merged baseUrl
  differs from the bundled one, so a bundled spec pinned to the
  personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
  rasterizer for any github-copilot model whose baseUrl is non-personal,
  catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
  pi-catalog/wire/github-copilot so catalog and coding-agent share one
  canonical check.

Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).

Fixes #3387
2026-06-24 16:27:14 +00:00
can1357 7408ee75f2 Merge PR #3193: fix(catalog): restore Umans GLM-5.2 max reasoning (@roboomp)
# Conflicts:
#	packages/ai/test/anthropic-alignment.test.ts
#	packages/catalog/test/umans-provider.test.ts
2026-06-24 18:26:00 +02:00
can1357 131d31f6ae feat(catalog): added cache invalidation for static model metadata
- Update `sakanaModelManagerOptions` to provide `dropCachedModelIdsOnStaticMismatch` configuration using identified static model IDs.
- Add a test case to verify that stale cached model metadata is correctly cleared when bundled model specs are updated.
2026-06-22 08:07:47 +02:00
can1357 4e54e557bc feat: improved context usage display and update model configurations
- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
2026-06-22 08:04:42 +02:00
can1357 89fe7b21a6 feat: added support for Sakana AI and Fugu provider
- Implemented Sakana AI and Fugu provider integration including authentication, API base URL resolution, and dynamic model discovery.
- Configured static model definitions and reasoning metadata for the Fugu model series within the catalog.
- Added environment variable support for API configuration and base URL overrides via `SAKANA_*` and `FUGU_*` variables.
- Verified service integration and provider registry registration through comprehensive test suites in both AI and catalog packages.
2026-06-22 06:55:18 +02:00
roboomp 2cbd26a15b test(catalog): used effort submodule import
Updated the Umans provider regression test to import Effort from the catalog effort submodule instead of the root catalog barrel, keeping the package-local test on the narrow dependency path.

Refs #3192
2026-06-21 12:45:34 +00:00
roboomp e277f3a310 fix(ai): emitted anthropic budget effort levels
Forward anthropic-budget-effort selections through mapOptionsForApi so buildParams serializes output_config.effort with budget-token thinking. Mark Umans GLM-5.2 as budget-effort and map the UI xhigh tier back to Umans's max wire value.

Refs #3192
2026-06-21 12:39:05 +00:00
roboomp 1e01d57d09 fix(catalog): dropped dead umans glm-5.2 effortmap
Removed the xhigh -> max effortMap baked onto the Umans GLM-5.2 spec; mode="budget" routes thinking depth via thinking.budget_tokens and never consults effortMap, so the wire shape is unchanged. The picker still surfaces high and xhigh via getModelDefinedEfforts, and the dynamic discovery still maps the upstream max level to Effort.XHigh.

Refs #3192
2026-06-21 12:26:20 +00:00
roboomp 3a888c0f59 fix(catalog): restored umans glm max reasoning
Mapped Umans GLM-5.2's upstream max reasoning level to the internal xhigh effort and preserved the max wire value in dynamic discovery and bundled catalog metadata.

Added resolver coverage for the high/max ladder and verified the xhigh request maps back to max.

Fixes #3192
2026-06-21 12:13:30 +00:00
roboomp 77b146518d style: bun run fix 2026-06-21 10:42:05 +00:00
roboomp 12ab84be91 fix(catalog/umans): dropped stale GLM cache rows
Upgraded installs can carry an Umans model cache written before the
GLM via-handoff catalog correction. If dynamic discovery is skipped or
fails, resolveProviderModels merges cache rows over the corrected static
catalog; mergeDynamicModel preserves image support when either side has
it, so a stale cached ["text", "image"] GLM row can re-add native image
support until the next successful refresh.

Add a provider-scoped cache drop hook for model ids whose cached rows
are unsafe across static fingerprint changes, opt Umans into it for
`umans-glm-5.1` and `umans-glm-5.2`, and cover the offline stale-cache
upgrade path with a regression test.

Fixes #3184
2026-06-21 10:41:58 +00:00
roboomp 565aba94bc fix(catalog/umans): treated supports_vision sentinels as text-only
`umans-glm-5.1` / `umans-glm-5.2` advertise themselves on the Umans
`models/info` endpoint with `supports_vision: "via-handoff"`. That
sentinel means image inputs are routed through a separate vision
handoff pre-analysis step; the GLM endpoint itself rejects raw image
blocks with `400 This model does not support image inputs`.

`umansSupportsVision` was returning `true` for any non-empty string,
so dynamic discovery mapped the GLM models to `input: ["text",
"image"]` and the agent sent images straight to GLM. The bundled
`umans-glm-5.1` / `umans-glm-5.2` rows in `models.json` carried the
same stale `["text","image"]` from a previous regen.

- Tighten `umansSupportsVision` to `value === true`; document the
  sentinel contract.
- Correct the two bundled rows to `input: ["text"]` so the vision
  handoff path runs.
- Add resolver- and bundle-level regression tests that cover
  `supports_vision: "via-handoff"` alongside the native-vision
  `umans-coder` case.

Fixes #3184
2026-06-21 10:32:53 +00:00
can1357 0b5e2d8276 fix(ai-providers): normalized Anthropic tool call IDs
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
2026-06-21 06:04:49 +02:00
roboomp 1495779e52 fix(catalog): preserved litellm discovery transport
- Kept provider discovery defaults authoritative for api and provider when layering bundled or models.dev metadata.
- Added a LiteLLM regression covering a deepseek-v4-flash collision with the ollama-cloud catalog.

Fixes #3162
2026-06-21 02:20:41 +00:00
can1357 320f514f77 fix(ai): adjusted llama-cpp base url and refactor token validation
- Update `llama.cpp` base URL to remove the `/v1` suffix.
- Consolidate local provider token validation in `coding-agent` using a `Set`.
- Add test coverage to ensure catalog model IDs are preserved verbatim on the wire.
2026-06-21 02:45:48 +02:00
oldschoola 13bfc7b9c0 Remove Wafer Pass provider 2026-06-21 02:45:48 +02:00
can1357 7fcb85f812 chore: updated changelogs & added moonshot tests 2026-06-21 02:29:57 +02:00
can1357 b128406030 Merge PR #2994: fix(catalog): clamp MiMo reasoning efforts (@riverpilot) 2026-06-20 22:12:49 +02:00
Alexander Kirilin eaf9248ec2 fix(catalog): merge main into MiMo efforts 2026-06-20 03:27:05 -04:00
can1357 221f4102fb fix: resolved Fireworks Qwen models to openai thinking format
- Updated `buildOpenAICompat` to override the `qwen` thinking format for Fireworks-hosted models, ensuring they use `openai` thinking parameters instead.
- Prevented invalid `enable_thinking` payload errors by ensuring Fireworks-hosted Qwen requests conform to their strict schema.
- Updated `AgentSession` to allow Fireworks fast-fallback logic to execute even when standard retries are disabled.
2026-06-20 09:25:36 +02:00
roboomp 02bd026d5e fix(catalog): pinned MiniMax-M3 contextWindow to 1M on minimax-code(-cn) providers
The MiniMax-M3 long-context policy in generated-policies.ts only
covered the anthropic-messages providers `minimax` and `minimax-cn`.
The MiniMax Coding/Token Plan (international and China) endpoints
serve the same model through `minimax-code` and `minimax-code-cn`
on openai-completions, and shipped with the upstream 512K pricing
boundary baked into models.json. Switching to MiniMax-M3 under the
Coding Plan therefore still showed a 512K context window in the
status bar.

Broadens the policy carve-out to all four providers, re-bakes both
affected entries in the bundled models.json, and extends the
generated-policies / bundled-catalog tests to assert 1M for the two
newly covered providers.

Fixes #3097
2026-06-20 03:07:50 +00:00
Alexander Kirilin 057b39fc46 fix(catalog): resolve MiMo branch conflicts 2026-06-19 20:28:05 -04:00
roboomp 47cc464962 fix(catalog): retired stale Claude 4.6 wire ids and healed the bundled Sonnet route
Follow-up to #3071 review (codex + @PGupta-Git): the family rewrite alone
did not heal users running off the bundled catalog or stale SQLite cache
rows, which still routed Sonnet 4.6 thinking efforts to the 404
`claude-sonnet-4-6-thinking` wire id. `collapseEffortVariants` treats
those collapsed snapshots as authoritative and `refreshCollapsedThinking`
exits early for families without `effortBudgets` (Claude pairs), so the
new empty routing in the hand table never reached the snapshot.

- Declared the dead wire ids as `retiredMembers` on the Claude 4.6
  families (`claude-sonnet-4-6-thinking` on the Sonnet family,
  `claude-opus-4-6` on the Opus family). This triggers
  `reconcileRetiredRouting` to rewrite every `effortRouting` entry that
  targets a retired id to a live wire id (Sonnet falls back to the bare
  member; Opus falls back to `-thinking`).
- Refreshed the bundled `packages/catalog/src/models.json` Sonnet 4.6
  entry so fresh installs do not boot with the dangling routing — the
  surgical diff matches what the generator would emit; the rest of the
  catalog is left untouched to keep the bug-fix PR scoped.
- Added regression tests in `variant-collapse.test.ts` for both
  reconciliation paths and a bundled-catalog test
  (`issue-3067-repro.test.ts`) that pins the live-wire-id resolution end
  to end through `buildModel` for every effort tier.

Fixes #3067
2026-06-19 22:25:48 +02:00
roboomp cbb9c96c64 fix(providers): aligned Antigravity Claude 4.6 routing with the live backend
The daily Cloud Code Assist backend (`daily-cloudcode-pa`) exposes Claude 4.6
asymmetrically: `claude-sonnet-4-6` has no `-thinking` twin and
`claude-opus-4-6` has only the `-thinking` twin. The shared
`thinkingPair("claude-sonnet-4-6", …)` family (with `preserveAbsentEffortRoutes`)
kept every effort routed to `claude-sonnet-4-6-thinking` even when discovery
only returned the bare id, so any reasoning-on request 404'd with
`Requested entity was not found`. Claude on Antigravity also caps
`maxOutputTokens` at 64000, while `ANTIGRAVITY_MODEL_WIRE_PROFILES` had no
Claude entries — discovery's 65536 propagated to the wire and 400'd with
`Request contains an invalid argument`.

- Replaced the two `thinkingPair` calls for Claude 4.6 in `SHARED_CCA_FAMILIES`
  with bespoke single-wire families. Sonnet collapses to the bare wire id,
  Opus collapses to the `-thinking` wire id, and per-effort thinking is
  carried by the request body's `thinkingBudget` on the single shared wire id.
  Listing both candidate ids in `members` (priority order) keeps the collapse
  correct if the backend mix ever rebalances.
- Added `claude-sonnet-4-6` and `claude-opus-4-6-thinking` entries to
  `ANTIGRAVITY_MODEL_WIRE_PROFILES` capping `maxOutputTokens` at 64000.
- Made `AntigravityModelWireProfile.modelEnum` optional — Anthropic-backed
  wire ids are accepted without a captured `labels.model_enum` token. The
  request builder now emits the label only when the profile defines one.
- Regression tests in `variant-collapse.test.ts` (routing/wire-id resolution
  for both 4.6 families across all three discovery permutations) and
  `google-gemini-cli-alignment.test.ts` (request builder caps Claude
  `maxOutputTokens` at 64000 and omits the unset `model_enum` label).

Fixes #3067
2026-06-19 18:33:57 +00:00
can1357 40101bd3eb Merge PR #3043: fix(catalog): omit Ollama Cloud output caps (@wolfiesch)
# Conflicts:
#	packages/catalog/src/models.json
2026-06-19 17:15:43 +02:00
Alexander Kirilin 7baa4e202e Merge remote-tracking branch 'origin/main' into fix/2864-mimo-efforts
# Conflicts:
#	packages/catalog/src/model-thinking.ts
#	packages/catalog/src/models.json
#	packages/catalog/src/variant-collapse.ts
#	packages/catalog/test/variant-collapse.test.ts
#	packages/coding-agent/test/model-registry.test.ts
2026-06-19 10:29:52 -04:00
can1357 ab6d1f916b feat(catalog): adjusted GLM-5.2 reasoning effort mapping per host
- Refactor GLM-5.2 effort mapping to accommodate specific requirements for Z.ai, OpenRouter, and general OpenAI-compatible hosts.
- Apply host-specific logic to ensure `xhigh` UI tiers are correctly resolved to the required `max` budget for supported providers.
- Add test coverage verifying expected effort mappings across different model hosts.
2026-06-19 16:06:16 +02:00
can1357 0dfeac8a75 feat: added auth discovery broker and expand model support
- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.
2026-06-19 16:06:16 +02:00
Wolfgang Schoenberger 409af331e7 fix(catalog): omit Ollama Cloud output caps 2026-06-19 02:51:23 -07:00
can1357 828699ffdf test(catalog): drop evaluator-worktree note from copilot signing test 2026-06-18 23:01:11 +02:00
can1357 2af64d5634 fix(catalog): treat github-copilot anthropic proxy as a signing endpoint (#2851)
GitHub Copilot's `anthropic-messages` proxy (api.githubcopilot.com) forwards
to signature-enforcing Anthropic and returns full thinking signatures, but the
compat builder classified it as a non-signing reasoning endpoint via the
generic `reasoning && !official` default (`replayUnsignedThinking: true`).

When a checkpoint/branch-return turn is an abandoned tool-use turn (adaptive
Opus emits a tool call then ends on `stop`/`end_turn`), `transformMessages`
correctly strips its end_turn-bound, unreplayable signature. On a
`replayUnsignedThinking` endpoint the encoder then re-emitted that block as
`{ type: "thinking", signature: "" }`. An empty signature is rejected by the
signature-enforcing backend with `400 Invalid signature`, which corrupts the
session and re-trips on every full history re-send (e.g. after toggling MCP
servers).

Exclude github-copilot from `replayUnsignedThinking` so unsigned/stripped
thinking degrades to text exactly like the official Anthropic API — wire-valid
and lossless of the tool_use pairing. Z.AI / DeepSeek / other 3p reasoning
endpoints (#2005) and cross-model preservation (#2257/#2265) are unaffected.

Tests:
- packages/catalog/test/anthropic-copilot-signing-compat.test.ts: copilot
  (incl. enterprise copilot-api.* hosts) -> replayUnsignedThinking false;
  generic 3p reasoning -> true; official -> false. Fails before / passes after.
- packages/ai/test/anthropic-copilot-checkpoint-thinking-signature.test.ts:
  a signing copilot model never emits an empty-signature thinking block for a
  historical checkpoint turn (demotes to text, keeps tool_use), and still
  replays a clean signed historical thinking block natively.
2026-06-18 22:52:29 +02:00