Commit Graph
293 Commits
Author SHA1 Message Date
can1357 24341e4abf chore: bump version to 16.2.2 2026-06-27 13:24:46 +02:00
can1357 1e2817d37d Merge remote-tracking branch 'origin/farm/329514e4/ttsr-hashline-edit-path-extraction' 2026-06-27 13:24:24 +02:00
can1357 053da98ddc feat: removed pi dialect and its associated infrastructure
- Removed the pi dialect implementation and associated source files.
- Updated dialect resolution, factory registration, and type definitions to exclude pi.
- Cleaned up settings schema and user options to remove pi-related configurations.
- Deleted corresponding test suites covering pi dialect functionality, in-band tools, and examples.
2026-06-27 08:37:25 +02:00
can1357 d738329485 chore: bump version to 16.2.1 2026-06-27 06:35:42 +02:00
can1357 7463818a32 chore: bump version to 16.2.0 2026-06-27 04:39:59 +02:00
can1357 26b3a22186 test: updated test suites and fix rendering test flakiness
- Update `umans-provider` test to remove references to deprecated GLM 5.1 model.
- Rename search tool reference to `grep` in `advisor` test.
- Improve test stability in TUI components by explicitly draining `setImmediate` queues before flushing terminal state.
2026-06-27 04:28:39 +02:00
can1357 5a044dc0da feat: consolidated and automate changelog management
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
2026-06-27 03:08:52 +02:00
can1357 f951fe3ca1 chore: bump models 2026-06-27 02:16:38 +02:00
can1357 667319adb4 fix(types): add os import, drop invalid id fields in tests after #3416/#3510 drop 2026-06-27 02:10:41 +02:00
can1357 14cc9cba0f Merge PR #3106: fix(compaction): enable custom provider remote compaction (@roboomp) 2026-06-27 01:39:34 +02:00
can1357 f6e7d8ebc7 Merge PR #3574: feat(coding-agent): discover rich LiteLLM proxy metadata (@jdavv) 2026-06-27 01:39:32 +02:00
can1357 3e5360ca4c Merge PR #3060: feat(provider): add GitLab Duo Agent provider (@jiwangyihao) 2026-06-27 01:39:31 +02:00
roboomp cfb3b60685 fix(ai): sent llama.cpp string tool choice
Downgraded named forced chat-completions tool_choice to required for llama.cpp so string-only parsers do not ignore the directive.

Added a compat flag and regression coverage for the capture-at-stop resolve path.

Fixes #3593
2026-06-26 18:03:27 +00:00
can1357 011f036ef5 chore: bump version to 16.1.23 2026-06-26 18:29:01 +02:00
Jean-Luc Davern 3dc9099058 feat(coding-agent): discover rich LiteLLM metadata 2026-06-26 11:26:27 -05:00
can1357 d7f05ac640 fix(catalog): scoped CoreWeave catalog metadata 2026-06-26 17:12:01 +02:00
Lance Tuller efdcadf0d5 feat(ai): add CoreWeave Serverless Inference provider 2026-06-26 09:13:41 -04:00
roboomp ac958d5905 fix(ai): emit Qwen preserve_thinking so local-server cache survives new user messages
The Qwen3 / Qwen3.6 chat template strips <think>...</think> from every
assistant turn whose loop.index0 <= ns.last_query_index, so the moment a
new user message (the user's real next prompt OR the auto-learn
capture-at-stop nudge) lands, every prior assistant turn becomes 'older'
and is re-rendered without its <think> block — diverging from the
generation tokens still in the local slot's KV cache and forcing full
prompt re-processing on SWA models.

Sending reasoning_content alone (the #3528 fix) does not help: the
template's older branch renders only `content`, never the
reasoning_content field. The official Qwen3.6 fix is
`preserve_thinking: true`, which makes the template render
<think>\n{reasoning_content}\n</think>\n\n{content} for every assistant
turn regardless of position.

- packages/catalog/src/compat/openai.ts: new qwenPreserveThinking shared
  compat flag, auto-enabled when the resolved thinkingFormat is `qwen`
  or `qwen-chat-template` AND replayReasoningContent is on (the four
  local provider ids plus loopback / RFC1918 / *.local baseUrls).
  Responses-API builder pins it false — it's a chat-template knob,
  irrelevant on the Responses surface.
- packages/ai/src/providers/openai-shared.ts: chat-completions encoder
  emits preserve_thinking: true alongside enable_thinking: true in both
  Qwen disable-mode branches (twin top-level + chat_template_kwargs
  emission so llama.cpp / vLLM / SGLang and Alibaba's compatible-mode
  wire shapes all pick it up). Stays off when thinking is disabled.
- packages/ai/test/issue-3528-repro.test.ts: nine new pins covering the
  auto-detection matrix, the wire emission on local Qwen + thinking, the
  cloud-Qwen / reasoning-disabled negative cases, and the
  explicit-override escape hatches in both directions.
- Backfilled qwenPreserveThinking: false on three hand-rolled
  ResolvedOpenAICompat fixtures so the required field stays satisfied.
- AI + catalog changelog entries under ## [Unreleased].

Verification:
  bun --cwd packages/ai test ./test/issue-3528-repro.test.ts ./test/openai-completions-compat.test.ts ./test/openai-completions-tool-result-images.test.ts ./test/issue-967-vision-guard.test.ts → 87 pass
  bun --cwd packages/ai test ./test/issue-3434-repro.test.ts ./test/deepseek-reasoning-content.test.ts ./test/ollama-thinking-disable.test.ts ./test/openai-compat-policy.test.ts → 37 pass
  bun --cwd packages/catalog test → 325/325 pass
  bun --cwd packages/ai check:types && bun --cwd packages/catalog check:types → clean

Fixes #3541
2026-06-26 08:44:28 +00:00
jiwangyihao 6d280d4349 fix(gitlab-duo): 命名空间发现翻页遍历顶层 group 并放宽 SSH 远程端口比较
处理两条 Codex 审阅意见:(1) fetchTopLevelGroupNamespaceCandidates 只取第一页,token 属于 >100 个顶层 group 时后续页的可用 Duo namespace 不可见,现在跟随 GitLab x-next-page 分页(受 GITLAB_DUO_WORKFLOW_MAX_GROUP_PAGES 限制)遍历所有页再校验候选;(2) 自管 GitLab 常把 SSH 暴露在独立端口(ssh://git@host:2222/...)而 web 是 https://host,原 host:port 严格比较会误拒该远程,现在 SSH/scp 远程按裸 hostname 比较,HTTP(S) 远程仍严格比 host:port 以区分同主机不同服务。各加 1 个发现测试。
2026-06-26 16:13:27 +08:00
jiwangyihao bf1b692df9 fix(gitlab-duo): direct_access 错误保留 HTTP 状态码以触发凭据轮换
机器人指出两处问题,本提交处理:

1) direct_access 带 body message 的错误丢弃了 response.status,导致
   streaming auth-retry 路径(extractStatusFromAssistantError ->
   extractHttpStatusFromError)无法恢复状态、无法刷新/轮换过期 OAuth 或
   配额受限的 broker 凭据。现在即便有 body message 也内嵌 HTTP <status>,
   与 create 路径既有约定一致。新增 401 Unauthorized 回归测试断言状态可恢复,
   并强化既有 403 配额测试。

2) generate-models 种子注释过度暗示生成期会跑 namespace-scoped 发现。
   实际上 gitlab-duo-agent descriptor 故意不带 catalogDiscovery,被
   isCatalogDescriptor 过滤排除在生成发现循环之外,因此生成期绝不会拉取
   某账号的 aiChatAvailableModels,只播种通用 namespace-free fallback。
   修正注释表述,新增两条针对 descriptor 的回归测试:断言 descriptor 无
   catalogDiscovery(永不参与生成发现),以及 fallback 模型不带
   gitlabDuoWorkflowRootNamespaceId。
2026-06-26 16:13:24 +08:00
jiwangyihao 0e78a24443 fix(catalog): 将 GitLab Duo Agent fallback 模型纳入 bundled models.json
机器人指出 gitlab-duo-agent 不在 models.json,fresh 安装(尚无带凭据的动态
发现/缓存)时内置 catalog 看不到默认模型。generator 现按 Sakana 同样方式播种
gitlab-duo-agent 的 fallback 模型(claude_sonnet_4_6_vertex):live aiChatAvailableModels
发现成功时其条目按 id 去重胜出,仅在无凭据/失败 regen 时落种子。

- generate-models.ts:导入 buildGitLabDuoWorkflowFallbackModel,在 authoritative
  发现未命中时 push 种子。
- 重新生成 models.json,仅保留 gitlab-duo-agent provider 块的新增,其余 provider
  数据维持基线不变(regen 在无凭据环境下未触碰其它 provider)。
- 新增针对 descriptor 的回归测试(非 bundled JSON):断言 manager options 暴露
  fallback 静态模型,符合 AGENTS.md 要求。
2026-06-26 16:13:24 +08:00
jiwangyihao 20e4c3bd5b fix(ai): 处理 Codex 对项目自动发现、设置启用与 changelog 归属的三项审阅意见
- 自动项目发现优先复用命名空间解析所依据的项目(工作区 git remote 或显式
  项目),而不是从 group 列表泛选;多项目 group 下不再把工作流 scope 到无关项目。
  命名空间选择沿用 remote/project 来源的 projectPath,运行时据此 scope。
- Duo 设置启用仅在确定性尝试(任意 HTTP 响应,含 4xx)后才标记 ensured;
  瞬时网络错误/5xx 返回 false,使后续 turn 可重试,避免 fresh 命名空间因一次
  瞬时失败而永久跳过 PUT。
- agent CHANGELOG 把 cwd 转发条目从已发布 [16.1.8] 段移至 [Unreleased];
  移动后 [16.1.8] 与上游 main 一致。
2026-06-26 16:13:24 +08:00
jiwangyihao 3298c35eec fix(catalog): 处理 Codex 对 Duo 模型缓存与远程端口比较的两项审阅意见
- 动态模型缓存键改为镜像 discoverGitLabDuoWorkflowNamespace 真实解析输入:
  内置发现只传 apiKey/baseUrl/fetch,原键里的 namespaceId/projectId/cwd 恒为
  空,退化为 apiKey+baseUrl,导致同 token 下两个不同 group 的工作区互相复用
  权威模型缓存。改为按 (凭据, baseUrl, 命名空间/项目 config 或同名环境变量,
  有效 cwd) 指纹分区。
- 远程 host 比较改用 host(含端口)而非 hostname:同主机不同端口的自托管 GitLab
  不再被误判为同实例,避免从错误远程推导项目路径。SCP 远程无端口概念,仍按裸
  host 比较。新增跨端口远程不被当作工作区项目的回归测试。
2026-06-26 16:13:23 +08:00
jiwangyihao 4fdf2f83a5 fix(ai): 处理 rebase 后 Codex 新增的三项审阅意见
- catalog CHANGELOG 删除 rebase 重放进已发布 [16.1.4] 段落的重复 Claude 4.6 条目,使该段落与上游 main 完全一致(已发布段不可变)
- Duo Agent finally 清理在最终 idle timeout(重试已耗尽)时也发送 stop PATCH,避免代理/LB 持续断连场景下服务端工作流残留
- auth-broker login 仅对 pasteCodeFlow provider 传入 onManualCodeInput,普通 loopback provider 不再让 readline 提示与 HTTP 回调竞争导致终端残留
2026-06-26 16:13:23 +08:00
jiwangyihao b2403bd5f9 fix(ai): address Codex review for Duo Agent provider
- 修复 special.ts 重复 return 片段导致 catalog 无法解析
- stop/settings-enablement 改用 gitLabApiUrl 保留自托管相对路径
- 重放命中 action 早返回前清除 paused,避免续帧被缓冲到超时
- setupForNamespace 返回值改用具名 GitLabDuoWorkflowNamespaceSetup(去除 ReturnType)
- toolChoice=none 的 side-request 不再向 Duo 暴露工具
- 套接字异常关闭(closed)也走 stop 清理,避免服务端工作流残留
- ProviderSessionState.close 关闭时发送 stop,避免会话销毁后残留工作流
- 命名空间/设置启用缓存按 (account, baseUrl, cwd) 分区
- 动态模型缓存按凭据+命名空间作用域分区
- resume 失败时丢弃 active 并 stop 工作流
2026-06-26 16:13:22 +08:00
jiwangyihao 65108346e9 feat(catalog): add GitLab Duo Agent model discovery 2026-06-26 16:13:20 +08:00
can1357 ff759ed850 chore: bump version to 16.1.22 2026-06-26 09:19:53 +02:00
roboomp b6d81f22e1 fix(catalog): replayReasoningContent must not gate on spec.reasoning
The discovery paths for llama.cpp, LM Studio, and openai-models-list
hardcode reasoning: false because the upstream /models endpoints don't
advertise the capability. The original gate

    Boolean(spec.reasoning) && (LOCAL_PROVIDER || loopback)

therefore left replayReasoningContent off for the common setup the bug
targets — a discovered Qwen / DeepSeek model on local llama.cpp — and
the stream parser still recorded the upstream's reasoning_content
deltas as thinking blocks, so #3528 reproduced unchanged.

Drop the spec.reasoning gate. The encoder's own
'if (nonEmptyThinkingBlocks.length > 0)' guard ensures the flag stays a
no-op for pure-text turns, so always-on for local hosts adds nothing to
non-reasoning histories. Flip the matching test and add an encoder pin
that mirrors the discovery setup (reasoning: false + actual thinking
block must still ride as reasoning_content).

Addresses chatgpt-codex review on #3532.
2026-06-26 06:43:23 +00:00
roboomp e59466cca5 fix(catalog): exclude proxy providers from replayReasoningContent auto-detect
LiteLLM defaults to http://localhost:4000/v1 and is the only built-in
`openai-completions` provider that forwards to an unrelated upstream
(OpenAI, Anthropic, …) rather than running a chat-template renderer
itself. The loopback auto-detection from the original #3528 fix would
push `reasoning_content` to those upstreams, which gain no KV-cache
benefit and may 400 on the extra field.

Add a `PROXY_OPENAI_COMPAT_PROVIDERS` deny set (currently just
`litellm`) that excludes proxy ids from both the provider allow-list
and the loopback heuristic. Users running a custom proxy in front of a
llama.cpp-style backend can still opt in via
`compat.replayReasoningContent: true`.

Addresses chatgpt-codex review on #3532.
2026-06-26 06:35:00 +00:00
roboomp 4842f00297 style: bun run fix 2026-06-26 06:23:38 +00:00
roboomp 5e006e7f81 fix(ai): replay reasoning_content on assistant turns for local llama.cpp servers
Local llama.cpp / LM Studio / vLLM / Ollama (openai-completions mode)
and custom providers on loopback/RFC1918 baseUrls re-tokenize the entire
chat-template prompt every request. Qwen3 / DeepSeek-R1 / GLM templates
reconstruct the prior assistant turn's `<think>…</think>` block from
`reasoning_content`; dropping the field re-renders the assistant turn
without thinking content, the rendered tokens diverge from the slot's
existing KV cache, and llama.cpp falls back to full prompt re-processing.

The auto-learn capture-at-stop nudge (#3504/#3505) made this reproduce
on every turn for thinking-enabled local models: the reporter's wire
captures show `cached_tokens` collapsing from 38650 (req11) to 0
(req12) the moment the assistant reply re-enters history as a context
message.

Add `OpenAICompat.replayReasoningContent`, auto-enabled in
`buildOpenAICompat` whenever `spec.reasoning` is set AND the model is on
a known local provider id or a loopback / RFC1918 / `*.local` host. The
`openai-completions` encoder gets a fourth thinking-block branch that
emits `reasoning_content` on every reasoning-engaged assistant turn
(not just tool-call turns), honoring the streamed signature when it
identifies a recognized wire field and falling back to the configured
`reasoningContentField` otherwise. `transformMessages` learns the new
flag so cross-API replays into local llama.cpp targets preserve
unsigned thinking blocks.

Fixes #3528
2026-06-26 06:23:27 +00:00
can1357 018a9638c0 chore: bump version to 16.1.21 2026-06-26 07:39:22 +02:00
can1357 0fc6d136c3 chore: bump version to 16.1.20 2026-06-25 22:45:00 +02:00
can1357 3188fc6bb9 Merge remote-tracking branch 'origin/farm/371a49e5/anthropic-claude-4-5-budget-effort' 2026-06-25 22:35:22 +02:00
roboomp 7a4cf0e30c fix(catalog): classified direct Anthropic Sonnet/Haiku 4.5 as plain budget thinking
Direct Anthropic Claude Sonnet 4.5 and Haiku 4.5 (plus their Cloudflare,
Vertex, GitLab-Duo, Copilot, OpenCode-Zen, and Bedrock cross-region passes)
were classified as anthropic-budget-effort, which made the Anthropic provider
serialize output_config.effort alongside the thinking.budget_tokens block.
Anthropic only honors output_config.effort on Opus 4.5 and adaptive (4.6+)
Messages-API models — Sonnet 4.5 and Haiku 4.5 reject every request with
HTTP 400 'This model does not support the effort parameter.', so the advisor
(and any agent on a Sonnet/Haiku 4.5 SKU) failed every turn.

inferThinkingControlMode now gates anthropic-budget-effort to
parsedModel.kind === 'opus' && semverGte(version, '4.5') on both
anthropic-messages and bedrock-converse-stream. Sonnet/Haiku 4.5 fall through
to mode: 'budget' (effort still scales the per-tier thinking budget via
ANTHROPIC_THINKING[reasoning]); Opus 4.5 keeps anthropic-budget-effort and
continues to emit output_config.effort. anthropic-budget-effort also remains
in use for Anthropic-compatible third-party backends that natively support
the field (Umans GLM 5.2).

Regenerated models.json so the 19 Sonnet/Haiku 4.5 first-party entries flip
to mode: 'budget' and the 12 Opus 4.5 entries stay on anthropic-budget-effort.

Regression tests cover both sides: anthropic-alignment.test.ts asserts the
Sonnet 4.5 wire body omits output_config and the Opus 4.5 wire body emits
output_config.effort: 'medium'.

Fixes #3497
2026-06-25 20:16:34 +00:00
roboomp 80862b79da fix(agent): handled ollama-cloud task backoff
Added ollama-cloud subagent concurrency limiting, role fallback-chain inheritance, and visible empty length errors for native Ollama responses.

Fixes #3464
2026-06-25 11:57:51 +00:00
can1357 451af61280 chore: bump version to 16.1.19 2026-06-25 13:14:44 +02:00
can1357 b2b818af6f chore: bump models 2026-06-25 06:46:36 +02:00
can1357 1a89d784e7 chore: bump version to 16.1.18 2026-06-25 05:17:47 +02:00
can1357 8d60776a6f chore: bump version to 16.1.17 2026-06-24 21:11:31 +02:00
can1357 90c8f8251a Merge remote-tracking branch 'origin/farm/3d137ac8/openrouter-anthropic-reasoning-replay' 2026-06-24 21:07:21 +02:00
can1357 23838ba80b chore: update changelogs 2026-06-24 21:03:29 +02:00
can1357 0eb21efa1a Merge remote-tracking branch 'origin/farm/7ef98714/snapcompact-copilot-vision-gate' 2026-06-24 21:00:37 +02:00
can1357 4c4fb21caa Merge remote-tracking branch 'origin/farm/cf20a2c7/clamp-ollama-cloud-max-output-tokens' 2026-06-24 21:00:03 +02:00
roboomp ebdce456dc fix(providers): stripped openrouter anthropic reasoning replay
Fixes #3399
2026-06-24 18:57:33 +00:00
roboomp 0edb3d32b8 fix(ai/ollama): clamped num_predict at the 65536 Ollama Cloud cap
Ollama Cloud rejects any chat request whose options.num_predict exceeds
65536 with HTTP 400, but the existing safety relied on the load-time
omitMaxOutputTokens policy in model-registry.ts. Stale models.db rows
predating that policy (and custom modelOverrides re-enabling output
caps) carried maxTokens: 1048576 forward to the wire layer untouched,
so every request to deepseek-v4-pro / deepseek-v4-flash 400'd with:

  max_tokens (1048576) exceeds model's maximum output tokens (65536)
  for model deepseek-v4-pro

createChatBody now resolves num_predict through resolveNumPredict,
which clamps every ollama-cloud request at the documented cap before
serialization — independent of the cached spec, on top of the existing
omitMaxOutputTokens path. Self-hosted ollama traffic is unaffected.

Fixes #3392
2026-06-24 17:29:28 +00:00
can1357 a4dbe6396c style: biome organize-imports on merged test files (#3193, #3312) 2026-06-24 19:03:32 +02:00
can1357 ff2280e6ee chore(changelog): normalize [Unreleased] after merge sweep (fix-changelogs --since baseline; dedup #3258) 2026-06-24 18:46:00 +02:00
roboomp 714051d795 fix(catalog,coding-agent): disable vision on non-personal copilot endpoints
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.

- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
  forces input=['text'] whenever the resolved baseUrl is not the
  canonical personal-Copilot host, so the upstream's vision flag is
  honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
  OR-upgrading with the bundled reference) when the merged baseUrl
  differs from the bundled one, so a bundled spec pinned to the
  personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
  rasterizer for any github-copilot model whose baseUrl is non-personal,
  catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
  pi-catalog/wire/github-copilot so catalog and coding-agent share one
  canonical check.

Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).

Fixes #3387
2026-06-24 16:27:14 +00:00
can1357 7408ee75f2 Merge PR #3193: fix(catalog): restore Umans GLM-5.2 max reasoning (@roboomp)
# Conflicts:
#	packages/ai/test/anthropic-alignment.test.ts
#	packages/catalog/test/umans-provider.test.ts
2026-06-24 18:26:00 +02:00