- Removed the pi dialect implementation and associated source files.
- Updated dialect resolution, factory registration, and type definitions to exclude pi.
- Cleaned up settings schema and user options to remove pi-related configurations.
- Deleted corresponding test suites covering pi dialect functionality, in-band tools, and examples.
- Update `umans-provider` test to remove references to deprecated GLM 5.1 model.
- Rename search tool reference to `grep` in `advisor` test.
- Improve test stability in TUI components by explicitly draining `setImmediate` queues before flushing terminal state.
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
Downgraded named forced chat-completions tool_choice to required for llama.cpp so string-only parsers do not ignore the directive.
Added a compat flag and regression coverage for the capture-at-stop resolve path.
Fixes#3593
The Qwen3 / Qwen3.6 chat template strips <think>...</think> from every
assistant turn whose loop.index0 <= ns.last_query_index, so the moment a
new user message (the user's real next prompt OR the auto-learn
capture-at-stop nudge) lands, every prior assistant turn becomes 'older'
and is re-rendered without its <think> block — diverging from the
generation tokens still in the local slot's KV cache and forcing full
prompt re-processing on SWA models.
Sending reasoning_content alone (the #3528 fix) does not help: the
template's older branch renders only `content`, never the
reasoning_content field. The official Qwen3.6 fix is
`preserve_thinking: true`, which makes the template render
<think>\n{reasoning_content}\n</think>\n\n{content} for every assistant
turn regardless of position.
- packages/catalog/src/compat/openai.ts: new qwenPreserveThinking shared
compat flag, auto-enabled when the resolved thinkingFormat is `qwen`
or `qwen-chat-template` AND replayReasoningContent is on (the four
local provider ids plus loopback / RFC1918 / *.local baseUrls).
Responses-API builder pins it false — it's a chat-template knob,
irrelevant on the Responses surface.
- packages/ai/src/providers/openai-shared.ts: chat-completions encoder
emits preserve_thinking: true alongside enable_thinking: true in both
Qwen disable-mode branches (twin top-level + chat_template_kwargs
emission so llama.cpp / vLLM / SGLang and Alibaba's compatible-mode
wire shapes all pick it up). Stays off when thinking is disabled.
- packages/ai/test/issue-3528-repro.test.ts: nine new pins covering the
auto-detection matrix, the wire emission on local Qwen + thinking, the
cloud-Qwen / reasoning-disabled negative cases, and the
explicit-override escape hatches in both directions.
- Backfilled qwenPreserveThinking: false on three hand-rolled
ResolvedOpenAICompat fixtures so the required field stays satisfied.
- AI + catalog changelog entries under ## [Unreleased].
Verification:
bun --cwd packages/ai test ./test/issue-3528-repro.test.ts ./test/openai-completions-compat.test.ts ./test/openai-completions-tool-result-images.test.ts ./test/issue-967-vision-guard.test.ts → 87 pass
bun --cwd packages/ai test ./test/issue-3434-repro.test.ts ./test/deepseek-reasoning-content.test.ts ./test/ollama-thinking-disable.test.ts ./test/openai-compat-policy.test.ts → 37 pass
bun --cwd packages/catalog test → 325/325 pass
bun --cwd packages/ai check:types && bun --cwd packages/catalog check:types → clean
Fixes#3541
The discovery paths for llama.cpp, LM Studio, and openai-models-list
hardcode reasoning: false because the upstream /models endpoints don't
advertise the capability. The original gate
Boolean(spec.reasoning) && (LOCAL_PROVIDER || loopback)
therefore left replayReasoningContent off for the common setup the bug
targets — a discovered Qwen / DeepSeek model on local llama.cpp — and
the stream parser still recorded the upstream's reasoning_content
deltas as thinking blocks, so #3528 reproduced unchanged.
Drop the spec.reasoning gate. The encoder's own
'if (nonEmptyThinkingBlocks.length > 0)' guard ensures the flag stays a
no-op for pure-text turns, so always-on for local hosts adds nothing to
non-reasoning histories. Flip the matching test and add an encoder pin
that mirrors the discovery setup (reasoning: false + actual thinking
block must still ride as reasoning_content).
Addresses chatgpt-codex review on #3532.
LiteLLM defaults to http://localhost:4000/v1 and is the only built-in
`openai-completions` provider that forwards to an unrelated upstream
(OpenAI, Anthropic, …) rather than running a chat-template renderer
itself. The loopback auto-detection from the original #3528 fix would
push `reasoning_content` to those upstreams, which gain no KV-cache
benefit and may 400 on the extra field.
Add a `PROXY_OPENAI_COMPAT_PROVIDERS` deny set (currently just
`litellm`) that excludes proxy ids from both the provider allow-list
and the loopback heuristic. Users running a custom proxy in front of a
llama.cpp-style backend can still opt in via
`compat.replayReasoningContent: true`.
Addresses chatgpt-codex review on #3532.
Local llama.cpp / LM Studio / vLLM / Ollama (openai-completions mode)
and custom providers on loopback/RFC1918 baseUrls re-tokenize the entire
chat-template prompt every request. Qwen3 / DeepSeek-R1 / GLM templates
reconstruct the prior assistant turn's `<think>…</think>` block from
`reasoning_content`; dropping the field re-renders the assistant turn
without thinking content, the rendered tokens diverge from the slot's
existing KV cache, and llama.cpp falls back to full prompt re-processing.
The auto-learn capture-at-stop nudge (#3504/#3505) made this reproduce
on every turn for thinking-enabled local models: the reporter's wire
captures show `cached_tokens` collapsing from 38650 (req11) to 0
(req12) the moment the assistant reply re-enters history as a context
message.
Add `OpenAICompat.replayReasoningContent`, auto-enabled in
`buildOpenAICompat` whenever `spec.reasoning` is set AND the model is on
a known local provider id or a loopback / RFC1918 / `*.local` host. The
`openai-completions` encoder gets a fourth thinking-block branch that
emits `reasoning_content` on every reasoning-engaged assistant turn
(not just tool-call turns), honoring the streamed signature when it
identifies a recognized wire field and falling back to the configured
`reasoningContentField` otherwise. `transformMessages` learns the new
flag so cross-API replays into local llama.cpp targets preserve
unsigned thinking blocks.
Fixes#3528
Direct Anthropic Claude Sonnet 4.5 and Haiku 4.5 (plus their Cloudflare,
Vertex, GitLab-Duo, Copilot, OpenCode-Zen, and Bedrock cross-region passes)
were classified as anthropic-budget-effort, which made the Anthropic provider
serialize output_config.effort alongside the thinking.budget_tokens block.
Anthropic only honors output_config.effort on Opus 4.5 and adaptive (4.6+)
Messages-API models — Sonnet 4.5 and Haiku 4.5 reject every request with
HTTP 400 'This model does not support the effort parameter.', so the advisor
(and any agent on a Sonnet/Haiku 4.5 SKU) failed every turn.
inferThinkingControlMode now gates anthropic-budget-effort to
parsedModel.kind === 'opus' && semverGte(version, '4.5') on both
anthropic-messages and bedrock-converse-stream. Sonnet/Haiku 4.5 fall through
to mode: 'budget' (effort still scales the per-tier thinking budget via
ANTHROPIC_THINKING[reasoning]); Opus 4.5 keeps anthropic-budget-effort and
continues to emit output_config.effort. anthropic-budget-effort also remains
in use for Anthropic-compatible third-party backends that natively support
the field (Umans GLM 5.2).
Regenerated models.json so the 19 Sonnet/Haiku 4.5 first-party entries flip
to mode: 'budget' and the 12 Opus 4.5 entries stay on anthropic-budget-effort.
Regression tests cover both sides: anthropic-alignment.test.ts asserts the
Sonnet 4.5 wire body omits output_config and the Opus 4.5 wire body emits
output_config.effort: 'medium'.
Fixes#3497
Ollama Cloud rejects any chat request whose options.num_predict exceeds
65536 with HTTP 400, but the existing safety relied on the load-time
omitMaxOutputTokens policy in model-registry.ts. Stale models.db rows
predating that policy (and custom modelOverrides re-enabling output
caps) carried maxTokens: 1048576 forward to the wire layer untouched,
so every request to deepseek-v4-pro / deepseek-v4-flash 400'd with:
max_tokens (1048576) exceeds model's maximum output tokens (65536)
for model deepseek-v4-pro
createChatBody now resolves num_predict through resolveNumPredict,
which clamps every ollama-cloud request at the documented cap before
serialization — independent of the cached spec, on top of the existing
omitMaxOutputTokens path. Self-hosted ollama traffic is unaffected.
Fixes#3392
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.
- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
forces input=['text'] whenever the resolved baseUrl is not the
canonical personal-Copilot host, so the upstream's vision flag is
honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
OR-upgrading with the bundled reference) when the merged baseUrl
differs from the bundled one, so a bundled spec pinned to the
personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
rasterizer for any github-copilot model whose baseUrl is non-personal,
catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
pi-catalog/wire/github-copilot so catalog and coding-agent share one
canonical check.
Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).
Fixes#3387