- Filtered recall, fact, vector, temporal, and polyphonic voices to session-owned or explicitly global memories.
- Implemented graph_query and graph_link MCP tool handlers via EpisodicGraph; wired annotations and graph into Mnemosyne's external-db path.
- Fixed restore to stage and integrity-check before replacing the live database, rolling back on failure.
- Deduped generateId with a per-process nonce to prevent batch duplicate-content collisions.
- Added `disablesParallelToolUse` predicate scoped to exactly Opus 4.8.
- Injected `disable_parallel_tool_use: true` into outgoing `tool_choice` when tools are present; synthesizes an `auto` choice if none was set.
- Leaves `tool_choice: none` untouched to avoid Anthropic API rejection.
Mapped GitHub Copilot OAuth discovery context windows from max_context_window_tokens before prompt budgets.
Kept output token mapping on max_output_tokens and covered the fallback order in Copilot model limit tests.
Fixes#1539
Kept valid signed thinking blocks intact on the latest abandoned Anthropic tool-use assistant turn while still downgrading historical abandoned turns and aborted/error turns.
Fixes#1531
- Added `wrapFetchForRequestDebug` to intercept fetch calls and write `rr-session-N.json` (request) and `rr-session-N.res.log` (response headers + raw body) when `PI_REQ_DEBUG=1`.
- Integrated debug fetch wrapping into `stream` and `streamSimple` entry points so all providers inherit recording automatically.
- Extended Cursor HTTP/2 and Codex WebSocket transports with explicit debug session hooks for non-fetch protocols.
- Added `fetch` field propagation through `mapOptionsForApi` and Bedrock options so the wrapped fetch reaches every provider.
- Prefixed model cache fingerprints with a merge-v2 marker so cache keys are regenerated after model fingerprinting changes.
- Normalized OpenCode base-path handling and mapped dynamic OpenCode models to descriptor metadata so discovered models preserve their intended `api` and `baseUrl`.
- Added a regression test for issue #887 that mocks `/v1/models` and verifies `qwen3.7-max` remains `anthropic-messages` on refresh.
- Updated Kimi completion detection to treat z.ai binary thinking only for native Moonshot/Kimi-code hosts.
- Removed the generic Kimi model-id fallback from zai selection so OpenAI-compatible proxies default to openai reasoning format.
- Added regression coverage for native Kimi, proxy Kimi hosts, and OpenRouter precedence in thinkingFormat detection.
- Removed Ghostty-specific hardware-cursor forcing from TUI preference resolution and dropped the redundant terminal-cursor marker flag.
- Updated interactive mode editors to use `ui.getShowHardwareCursor()` so cursor mode now follows actual hardware-cursor visibility.
- Reworked terminal regressions tests to assert Ghostty respects the requested cursor preference and only emits cursor-show output when enabled.
- Fixed agent loop abandoning tool_use blocks on `stop`/`end_turn` turns; only `length` (truncation) now skips execution.
- Verified against live Anthropic API: stop_reason is never replayed on wire and doesn't gate continuation validity.
- Added tests pinning the wire-safety contract: thinking blocks with stale/missing signatures are downgraded to text before replay.
- Added prompt_cache_hit_tokens / prompt_cache_miss_tokens parsing in
parseChunkUsage for DeepSeek's prompt cache format where
prompt_tokens = hit_tokens + miss_tokens.
- DeepSeek formula: input = prompt_tokens - hit_tokens (= miss, billed input),
total = input + output + hit_tokens (avoid double-counting miss in cacheWrite).
- Added cache_hit status line segment showing cache hit rate:
rate = cacheRead / (cacheRead + cacheWrite) x 100%.
- Added cache_hit to all status line presets.
The :tools model suffix engages NanoGPT's server-side tool-call parser,
which 502s with code malformed_tool_call on complex DeepSeek payloads
(observed reliably on todo_write). The default route forwards
delta.content (including DSML envelope leaks) which our StreamMarkupHealing
already heals into a structured tool call.
The DSML allowlist still covers nanogpt and the parallel-index fix
remains; only the :tools suffix is removed.
Fixes#1488
- Updated the agent loop to execute tool calls only when the assistant stop reason was `toolUse`.
- Added skipped placeholder `tool_result` messages for leftover `toolCall` blocks when a turn ended without `toolUse`.
- Promoted OpenAI/Ollama `stop` tool-call turns to `toolUse` and stripped thinking signatures on abandoned tool-use turns during message transforms.
- Added a `historyRebuild` render intent to clear viewport and scrollback and emit a full repaint when geometry change invalidated terminal history.
- Updated render planning so width and height changes now rebuild native history immediately, while non-size content-only shrink updates only repaint the viewport to avoid yanking users in existing scrollback.
Tracked OpenAI-compatible streaming tool calls by provider index so parallel NanoGPT read calls keep their argument deltas attached to the matching tool-call block instead of falling through to the most recently opened call.
Added a NanoGPT DeepSeek regression covering two parallel read calls whose argument chunks arrive by index after the start chunk.
Fixes#1488
NanoGPT-hosted DeepSeek models (e.g. `nanogpt/deepseek/deepseek-v4-pro`
with reasoning enabled) emit `<|DSML|tool_calls>` envelopes inside
`delta.content` rather than the structured `tool_calls` array. The
healing pass that turns those leaks into real tool calls only engaged
when `modelMayLeakDsmlToolCalls` saw a known DeepSeek-hosting provider
in its allowlist; NanoGPT was missing, so the markup fell through as
visible text and the turn failed with malformed tool-call data.
Add `nanogpt` to the allowlist so the existing dsml grammar parses
the envelope into a structured tool call, mirroring `deepseek`,
`ollama`, `openrouter`, etc. Covered by a new repro test that
streams the reporter's verbatim DSML leak through the NanoGPT model.
Fixes#1488
Updated zhipu-coding-plan discovery and credential validation to use the dedicated Coding Plan API base URL instead of the general BigModel endpoint.
Added regression coverage for the default discovery URL.
Fixes#1494
- Added AnthropicCompat.supportsMidConversationSystem and AnthropicMessageParam system-role support for per-model control.
- Added supportsMidConversationSystemMessages(modelId) to return false on unknown models and true only for opus >=4.8.
- Updated convertAnthropicMessages/buildParams to map eligible developer turns to system role and keep unsupported turns as user.
- Added tests and fixtures for mid-conversation handling and documented `claude-opus-4-8` Bedrock metadata in changelog.
- Mocked the Vertex stream E2E test to override the home directory and clear GOOGLE_APPLICATION_CREDENTIALS so token resolution uses metadata credentials instead of local ADC files.
- Updated wafer and model-registry test expectations to match current model metadata values (Qwen3.7 Max and claude-opus-4-8).
- Added Opus 4.8 metadata and expanded provider entries, including Bedrock and other model variants.
- Updated `models.json` with additional Claude, Grok, Gemini, and Qwen entries carrying reasoning or image support.
- Adjusted Anthropic effort mapping for Opus 4.7+ to a five-tier scale and cached `requireSupportedEffort` results.
- Adjusted xhigh handling so legacy models now map to max and Opus 4.7+ map to low/med/high/xhigh/max.
- Updated thinking and alignment tests to expect revised xhigh/max mappings for opus47 and legacy models.
@yofriadi flagged that DeepSeek V4, GLM-5.x, Qwen3.x, MiMo, and MiniMax
on opencode-go front the same Zen gateway as Kimi (https://opencode.ai/go).
The thinking-state invariant — reasoning_content required on assistant
tool-call history when thinking is enabled, rejected when off — is
gateway-wide, not Kimi-specific.
Drop the `isKimiModelId` gate from the buildParams override so every
opencode request in thinking mode flips
requiresReasoningContentForToolCalls=true (with
allowsSyntheticReasoningContentForToolCalls=false and
reasoningContentField="reasoning_content" pinned the same way). The
existing #1071 guard for thinking-off and the
disableReasoningOnForcedToolChoice guard remain intact, so non-thinking
and forced-tool turns never inject the field.
Added parametrized regression tests covering GLM-5.1, Qwen3.7 Max, and
MiMo-V2 Pro through `streamOpenAICompletions` + `onPayload` — three
distinct thinkingFormat paths ("openai", "qwen", "openai") — to
prove the override fires regardless of model family. Expanded the
CHANGELOG entry to list every catalog model now covered.
@yofriadi reports the same OpenCode Zen 'reasoning_content is missing'
400 on `opencode-go/deepseek-v4-flash` and `deepseek-v4-pro`. The
gateway is the same Zen instance and the wire-body invariant matches:
the streamed-signature path in convertMessages previously emitted both
`reasoning` and `reasoning_content` on DeepSeek turns, which the
strict schema flags exactly like the Kimi case.
The line-1488 fix from 4215228b8 already coerces the replay onto the
configured `reasoningContentField` whenever
`allowsSyntheticReasoningContentForToolCalls=false` — DeepSeek family
sets that flag — so deepseek-v4 payloads now carry only
`reasoning_content`. Add a streamOpenAICompletions + onPayload
regression test that pins the shape.
The opencode-kimi thinking-mode override flipped
`requiresReasoningContentForToolCalls` on but left the rest of compat
at its default: `allowsSyntheticReasoningContentForToolCalls=true` and
the streamed-signature path in `convertMessages` echoed whichever
recognized field the upstream emitted. Opencode Kimi streams reasoning
under `reasoning`, so a follow-up replay landed `reasoning` on the
assistant message and left `reasoning_content` empty — the gateway
still 400s with 'reasoning_content is missing in assistant tool call
message at index N'.
Force `reasoning_content` as the wire field for this override:
- `buildParams` now also sets
`allowsSyntheticReasoningContentForToolCalls=false` and
`reasoningContentField="reasoning_content"` on the same gated
branch (kimi + opencode + thinking-on + not forced-tool).
- `convertMessages` thinking-block branch now respects
`allowsSyntheticReasoningContentForToolCalls`: when false, replay
always uses the configured `reasoningContentField` instead of the
streamed signature, so we never simultaneously write to both
`reasoning` and `reasoning_content`. DeepSeek already runs through
the same code with `allowsSynthetic=false` and existing tests
continue to pass under the cleaner output.
Updated the #1484 regression test to use the upstream's actual
`thinkingSignature: "reasoning"` shape and to additionally assert
that `reasoning` is absent from the wire body so a future regression
to dual-key emission would fail.
When OpenCode Kimi receives a forced `tool_choice`, the existing
`disableReasoningOnForcedToolChoice` guard at the bottom of
`buildParams` strips `reasoning_effort` (and sets `thinking: disabled`
for zai-format paths) from the wire body so Kimi does not 400 with
'tool_choice specified is incompatible with thinking enabled'. The
per-request reasoning_content override added for #1484 ran earlier in
`buildParams` and ignored that suppression, so the resulting payload
combined a thinking-disabled request with replayed reasoning_content on
prior assistant tool-call turns — reviving the #1071 'Extra inputs are
not permitted' failure.
Mirror the suppression in the override: skip flipping
`requiresReasoningContentForToolCalls` whenever
`compat.disableReasoningOnForcedToolChoice` would erase thinking on
this turn. New regression test exercises the forced-tool path through
`streamOpenAICompletions` + `onPayload` and asserts both the absence
of `reasoning_content` on the assistant message and the absence of
`reasoning_effort` on the request body.
OpenCode Zen's Kimi gateway gates reasoning_content on the request's
thinking state: it 400s with 'Extra inputs are not permitted' when
thinking is off but the field is supplied (#1071), and 400s with
'thinking is enabled but reasoning_content is missing in assistant tool
call message at index N' (#1484) when thinking is on and the field is
absent. Static compat detection from #1071 kept the field off
unconditionally, so reasoning-mode requests now 400 on the follow-up
turn after any tool call.
Override compat.requiresReasoningContentForToolCalls in buildParams when
the request itself is in thinking mode (options.reasoning set,
!options.disableReasoning, model.reasoning true), so prior tool-call
turns replay reasoning_content while thinking-disabled requests keep
the #1071 guard intact. Regression tests exercise both states through
streamOpenAICompletions + onPayload to assert on the wire body.
Fixes#1484
isMoonshotKimi only matched direct Moonshot/Kimi-code providers,
leaving kimi-* models routed through OpenCode Go, Kilo, and other
proxies with thinkingFormat: 'openai' → reasoning_effort instead
of the zai binary format the Kimi backend expects.
OpenRouter's normalized reasoning path (thinkingFormat: 'openrouter')
takes precedence over the generic Kimi model-id match so bundled
openrouter/xiaomi/kimi-* models keep their existing API shape.
isMoonshotKimi only matched direct Moonshot/Kimi-code providers,
leaving kimi-* models routed through OpenCode Go, OpenRouter, Kilo,
and other proxies with thinkingFormat: 'openai' → reasoning_effort
instead of the zai binary format the Kimi backend expects.
Extend thinkingFormat detection to isKimiModel so the zai format
applies regardless of which provider routes the model. Compat
blocks already set by individual providers (kimi-code mapModel,
descriptor transformModel) take precedence via resolveOpenAICompat
overlay, so existing bundled entries are unchanged.
The models.dev upstream returns Xiaomi models; the descriptor was set to
anthropicMessagesDescriptor causing every `bun run generate-models` to
regenerate models.json with api=anthropic-messages and baseUrl=/anthropic,
overwriting the intentional switch to OpenAI Chat Completions API.
Changed to openAiCompletionsDescriptor("/v1") so Xiaomi models consistently
use the OpenAI-compatible endpoint that supports reasoning_effort, avoids
the unsigned thinking-signature leak problem, and matches how Kilo,
OpenRouter, NanoGPT, and ZenMux route Xiaomi traffic.
Constraint: models.dev upstream data cannot be changed externally.
Rejected: Patching models.json manually after each generate-models run | Would break on every models bump.
Confidence: high
Scope-risk: narrow
Directive: If models.dev ever adds Xiaomi models with different IDs, verify
the openAiCompletionsDescriptor mapping still covers them.
Tested: `bun run generate-models` produces xiaomi/* models with api=openai-completions
Not-tested: Actual API call against Xiaomi endpoint