- Updated the Gemini 3.1 Pro variant collapse so high effort now routes to `gemini-pro-agent` instead of `gemini-3.1-pro-high`.
- Kept `gemini-3.1-pro-high` in the collapse members list so the raw discovery ID remains handled.
- Lowered `contextWindow` from `163840` to `131072` tokens in `models.json`.
Individual conflict resolution now bypasses the LSP writethrough to prevent formatting from corrupting other unresolved marker blocks and to avoid noisy diagnostics in partially resolved files.
- Recomputed `tools.discoveryMode: "auto"` in the deferred MCP closure in `sdk.ts` once the real tool count is known: a toolset crossing the threshold now flips discovery on, registers and activates `search_tool_bm25`, and skips `activateAll` instead of force-activating every MCP tool.
- Guarded the deferred MCP task against disposed sessions: added `AgentSession.isDisposed` and `enableMCPDiscovery()`, and the late connect now calls `disconnectAll()` instead of refreshing tools onto a dead session.
- Cleared `#fastPathKey`/`#fastPathItems` in `AssistantMessageComponent.invalidate()` so theme/symbol changes rebuild reused Markdown children instead of keeping stale captured themes.
- Memoized unusable read summaries as a `false` sentinel in `read.ts` so the per-session LRU no longer retains full sources of unsummarizable files.
- Broadened `HAS_REF_DEF` in `markdown.ts` to match backslash-escaped reference labels (`[a\]b]: x`) and cleared frozen stream-lex state on blank `setText()`.
- Added regression tests: deferred auto-discovery flip and mid-connect dispose (`sdk-mcp-auto-discovery.test.ts` + `many-tools-mcp.ts` fixture), fast-path child rebuild on invalidate, and escaped-ref-def incremental-lex equivalence.
- Removed stateful SSE chain controls from Codex provider options.
- Changed SSE transport to send full request bodies without `previous_response_id` deltas.
- Updated stale-chain recovery to treat `unsupported` as stale and recover via websocket.
- Scoped `x-codex-turn-state` to in-turn follow-ups and cleared turn state otherwise.
- Adjusted Anthropic tool-result conversion to keep error results text-only while collecting any image blocks separately.
- Reattached collected images after the tool-result run in the same user message with a separator note for Anthropic compatibility.
- Added tests covering image-only error tool_results, image hoisting order, and successful tool_result image preservation.
Reviewer flagged that the bracketed-paste path is untrusted terminal input —
ANSI escapes, control chars, newlines/tabs, or a multi-hundred-char path
would corrupt the status line (per AGENTS.md TUI sanitization rules) and
leak the absolute home-dir path.
The new ENOENT diagnostic now feeds the path through sanitizeText (strip
ANSI/C0/C1 controls), collapses CR/LF/TAB to single spaces, runs it
through shortenPath (collapse home → '~'), and truncateToWidth-clamps it
to TRUNCATE_LENGTHS.CONTENT (80) before interpolating into either the SSH
or local status string. Added a third assertion to the repro test
defending the contract: ANSI/control bytes never reach the status, and
the displayed path is bounded below the input length.
Refs PR #2376
- Added optional `isError` markers on inline tool-result and message models.
- Skipped error-marked tool results during swap planning, savings estimation, and live transformations.
- Added tests confirming large error tool results stay text-only and are not considered for inline swaps.
When a local terminal forwards a bracketed-paste containing an image-file
path and the omp process is itself running on the remote end of an SSH
session, the path is on the user's local machine and unreachable from the
remote filesystem. The path-as-text fallback in handleImagePathPaste then
made it look like the image was attached when in fact only a useless
absolute path entered the editor and the bytes never crossed.
Distinguish ENOENT from other read failures: drop the misleading
path-as-text degrade, and surface a clear status that names the file and
— when SSH_CONNECTION/SSH_TTY/SSH_CLIENT indicate the remote end of an
SSH session — points the user at the clipboard image-paste shortcut so
the bytes actually cross. The ImageInputTooLargeError and unknown-error
branches keep their existing fallback so genuine type issues still
yield the user's input back to them.
Fixes#2375
Dynamically assigns segment colors by sweeping across the full HSV hue spectrum based on their track position. This ensures each segment has a distinct, predictable hue and removes the need for explicit `color` assignments.
- Updated the Windows npm shim resolver to resolve shim files against `cwd` and inspect `_prog` before treating a batch wrapper as a node launcher.
- Added a node-only guard so non-node wrappers, such as python shims, fall back to standard cmd.exe execution.
- Adjusted mcp stdio tests with a new non-node shim fixture and a simplified notify race case to validate transport teardown behavior.
Pressing Ctrl+T (toggle thinking-block visibility) — or any other path that calls rebuildChatFromMessages, e.g. the theme/preset selector — during the pre-streaming window after a submission cleared the just-submitted user message until the first assistant token arrived.
startPendingSubmission paints the user message optimistically before session.prompt(...) lands it in session entries; rebuildChatFromMessages reads buildTranscriptSessionContext() which has no record of it yet, so the rebuild erased it and the streaming-component re-add block in toggleThinkingBlockVisibility was a no-op (no stream yet).
Centralize the fix in rebuildChatFromMessages: after rendering the transcript, replay the in-flight optimistic user submission. EventController#handleMessageStart already clears optimisticUserMessageSignature once the real user message_start lands, so the replay is a no-op post-streaming and cannot duplicate. cancelPendingSubmission clears the signature and marks the submission cancelled before its own rebuild, so cancellation paths still wipe the message.
Fixes#2372
- Added SGR mouse routing to wizard scenes for hover, click, and wheel interactions.
- Added pointer-based tab selection and panel wheel scrolling in providers tabs.
- Added splash/outro left-click handling to start and complete the wizard flow.
- Added setup-wizard tests for mouse routing, splash click entry, and local coordinates.
- Used per-frame palette usage statistics to remap global indices into a compact local palette before PNG encoding.
- Encoded indexed outputs with dynamic 1-, 2-, and 4-bit depths based on used color count.
- Raised PNG compression from Balanced to High for indexed and RGB paths and updated parity and decode tests for variable-bit palette decoding.
Updated thinking visibility toggles to refresh assistant blocks in place instead of rebuilding the transcript, preserving pending user submissions and loaders before streaming starts. Added a regression test for the Ctrl+T pre-stream gap.\n\nFixes #2370
Restored cmd.exe's lookup order on Windows so an unqualified MCP command (e.g. server.cmd) checks the configured cwd before iterating PATH, keeping a project-local shim from being shadowed by a same-named global one.
- Implemented model-specific frame-size billing for Anthropic, OpenAI, and Google.
- Changed compaction shape resolution to bind model variants to ideal frame sizes.
- Updated tests and docs to reflect new frame-size and budget behavior.
Resolved Windows npm-generated .cmd MCP shims to their Node entrypoint before spawning so CodeGraph keeps ownership of stdio instead of disconnecting behind a transient cmd.exe wrapper.
Fixes#2367
Widened Kimi K2.6 OpenAI-compatible stream watchdog defaults so long reasoning starts do not hit the generic first-event timeout. Covered Fire Pass public and router ids with a regression test.\n\nFixes #2366
- Adjusted snapcompact rendering to compute used rows from text, dim toggles, and doc line breaks, then derive output height from usedRows x lineRepeat x cellHeight.
- Updated indexed and RGB PNG encoders to accept explicit canvas width and height so native renders now emit non-square frames matching actual content.
- Expanded Rust and TypeScript tests and updated docs/changelogs to assert and describe variable-height frame behavior.
- Updated snapcompact selectors and previews to pass `ShapeTarget` into `resolveShape` so `auto` is model-tuned.
- Applied `providerFrameBudget` in session compaction so generated archives stay within provider image caps.
- Replaced inline hard-coded image limits with `providerImageBudget` and skipped rasterization at cap.
- Added OpenRouter inline-transformer tests for cap exhaustion and existing-image budget exhaustion behavior.
- Added chunked and cached Kimi mono-prod probing with hit/miss diagnostics.
- Added parse_robust fallback and cache re-score path for numbered QA answer parsing.
- Added multi-route probe execution modes including smoke, bill, frame, AB, and last-line checks.
- Changed mono_prod CLI to support custom shape JSON/name and pricing overrides.
- Added support for `{api,id}` `ShapeTarget` in `resolveShape`, deriving variants from model IDs.
- Added provider budget APIs/constants and `providerFrameBudget` clamping with `MAX_FRAMES`.
- Added `Archive.textTail` and moved overflow pages into plain-text tail folding across frames.
- Updated summary prompt rendering to show continuation only when `textTail` exists and append it as text.
Sent Hindsight retain timestamps with local timezone offsets and supplied timestamps for automatic and queued retains. Added regression coverage for client serialization and backend timestamp propagation.\n\nFixes #2363
- Previously, `anthropic-messages` requests using `resolveWireModelId` would always derive the non-thinking variant for `requestModelId`, even when reasoning was explicitly enabled.
- This change ensures the `reasoning` state is correctly passed to `resolveWireModelId`, allowing the API request to include the appropriate `X-thinking` model variant when thinking is active, standardizing behavior across providers.
- Repointed `SHAPES.anthropic` and the `anthropic-messages`/unknown-API fallback in `resolveShape` at the `6x12-dim` variant: production mono eval on claude-fable scored f1 .840 vs .877 for the repeated grid (within noise at n=25) at 37% lower cost, with no refusals.
- Reworded the `snapcompact.shape` descriptions in `settings-schema.ts` to drop per-provider eval-winner claims from the variant help text.
- Updated `snapcompact.test.ts` (new `6x12-dim` default render/compact assertions, `8x8r-bw` exercised via `resolveShape`), `snapcompact-inline.test.ts` frame math for the new geometry, and the `docs/compaction.md` shape sentence.
- Amended the snapcompact `[Unreleased]` entry that said the Anthropic default stayed `8x8r-bw` and added the default-switch entry.
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
- Added `normalizeMandatoryReasoningOptions` to `stream.ts`: models baked with `thinking.requiresEffort` floor omitted or disabled reasoning to `minimumSupportedEffort` instead of sending an explicit disable, fixing "Reasoning is mandatory for this endpoint and cannot be disabled" 400s on OpenRouter Gemini 3.x.
- Skipped the clamp for `suppressWhenOff` models, which handle off provider-side via an explicit wire config.
- Added `requires-effort.test.ts` covering omitted-reasoning clamping, `disableReasoning` suppression, untouched explicit efforts, flag-free pair routing, and flagged collapsed pairs.
- Added the pi-ai changelog entry.
- Added a `requestModelId` option to `AnthropicOptions`, serialized as `requestModelId ?? model.requestModelId ?? model.id` in `buildParams`.
- Replaced `resolveOpenAICompletionsModelId`'s pinned-id short-circuit with `resolveWireModelId`-based resolution ahead of the provider-specific id transforms.
- Passed `requestModelId: resolveWireModelId(model, reasoning)` from every `mapOptionsForApi` case in `stream.ts`, so collapsed `X`/`X-thinking` pairs on aggregators and custom providers switch to the thinking SKU when reasoning is enabled.
- Added the pi-ai changelog entry.
- Added `thinking.requiresEffort` to `ThinkingConfig` and baked it in `deriveThinking`/`fillThinkingWireDefaults` via `impliesMandatoryReasoning` for reasoning-only upstreams: Gemini 3.x, Gemini 2.5 Pro, the OpenAI o-series, MiniMax M2, and thinking-only `-reasoner`/`-reasoning` SKUs.
- Added `minimumSupportedEffort()` to `model-thinking.ts` as the clamp target for thinking-off requests on flagged models.
- Moved `stripThinkingVariantToken`/`findThinkingVariantToken` into `identity/family.ts`, taught them the `-reasoning`/`-reasoner` spellings, and re-pointed the `variant-collapse` and coding-agent `model-resolver` imports.
- Dropped `requiresEffort` (with `effortRouting`/`suppressWhenOff`) from collapsed-pair thinking surfaces in `derivePairThinkingSurface`, since the collapsed pair routes off to the bare backing id.
- Regenerated `models.json` and covered derivation, backfill, and reasoning-token pairing in `model-thinking.test.ts` and `variant-collapse.test.ts`.
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.
- Resolved OpenAI shape resolution to `openai` and default to `8on16-bw`.
- Fixed catalog generation to collapse effort tiers before provider grouping.
- Updated help text and schemas to describe `8on16-bw` as the OpenAI auto default.
- Added production `render_pages` and `mono_prod` scripts for end-to-end QA output.
- Added the `6x12-dim`, `8x13-bw`, `8on16-bw`, `doc-8on16-bw`, `doc-8on16-sent`, and `doc-8on16-sent-dim` options to the `snapcompact.shape` enum in `settings-schema.ts` with per-variant descriptions naming each eval winner.
- Updated the `snapcompact.shape` paragraph in `docs/compaction.md` to list the new grid and two-column doc variants.
- Updated the coding-agent changelog: rewrote the `snapcompact.shape` Added entry for the new variants and recorded the Changed/Fixed entries for the already-landed variant-alias and pending-preview commits (one contiguous changelog run).
- Added `ToolRenderer.provisionalPendingPreview` in `renderers.ts` and consulted it from `ToolExecutionComponent.isTranscriptBlockCommitStable`, replacing the 15.11.6 blanket gate that marked every collapsed pending preview commit-unstable.
- Marked only the tail-window previews the result render re-anchors as provisional: the edit streamed-diff tail (`edit/renderer.ts`), bash/ssh command caps (`createShellRenderer`, `sshToolRenderer`), and eval cells with interleaved outputs (`evalToolRenderer`).
- Restored mid-stream scrollback commits for every other pending preview, so tool calls taller than the viewport (e.g. a task call's context/assignment markdown) no longer read as cut off until the result lands.
- Added a regression test in `tool-live-region-scrollback.test.ts` scroll-appending a tall collapsed streaming task call into native scrollback mid-stream.
- Regenerated the bundled catalog via `generate-models` to pick up effort-tier variant collapsing (raw `-low`/`-high`/`-thinking` member ids folded into logical entries with `thinking.effortRouting`) and display-name cleaning (gateway author prefixes and alias/price/promo decorations dropped).
- Added `cleanModelName` to `utils.ts`, dropping gateway author prefixes (`OpenAI: …`), `(latest)` alias markers, `(Antigravity)` attribution, price tiers (`($$$$)`), and promo/lifecycle tags (`(20% off)`, `(retires …)`) while preserving variant tags that map to distinct wire ids (`(Thinking)`, `(free)`, `(Fast)`, dates, regions).
- Applied it in `buildModel` (covers live discovery and stale caches) and as a display-name normalization pass in `generate-models.ts`; Antigravity discovery no longer appends `(Antigravity)` to display names.
- Added name-cleaning coverage to `build.test.ts`.
- Changelog entry for this change landed with the variant-collapse commit (same contiguous `CHANGELOG.md` run).