Ports only the thinking double-format fix: resolveThinkingDisplay reuses block.thinking when rawThinking is set (buildDisplayMessage already formatted it), plus a single-entry memo in formatThinkingForDisplay and a rawThinking regression test. The PR's incremental reveal slicing is superseded by the already-merged #3848 (memoized grapheme slicing).
- Filtered dot-only or blank thinking blocks so they no longer render as assistant thought.
- Adjusted assistant-message and streaming-reveal logic to use visible-thinking helpers for consistency.
- Recomputed `tools.discoveryMode: "auto"` in the deferred MCP closure in `sdk.ts` once the real tool count is known: a toolset crossing the threshold now flips discovery on, registers and activates `search_tool_bm25`, and skips `activateAll` instead of force-activating every MCP tool.
- Guarded the deferred MCP task against disposed sessions: added `AgentSession.isDisposed` and `enableMCPDiscovery()`, and the late connect now calls `disconnectAll()` instead of refreshing tools onto a dead session.
- Cleared `#fastPathKey`/`#fastPathItems` in `AssistantMessageComponent.invalidate()` so theme/symbol changes rebuild reused Markdown children instead of keeping stale captured themes.
- Memoized unusable read summaries as a `false` sentinel in `read.ts` so the per-session LRU no longer retains full sources of unsummarizable files.
- Broadened `HAS_REF_DEF` in `markdown.ts` to match backslash-escaped reference labels (`[a\]b]: x`) and cleared frozen stream-lex state on blank `setText()`.
- Added regression tests: deferred auto-discovery flip and mid-connect dispose (`sdk-mcp-auto-discovery.test.ts` + `many-tools-mcp.ts` fixture), fast-path child rebuild on invalidate, and escaped-ref-def incremental-lex equivalence.
- Defer MCP server discovery off the first-paint critical path for UI
sessions; tools and slash commands stream in through the existing
live-refresh channel once each server connects (non-UI modes keep the
blocking path). ~290ms off first paint with MCP servers configured.
- Build the model catalog's canonical-equivalence index lazily on first
read instead of eagerly in the ModelRegistry constructor. A default
interactive launch never reads it pre-paint, moving the ~210ms build
(over ~3,200 models) off the critical path: ~244ms (~16%) off cold boot.
- Memoize per-session read summaries (tree-sitter parse) on the content
hash of the freshly-read bytes; the file is still read fresh each call
so results stay correct. Repeat same-file summary read 17ms -> 2.4ms.
- Reuse the Markdown subtree across streaming reveal ticks, memoize
grapheme counting, and stop re-highlighting finalized thinking blocks.
- Attribute the previously-unlabeled synchronous boot region in the
PI_TIMING table and add a bench:guard boot-regression target.