- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Added Silver.ttf TrueType font support to `pi-natives` with automated fallback logic for bitmap font rendering.
- Implemented wide code point detection and cell-width calculation to improve CJK character handling and layout.
- Introduced dynamic font-aware preflight probing via `resolveShapeForText` and `renderabilityProbeText` for better font selection.
- Enabled semantic emoji folding and improved text normalization to handle non-Latin characters and emoji filtering.
AgentSession.switchSession() eagerly called buildDisplaySessionContext()
before setSessionFile, walking the previous session's branch and expanding
every compaction entry's snapcompact archive and openaiRemoteCompaction
replacementHistory into messages. For huge pre-fix sessions that materialized
GBs of data and OOMed in-TUI /resume even after the streaming loader fix.
The snapshot is only needed for same-session reloads, where
#didSessionMessagesChange compares the pre/post message arrays to detect
rollback edits. Different-session switches skip the call entirely; the
error-recovery path rebuilds the previous context on demand from the
restored state so MCP-selection restoration still has its inputs.
Added a regression test (test/agent-session-switch-prev-context.test.ts)
that spies on sessionManager.buildSessionContext across switchSession and
asserts the expected call count and target file per branch.
Fixes#3846
- Always remove surfaced irc:incoming records from the pending-aside queue; the inbox tool result already injects the body, so leaving them queued would auto-inject a duplicate at the next step.
- Updated the inbox tool to drain pending asides regardless of peek.
- Added a regression test asserting a peeked pending aside does not auto-inject.
- Drained running-session IRC asides through the inbox tool before the model step consumes them.
- Added a regression test for messages delivered while the recipient is already running.
Fixes#3834
Restrict the supersede sweep to compactions on the path from the current leaf so a newer compaction never rewrites a sibling branch's still-current summary or drops its preserveData. Streaming load now collects the active-branch ids before eliding instead of trampling sibling compactions encountered in file order.
Refs #3789
Stream large session loads, elide superseded compaction payloads, skip synchronous rewrites when the append-only file is already current, and provide usable picker previews for developer-started forks.
Fixes#3789
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
Drop the persisted assistant error before #runAutoCompaction so the kept region is clean, then re-append it on COMPACTION_CHECK_NONE without a fresh compaction entry — covers no-model, hook-cancel, and compaction-error paths. Same rollback for response.incomplete recovery.
Fixes#3747
Split active-context vs persisted-history removal in #checkCompaction so the persisted assistant error stays on the branch unless context promotion or compaction is actually scheduled.
Fixes#3747
Removed recoverable context-overflow and incomplete-response assistant errors from both active context and persisted session history before compaction/promotion schedules the retry.
Fixes#3747
- Implement cleanup logic to drop `thinkingSignature` values from assistant thinking blocks during persistence.
- Identify and drop signatures only when the underlying reasoning data is already recoverable via the `providerPayload` items.
- Ensure orphaned signatures that cannot be reconstructed from the payload are preserved during serialization.
- Add comprehensive test coverage to verify deduplication safety and edge-case handling for missing payloads.
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
- Removed the architectural restriction limiting advisors to read-only tools.
- Updated advisor configuration to permit any built-in tool, including `edit`, `write`, and `bash`.
- Defaulted advisor toolsets to `read`, `grep`, and `glob`, while maintaining strict session isolation for each advisor.
- Introduced comprehensive support for multiple concurrent, independently-configured advisors via `WATCHDOG.yml` files.
- Implemented a full-screen TUI overlay for managing advisor rosters, models, tools, and instructions.
- Added session-wide advisor initialization, telemetry aggregation, and named transcript isolation.
- Enhanced advisor security and observability with secret redaction in tool results and secure XML attribute encoding.
- Introduced normalization for title overrides to handle empty strings as null.
- Enabled atomic persistence for session title updates via `SessionManager`.
- Implemented conditional storage index restoration for failed title updates.
- Moved title persistence logic to a dedicated helper method to ensure consistency.
- Implement provider-native replay logic to enable reuse of remote compaction data across compatible models.
- Enhance compaction logic to re-expand and locally summarize remote history when provider-native replay is unavailable.
- Update OpenAI request setup to include session and routing identifiers for improved traceability.
- Refine token estimation for image content during truncation to ensure more accurate budget management.
- Update agent-session to resolve compaction model candidates before persistence, ensuring authentication availability.
Replaced the controller-side switchActiveModel flag with a currentContextTokens hint on AgentSession.setModel, so the over-context decision is computed against the refreshed candidate metadata. setModel returns whether the live switch happened, and the Alt+M default-role path uses that to gate the live side effects.
Decoupled default-role persistence from live model switching when the selected model is below the current session context window.
Updated the model selector regression coverage so the Alt+M Default action remains selectable and advances to thinking selection.
Fixes#3708
- Updated the demotion logic to treat complete, signed thinking blocks as stable content that terminates an interrupted stream.
- Protected signed thinking runs from being stripped when processing interrupted agent messages.
- Updated `convertToLlm` to detect and strip trailing thinking runs from user-interrupted assistant messages.
- Retained original thinking content on the persisted assistant message to ensure UI state (render, reload, and rebuild) remains intact.
- Added `followedByInterruptedThinking` helper to verify continuity message presence before stripping for the provider request.
- Updated persistence tests to confirm thinking remains in the session state but is excluded from LLM context headers.
- Introduced `resolveToolEventInput` to enable mode-specific transformation of editor tool inputs before normalization.
- Updated `ExtensionToolWrapper` and `HookToolWrapper` to utilize the new resolver during tool execution.
- Added support for tracking `reasoningTokens` in `SessionStats` and `AdvisorStats` within the agent session.
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
- Added `INTERRUPTED_THINKING_MESSAGE_TYPE` and `demoteInterruptedThinking` in `session/messages.ts` for stripping unfinished thinking into durable context.
- Persisted hidden interrupted-thinking context from `session/agent-session.ts` immediately after user-interrupted assistant turns.
- Added the `prompts/system/interrupted-thinking.md` envelope used for replaying interrupted reasoning.
- Added focused interrupted-thinking coverage in `test/agent-session-interrupted-thinking.test.ts` and `test/session/interrupted-thinking-demote.test.ts`.
- Introduced `TitleChangeEntry` type and `TITLE_CHANGE_ENTRY_TYPE` to record title modifications in session logs.
- Added `titleUpdatedAt` and `hasTitleSlot` to `SessionManager` state for persistent tracking of metadata.
- Updated `setSessionName` to serialize title changes into the session file via dedicated title slots.
- Modified session file output to include a title slot header for improved auditability.
- Added a fixed-width title slot system to serialize and persist session titles across physical files and backend storage.
- Integrated automated session title updates triggered by todo replan operations using conversation history context.
- Extended storage interfaces across memory, file-system, Redis, and SQL backends to support independent title metadata updates.
- Implemented title metadata parsing within session loaders and list utilities to ensure accurate retrieval and display.
Split image-bearing custom messages into developer text plus user image content before provider conversion so queued skill images stay out of developer slots.
Mapped user-attributed skill prompt custom messages to user-role model messages while preserving developer framing for auto-applied skills and other custom messages.
Fixes#3698
Auto-compaction now falls back to context-full on any snapcompact preflight or post-render rejection (non-ASCII, kept history too large, projection overflow), matching the text-only-model handling. Manual /compact snapcompact keeps its local-only failure contract per #3599.
Fixes#3659