- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Added a cross-turn tool-call loop guard that hashes canonical tool names and arguments, ignores intent metadata, and injects a hidden redirect when identical calls reach the configured threshold.
Fixes#3971
Two termination boundaries in the browser tool leaked browser-owned OS resources into the long-lived coding-agent process.
1. Aborted 'open' published an orphan. #open wrapped acquisition in untilAborted, which rejects its outer wrapper on abort but lets the inner launch resolve in the background; acquireBrowser then unconditionally stored the resolved handle in the module-global browsers map. releaseAllTabs walks tabs, not browsers, so the refCount:0 handle stayed alive to process exit.
2. Session dispose had no browser teardown. Browser/tab state lives in module-global maps, and AgentSession.dispose() had no hook to walk them, so headless/spawned Chromium the session opened survived it.
acquireBrowser now short-circuits before launch on a pre-aborted signal and disposes the handle when the launch completes after abort. TabSession records the creating session's id (opts.ownerSessionId, threaded through BrowserTool.#open), preserved across reuse so a subagent re-driving an existing tab does not yank teardown responsibility. AgentSession.dispose() invokes releaseTabsForOwner bounded by withTimeout(3s), mirroring the async-job/MCP disposal pattern.
Regression tests exercise both boundaries via spied CmuxSocketClient (no real puppeteer/socket) and cover: pre-aborted open short-circuit, aborted-mid-launch cleanup, releaseTabsForOwner reaping only owned tabs, and reuse preserving original ownership.
Fixes#3963
- Replaced leaf-to-root unshift path assembly with push plus one reverse in buildSessionContext and SessionEntryIndex.pathTo.
- Added regression coverage that keeps deep linear context and branch paths root-to-leaf without Array.unshift work.
Fixes#3961
- Replaced the map-based persisted message index with a Set of key identities.
- Removed the message content equality check during branch verification to avoid false out-of-order skips when display-side content variants are present.
- Added a hidden system notice prompt to instruct the model to break repetitive behaviors when a thinking or response loop is detected.
- Injected the redirect notice into the retried turn's context when resetting the active context after a loop retry.
- Added unit tests to verify the custom redirect message is appended, configured as non-displaying, and visible in subsequent LLM contexts.
Captured the configured thinking selector when entering plan mode so approving a plan restores auto instead of the provisional concrete effort. Reloaded DEFAULT(auto) badges from defaultThinkingLevel and covered the plan-approval handoff plus /model display.
Fixes#3901
- Added a last-resort recovery step to run `shake("elide")` on oversized message tails when auto-compaction cannot otherwise free enough context.
- Re-tests the context headroom and auto-continue predicates after a successful rescue before falling back to pausing maintenance.
- Updated the dead-end warning message to suggest running `/shake images` for irreducible, image-only tails.
fix(compaction): cap snapcompact frame payloads (#3866)
Bound rebuilt snapcompact image payloads by a per-request base64 byte
budget so long sessions stop re-sending multi-megabyte standing image
archives on every provider request; auto-compaction falls back to
context-full summaries when snapcompact output is too large.
Resolved snapcompact.ts conflict against the main font-rendering refactor
by keeping both renderabilityProbeText and the frame-budget helpers.
Fixed historyBlocks to emit the omitted-frame notice before the kept
(newer) images, since the byte budget drops the oldest frames — keeping
reconstructed blocks oldest-to-newest (addresses Codex P2 review).
Fixes#3792
Forwarded persisted provider stream timeout settings into model requests so slow local LLM streams can widen or disable first-event and idle watchdogs without environment variables.
Fixes#3878
Apply the snapcompact frame byte-budget cap even when the active model has no known context window, avoiding 80-frame archives on custom vision models.
Bounded persisted snapcompact image archives by base64 byte size so large sessions stop re-sending multi-megabyte frame walls on every provider request.
Auto snapcompact now falls back to context-full summaries when rendered frame payloads exceed the byte budget, and legacy oversized archives omit over-budget frames during LLM context rebuilds.
Fixes#3792
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Added Silver.ttf TrueType font support to `pi-natives` with automated fallback logic for bitmap font rendering.
- Implemented wide code point detection and cell-width calculation to improve CJK character handling and layout.
- Introduced dynamic font-aware preflight probing via `resolveShapeForText` and `renderabilityProbeText` for better font selection.
- Enabled semantic emoji folding and improved text normalization to handle non-Latin characters and emoji filtering.
AgentSession.switchSession() eagerly called buildDisplaySessionContext()
before setSessionFile, walking the previous session's branch and expanding
every compaction entry's snapcompact archive and openaiRemoteCompaction
replacementHistory into messages. For huge pre-fix sessions that materialized
GBs of data and OOMed in-TUI /resume even after the streaming loader fix.
The snapshot is only needed for same-session reloads, where
#didSessionMessagesChange compares the pre/post message arrays to detect
rollback edits. Different-session switches skip the call entirely; the
error-recovery path rebuilds the previous context on demand from the
restored state so MCP-selection restoration still has its inputs.
Added a regression test (test/agent-session-switch-prev-context.test.ts)
that spies on sessionManager.buildSessionContext across switchSession and
asserts the expected call count and target file per branch.
Fixes#3846
- Always remove surfaced irc:incoming records from the pending-aside queue; the inbox tool result already injects the body, so leaving them queued would auto-inject a duplicate at the next step.
- Updated the inbox tool to drain pending asides regardless of peek.
- Added a regression test asserting a peeked pending aside does not auto-inject.
- Drained running-session IRC asides through the inbox tool before the model step consumes them.
- Added a regression test for messages delivered while the recipient is already running.
Fixes#3834
Restrict the supersede sweep to compactions on the path from the current leaf so a newer compaction never rewrites a sibling branch's still-current summary or drops its preserveData. Streaming load now collects the active-branch ids before eliding instead of trampling sibling compactions encountered in file order.
Refs #3789
Stream large session loads, elide superseded compaction payloads, skip synchronous rewrites when the append-only file is already current, and provide usable picker previews for developer-started forks.
Fixes#3789
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
Drop the persisted assistant error before #runAutoCompaction so the kept region is clean, then re-append it on COMPACTION_CHECK_NONE without a fresh compaction entry — covers no-model, hook-cancel, and compaction-error paths. Same rollback for response.incomplete recovery.
Fixes#3747
Split active-context vs persisted-history removal in #checkCompaction so the persisted assistant error stays on the branch unless context promotion or compaction is actually scheduled.
Fixes#3747
Removed recoverable context-overflow and incomplete-response assistant errors from both active context and persisted session history before compaction/promotion schedules the retry.
Fixes#3747
- Implement cleanup logic to drop `thinkingSignature` values from assistant thinking blocks during persistence.
- Identify and drop signatures only when the underlying reasoning data is already recoverable via the `providerPayload` items.
- Ensure orphaned signatures that cannot be reconstructed from the payload are preserved during serialization.
- Add comprehensive test coverage to verify deduplication safety and edge-case handling for missing payloads.
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
- Removed the architectural restriction limiting advisors to read-only tools.
- Updated advisor configuration to permit any built-in tool, including `edit`, `write`, and `bash`.
- Defaulted advisor toolsets to `read`, `grep`, and `glob`, while maintaining strict session isolation for each advisor.
- Introduced comprehensive support for multiple concurrent, independently-configured advisors via `WATCHDOG.yml` files.
- Implemented a full-screen TUI overlay for managing advisor rosters, models, tools, and instructions.
- Added session-wide advisor initialization, telemetry aggregation, and named transcript isolation.
- Enhanced advisor security and observability with secret redaction in tool results and secure XML attribute encoding.
- Introduced normalization for title overrides to handle empty strings as null.
- Enabled atomic persistence for session title updates via `SessionManager`.
- Implemented conditional storage index restoration for failed title updates.
- Moved title persistence logic to a dedicated helper method to ensure consistency.
- Implement provider-native replay logic to enable reuse of remote compaction data across compatible models.
- Enhance compaction logic to re-expand and locally summarize remote history when provider-native replay is unavailable.
- Update OpenAI request setup to include session and routing identifiers for improved traceability.
- Refine token estimation for image content during truncation to ensure more accurate budget management.
- Update agent-session to resolve compaction model candidates before persistence, ensuring authentication availability.
Replaced the controller-side switchActiveModel flag with a currentContextTokens hint on AgentSession.setModel, so the over-context decision is computed against the refreshed candidate metadata. setModel returns whether the live switch happened, and the Alt+M default-role path uses that to gate the live side effects.
Decoupled default-role persistence from live model switching when the selected model is below the current session context window.
Updated the model selector regression coverage so the Alt+M Default action remains selectable and advances to thinking selection.
Fixes#3708
- Updated the demotion logic to treat complete, signed thinking blocks as stable content that terminates an interrupted stream.
- Protected signed thinking runs from being stripped when processing interrupted agent messages.
- Updated `convertToLlm` to detect and strip trailing thinking runs from user-interrupted assistant messages.
- Retained original thinking content on the persisted assistant message to ensure UI state (render, reload, and rebuild) remains intact.
- Added `followedByInterruptedThinking` helper to verify continuity message presence before stripping for the provider request.
- Updated persistence tests to confirm thinking remains in the session state but is excluded from LLM context headers.