Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.
Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).
Fixes#3749
Split active-context vs persisted-history removal in #checkCompaction so the persisted assistant error stays on the branch unless context promotion or compaction is actually scheduled.
Fixes#3747
Removed recoverable context-overflow and incomplete-response assistant errors from both active context and persisted session history before compaction/promotion schedules the retry.
Fixes#3747
Use the pre-session memory prompt snapshot as the learned.md baseline when startup consolidation refreshes before a session-scoped cache exists. This keeps active-session learn writes out of the refreshed prompt while still surfacing the new consolidated summary.
Refs #3743
Refresh the startup consolidation summary without rereading learned.md for the active session, so lessons captured while startup is still running remain deferred to the next session.
Refs #3743
Background memory startup writes memory_summary.md after the initial system-prompt build has already cached an empty value for the session. Drop the cached snapshot before refreshBaseSystemPrompt so the active session actually picks up the freshly consolidated summary instead of returning the stale cached one.
Refs #3743
Stopped passive autolearn from adding hidden conversation messages and froze local memory developer instructions per session so learn writes land in future sessions instead of mutating the active Anthropic prompt prefix.
Fixes#3743
Per #3740 review: many short strings (e.g. a tool result whose content array holds thousands of small text blocks) could sum past MAX_REPLICATED_PAYLOAD_BYTES without any individual field crossing the per-string floor, so the helper exited the truncation loop and shipped an oversized frame — the relay close/reconnect loop the helper was meant to prevent.
Replace the string-only truncation pass with a single walker that head-truncates strings AND head-clips arrays in one descent, driven by a concrete SHRINK_PASSES schedule that tightens both axes together. The final pass clamps every string to 64 B and every array to one element, so any payload converges. Add direct unit tests for shrinkForReplication covering: identity for small values, single-giant-string clamp, many-short-strings array clamp (no field above floor), and discriminator preservation on a fully-shrunk payload.
CollabHost shipped the first entry of every snapshot-chunk batch unconditionally, and broadcast live entry/event frames verbatim, so a single multi-megabyte tool result (read/bash/search) overflowed the relay's per-frame maxPayloadLength. The relay closed the host's WebSocket with 1006 ("Received too big message"), CollabSocket treated 1006 as non-fatal and reconnected, the next guest hello triggered the same oversized send, and the host status line cycled "Collab relay connection lost, reconnecting…" indefinitely.
Add shrinkForReplication: any host->guest payload whose JSON exceeds MAX_REPLICATED_PAYLOAD_BYTES (1 MB) is deep-cloned with long strings head-truncated and an "[…N chars elided for collab session]" marker; otherwise the original reference passes through. Apply it to snapshot chunk entries, live entry broadcasts, and live event broadcasts (including large tool_execution_end results). Regression test stands up a Bun.serve relay with 8 MB maxPayloadLength and a snapshot containing a 5 MB entry; asserts the host stays connected, the snapshot train finalizes, and the guest sees the entry with the elision marker.
Fixes#3739
InputController.handleFollowUp read raw editor text via getText(),
bypassing the paste-store expansion the Enter path applies through
Editor.getExpandedText(). A large paste collapsed into a [Paste #N, +X
lines] marker was therefore sent verbatim to the model when queued with
Ctrl+Q / Ctrl+Enter, silently dropping the pasted content.
Switch the follow-up path to getExpandedText() so queued submissions
match the Enter path. Image markers are untouched; pendingImages
forwarding is unchanged.
Updated existing input-controller stubs (skill-queue, followup-image,
keybindings) to implement getExpandedText, matching the production
CustomEditor surface.
Fixes#3737
- Added `recoverSectionPathFromTag` to reconcile bare or mismatched `[basename#tag]` paths using existing session snapshots.
- Implemented `readSectionForPreview` to fallback to recovered file paths when the authored path is absent.
- Updated `computeHashlineSectionDiff` to use the path recovery logic during preview content acquisition.
- Removed obsolete regression test file `issue-1765-repro.test.ts`.
- Updated temporary directory prefixes to include an @ symbol in multiple test suites.
- Standardized file paths for temporary agent storage across test environments.
- Implemented a queueing mechanism in `ToolExecutionComponent` to prevent starvation of edit previews during high-frequency argument updates.
- Replaced eager cancellation of in-flight diff computations with a drain loop that ensures every update is processed once the current compute settles.
- Added `partialJsonOf` helper to safely narrow streamed JSON buffers from tool arguments.
- Added regression test to verify that slow diff computations are not aborted by incoming stream chunks and instead queue a subsequent re-run.
- Implement cleanup logic to drop `thinkingSignature` values from assistant thinking blocks during persistence.
- Identify and drop signatures only when the underlying reasoning data is already recoverable via the `providerPayload` items.
- Ensure orphaned signatures that cannot be reconstructed from the payload are preserved during serialization.
- Add comprehensive test coverage to verify deduplication safety and edge-case handling for missing payloads.
- Removed support for `history://` URI schemes used to read agent transcripts from system and tool prompts.
- Updated IRC tool instructions to remove references to reading agent history for peer information.
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.