The per-entry 32 KB cap let a many-file batch (apply_patch / hashline
touching N files) accumulate unbounded snapshot bytes because each
`perFileResults` entry was checked independently. 100 files × 30 KB
each kept everything (~3 MB) even though the whole array still
serializes into one session JSONL line.
`capPerFileSnapshots` walks entries left-to-right with one shared
`MAX_EDIT_SNAPSHOT_TEXT_CHARS` budget. Each per-entry payload is still
capped individually by `pruneSnapshot`; if an entry's surviving bytes
would push the running aggregate past the cap, the entry is stripped
and stamped with `snapshotsPruned: true`. Early entries keep their ACP
diff visualization; later entries in a large batch degrade to text-only
exactly like over-sized single edits.
Regression test exercises five equal-size entries that each fit the
per-entry budget but bust it cumulatively, asserting only the first two
keep snapshots and the trailing three carry the pruned marker.
When a multi-entry single-path edit prunes the first entry's snapshots
(large pre-image) and keeps a later entry's snapshots (file shrunk
between entries), the aggregator at `executeSinglePathEntries` recorded
the later entry's small `oldText`/`newText` as the whole-file
transition. ACP clients would then render a misleading partial diff
instead of degrading to text-only for the over-budget edit.
Add an explicit `snapshotsPruned` marker on `EditToolDetails` /
`EditToolPerFileResult`, set by `pruneSnapshot` whenever it strips a
payload. `executeSinglePathEntries` tracks the flag across child
results and suppresses aggregate `oldText`/`newText` (re-stamping the
marker on the aggregate) the moment any child was pruned;
`executeApplyPatchPerFile` propagates the flag onto each per-file entry.
Regression test exercises the exact scenario the reviewer raised on
#3787: replace mode where entry 1 collapses a >1 MB file to a single
line (pruned) and entry 2 trivially renames the now-tiny result; the
aggregate result now carries `snapshotsPruned: true` with both
snapshot fields omitted.
Edit-tool results carried the full pre/post file content in
`details.oldText` / `details.newText`. For large files this bloated each
per-turn JSONL line by hundreds of KB even though the snapshots are
never sent to the LLM (provider serializers send only `content`) and
only consumed by the ACP event mapper for diff visualization.
Add `pruneOversizedEditSnapshots` and apply it at every site that
constructs an `EditToolDetails` / `EditToolPerFileResult`:
`executePatchSingle`, `executeReplaceSingle`, hashline `renderSection`
(delete + update branches), and both aggregators in `edit/index.ts`.
When combined `oldText` + `newText` exceeds 32 KB the helper returns a
shallow copy with both fields omitted; smaller edits pass through
unchanged. The diff, path, firstChangedLine, op, move, and diagnostics
fields are preserved, and ACP returns no diff content for over-budget
files (the text content still flows — graceful degradation).
Fixes#3786
- Optimized TUI tool argument previews by throttling JSON re-parsing to prevent frame starvation during high-frequency streaming.
- Suppressed redundant component updates for unchanged parsed fields while maintaining raw preview integrity for bash and patch renderers.
- Added adaptive parsing logic to `ToolArgsRevealController` that distinguishes between renderers requiring continuous raw JSON streams and those consuming parsed arguments.
- Updated `EventController` to dynamically determine exposure requirements based on tool type and wire-format metadata.
- Added test suite to validate service tier resolution logic in benchmark commands.
- Verified that command-line flags correctly override application setting values.
- Confirmed default behavior when neither flags nor configuration settings are provided.
- Optimized instruction sets for core agent tools including task, lsp, job, and irc.
- Standardized tool documentation structure by replacing parameter listings with structural instruction blocks.
- Mandated new communication and technical workflows for subagent results, symbol-aware code intelligence, and background task management.
- Refined messaging and coordination guidelines to prioritize inter-agent communication and direct operations.
- Added --par flag to execute benchmark runs concurrently with a default degree of 4.
- Added --service-tier flag to allow overriding the provider service tier per benchmark.
- Increased default benchmark run count from 1 to 10 to provide more robust averaging.
- Updated benchmarking logic to process requests in a concurrency-limited pool while preserving output order.
- Implemented pre-flight credential checks to prevent unnecessary worker spawning when authentication is missing.
llama.cpp /props.default_generation_settings.params.{max_tokens,n_predict} are per-request defaults the server applies when a client omits the field, not a hard model cap. Only the -1 unlimited sentinel is promoted to the runtime context window now; positive values fall back to the discovery default so client-side per-request overrides remain unconstrained.
Fixes#3781
Resolved selected-model refresh maxTokens against the effective context window, including live contextWindow overrides, so unlimited llama.cpp caps cannot exceed the configured context.
Fixes#3781
Gated discoverLlamaCppModelRuntimeMetadata's /props context fallback on the selected entry being present in /models, so refreshSelectedModelMetadata never patches a stale cached id with a different model's runtime metadata.
Fixes#3781
Mapped llama.cpp -1 generation limits from /props to the discovered runtime context window instead of the generic discovery default, including selected-model metadata refresh.
Fixes#3781
The auto-snapcompact non-ASCII fallback test seeded a single CJK turn with a
245k-token usage. claude-sonnet-4-5's 200k context window made that an input
overflow, so the recovery path dropped the assistant turn and left nothing to
summarize; the default 20k keep-recent window also kept both tiny messages.
Either way prepareCompaction returned undefined, the renderability preflight
never ran, and the strategy stayed "snapcompact" instead of downgrading to
context-full.
Derive the synthetic prompt size from the active model's context window (above
the auto-compaction threshold, below the window) and force keepRecentTokens=1
so the CJK history is always summarized and scanned, regardless of model
metadata.
Fixed the bash interceptor rule so allowed /dev sink redirects are skipped while scanning for later real file redirects in the same command.
Fixes#3763
Fixed the bash interceptor's echo/printf redirect rule so device sinks under /dev/null, /dev/tty, /dev/stdout, and /dev/stderr remain executable while real file redirects are still blocked.
Fixes#3763
Used compact hashed isolation directory segments and the short m mount dir so long task ids are not copied into subagent working paths.
Kept worktree cleanup compatible with legacy merged task-isolation directories.
Fixes#3756
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Ollama and llama.cpp discovery now prefer runtime context settings over model training metadata, so compaction thresholds match the window local servers actually accept.
Fixes#3752
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.
Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).
Fixes#3749
- Collapsed git worktree path in status line to project name with icon.
- Fixed out-of-workspace file edits by including the full path in headers.
- Fixed structured output schema violations by correctly handling payload nesting in terminal yields.
- Added test suites for git worktree logic, out-of-cwd reading, and subagent output serialization.
- Adjusted hashline header formatting to preserve absolute file paths instead of truncating them to basenames.
- Ensured absolute paths are passed through shortenPath to allow resolution while keeping home directory references concise.
- Prevented edit failures when reading files outside the workspace by ensuring tags remain resolvable.
- Refactored payload assembly to distinguish between incremental sections and terminal results.
- Prevented terminal markers from being incorrectly treated as section labels to avoid payload nesting issues.
- Updated section processing to ignore non-incremental terminal items, resolving incorrectly missing data in output-schema validation.
- Improved terminal item resolution to correctly fallback to the last assistant text when no explicit data is provided.
- Added `git.repo.linkedWorktreeSync` to identify and resolve git worktree metadata without spawning subprocesses.
- Updated `StatusLineComponent` to detect linked worktrees and resolve project/worktree context names.
- Modified path segment rendering to collapse nested git worktree paths and display the worktree name when it diverges from the active branch.
- Introduced `icon.worktree` symbol across themes to visually distinguish git worktree paths.
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.
Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).
Fixes#3749
Drop the persisted assistant error before #runAutoCompaction so the kept region is clean, then re-append it on COMPACTION_CHECK_NONE without a fresh compaction entry — covers no-model, hook-cancel, and compaction-error paths. Same rollback for response.incomplete recovery.
Fixes#3747
Split active-context vs persisted-history removal in #checkCompaction so the persisted assistant error stays on the branch unless context promotion or compaction is actually scheduled.
Fixes#3747
Removed recoverable context-overflow and incomplete-response assistant errors from both active context and persisted session history before compaction/promotion schedules the retry.
Fixes#3747
Use the pre-session memory prompt snapshot as the learned.md baseline when startup consolidation refreshes before a session-scoped cache exists. This keeps active-session learn writes out of the refreshed prompt while still surfacing the new consolidated summary.
Refs #3743
Refresh the startup consolidation summary without rereading learned.md for the active session, so lessons captured while startup is still running remain deferred to the next session.
Refs #3743
Background memory startup writes memory_summary.md after the initial system-prompt build has already cached an empty value for the session. Drop the cached snapshot before refreshBaseSystemPrompt so the active session actually picks up the freshly consolidated summary instead of returning the stale cached one.
Refs #3743
Stopped passive autolearn from adding hidden conversation messages and froze local memory developer instructions per session so learn writes land in future sessions instead of mutating the active Anthropic prompt prefix.
Fixes#3743
Per #3740 review: many short strings (e.g. a tool result whose content array holds thousands of small text blocks) could sum past MAX_REPLICATED_PAYLOAD_BYTES without any individual field crossing the per-string floor, so the helper exited the truncation loop and shipped an oversized frame — the relay close/reconnect loop the helper was meant to prevent.
Replace the string-only truncation pass with a single walker that head-truncates strings AND head-clips arrays in one descent, driven by a concrete SHRINK_PASSES schedule that tightens both axes together. The final pass clamps every string to 64 B and every array to one element, so any payload converges. Add direct unit tests for shrinkForReplication covering: identity for small values, single-giant-string clamp, many-short-strings array clamp (no field above floor), and discriminator preservation on a fully-shrunk payload.