AgentSession.switchSession() eagerly called buildDisplaySessionContext()
before setSessionFile, walking the previous session's branch and expanding
every compaction entry's snapcompact archive and openaiRemoteCompaction
replacementHistory into messages. For huge pre-fix sessions that materialized
GBs of data and OOMed in-TUI /resume even after the streaming loader fix.
The snapshot is only needed for same-session reloads, where
#didSessionMessagesChange compares the pre/post message arrays to detect
rollback edits. Different-session switches skip the call entirely; the
error-recovery path rebuilds the previous context on demand from the
restored state so MCP-selection restoration still has its inputs.
Added a regression test (test/agent-session-switch-prev-context.test.ts)
that spies on sessionManager.buildSessionContext across switchSession and
asserts the expected call count and target file per branch.
Fixes#3846
The initial PR update changed the MCP OAuth default prompt to 'login consent'.
The reporter verified Cloudflare's flow actually matches the reference MCP SDK
when no prompt parameter is sent: Cloudflare then reuses the existing account
grant and opens the scope/permission picker first. Forcing any prompt keeps the
flow on Cloudflare's account/consent page instead.
Match the reference SDK behavior: omit prompt by default and send
'prompt=consent' only when the requested scope contains offline_access, where
OIDC Core requires re-consent for offline access. Explicit oauth.prompt values,
including the empty-string omit escape hatch, still take precedence.
Also rename the dynamically registered MCP OAuth client from Codex to oh-my-pi
so Cloudflare consent screens show the current product name.
Fixes#3817
The MCP OAuth flow defaulted the authorization-request prompt parameter
to 'consent'. Per OpenID Connect Core 1.0 §3.1.2.1 that asks the
authorization server to re-prompt for consent only while reusing the
existing browser authentication session. Cloudflare's MCP OAuth server
(and other strict OIDC providers) honor that literally, so /mcp reauth
landed on the consent screen attached to whichever account the browser
cookie was for, leaving no way to switch the signed-in account.
Default to 'login consent' instead so the provider first re-prompts for
authentication (the page Claude Code shows on its reauth flow) and then
re-confirms consent, preserving the original intent of always
re-displaying the authorize screen. RFC 6749 §3.1 requires providers to
ignore prompt values they do not support, so the two-value form is safe
for non-OIDC servers. Existing per-server overrides via mcp.json's
`oauth.prompt` (including the empty-string escape hatch) are unchanged.
Fixes#3817
- Added management of provider session states during benchmark execution.
- Implemented a teardown process to close and clear session states after request completion.
Enabled the Gemini web search provider to use standard Google developer API credentials when Cloud Code Assist OAuth is absent.
Added developer API request coverage for native Google Search grounding and preserved existing OAuth request serialization.
Fixes#3810
Rejoined split Windows extension module paths before launch parsing finishes and stripped extended-length Win32 prefixes before Bun import and worker spawn APIs see them.
Fixes#3804
The DuckDuckGo provider hit api.duckduckgo.com (the Instant Answer API),
which only serves Wikipedia / Wolfram-Alpha-style topics — empty
AbstractText / Results / RelatedTopics for the vast majority of agent
queries. The orchestrator then rejected the empty response and surfaced
'DuckDuckGo returned no renderable search content', leaving users with
no working free fallback.
Switch the provider to POST html.duckduckgo.com/html/ (the no-JS HTML
frontend) with a browser User-Agent, parse the result blocks (unwrapping
//duckduckgo.com/l/?uddg=… redirect URLs), and map recency to the df
form field (d/w/m/y). When DuckDuckGo serves the bot-detection modal
(HTTP 200/202 with anomaly-modal body) we surface a clear
SearchProviderError so the orchestrator can fall through to the next
provider with cause attached.
Fixes#3799
Kept PATH-resolved Windows npx.cmd shims on the cmd.exe wrapper path so npm owns subprocess stdio exactly like the reporter's working cmd /c configuration.
Fixes#3794
Distinguished an absent provider (use configured preferred provider) from an explicit `--provider auto` (one-shot bypass that still respects exclusions) in executeSearch.
Fixes#3793
Restrict the supersede sweep to compactions on the path from the current leaf so a newer compaction never rewrites a sibling branch's still-current summary or drops its preserveData. Streaming load now collects the active-branch ids before eliding instead of trampling sibling compactions encountered in file order.
Refs #3789
Stream large session loads, elide superseded compaction payloads, skip synchronous rewrites when the append-only file is already current, and provide usable picker previews for developer-started forks.
Fixes#3789
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
The hashline multi-section path in `executeHashlineSingle` returned
`perFileResults: rendered.map(r => r.perFileResult)` directly. Each
per-section result had already been individually pruned by
`renderSection`, but the whole array bypassed the shared aggregate
budget added in 3987969 — a single hashline payload touching many files
with sub-32 KB snapshots each could still serialize unbounded snapshot
bytes into one session JSONL line.
Wrap the multi-section return in `pruneOversizedEditSnapshots`, which
delegates to `capPerFileSnapshots` and enforces the shared cap walking
left-to-right; early sections keep their ACP diff visualization, later
sections in a many-file batch degrade to text-only.
End-to-end regression test seeds five real on-disk files (~10 KB
combined snapshots each), runs a multi-section hashline SWAP via
`executeHashlineSingle`, and asserts the aggregate result holds the
cumulative kept snapshot bytes under `MAX_EDIT_SNAPSHOT_TEXT_CHARS`
with at least one section pruned.
The per-entry 32 KB cap let a many-file batch (apply_patch / hashline
touching N files) accumulate unbounded snapshot bytes because each
`perFileResults` entry was checked independently. 100 files × 30 KB
each kept everything (~3 MB) even though the whole array still
serializes into one session JSONL line.
`capPerFileSnapshots` walks entries left-to-right with one shared
`MAX_EDIT_SNAPSHOT_TEXT_CHARS` budget. Each per-entry payload is still
capped individually by `pruneSnapshot`; if an entry's surviving bytes
would push the running aggregate past the cap, the entry is stripped
and stamped with `snapshotsPruned: true`. Early entries keep their ACP
diff visualization; later entries in a large batch degrade to text-only
exactly like over-sized single edits.
Regression test exercises five equal-size entries that each fit the
per-entry budget but bust it cumulatively, asserting only the first two
keep snapshots and the trailing three carry the pruned marker.
When a multi-entry single-path edit prunes the first entry's snapshots
(large pre-image) and keeps a later entry's snapshots (file shrunk
between entries), the aggregator at `executeSinglePathEntries` recorded
the later entry's small `oldText`/`newText` as the whole-file
transition. ACP clients would then render a misleading partial diff
instead of degrading to text-only for the over-budget edit.
Add an explicit `snapshotsPruned` marker on `EditToolDetails` /
`EditToolPerFileResult`, set by `pruneSnapshot` whenever it strips a
payload. `executeSinglePathEntries` tracks the flag across child
results and suppresses aggregate `oldText`/`newText` (re-stamping the
marker on the aggregate) the moment any child was pruned;
`executeApplyPatchPerFile` propagates the flag onto each per-file entry.
Regression test exercises the exact scenario the reviewer raised on
#3787: replace mode where entry 1 collapses a >1 MB file to a single
line (pruned) and entry 2 trivially renames the now-tiny result; the
aggregate result now carries `snapshotsPruned: true` with both
snapshot fields omitted.
Edit-tool results carried the full pre/post file content in
`details.oldText` / `details.newText`. For large files this bloated each
per-turn JSONL line by hundreds of KB even though the snapshots are
never sent to the LLM (provider serializers send only `content`) and
only consumed by the ACP event mapper for diff visualization.
Add `pruneOversizedEditSnapshots` and apply it at every site that
constructs an `EditToolDetails` / `EditToolPerFileResult`:
`executePatchSingle`, `executeReplaceSingle`, hashline `renderSection`
(delete + update branches), and both aggregators in `edit/index.ts`.
When combined `oldText` + `newText` exceeds 32 KB the helper returns a
shallow copy with both fields omitted; smaller edits pass through
unchanged. The diff, path, firstChangedLine, op, move, and diagnostics
fields are preserved, and ACP returns no diff content for over-budget
files (the text content still flows — graceful degradation).
Fixes#3786
- Optimized TUI tool argument previews by throttling JSON re-parsing to prevent frame starvation during high-frequency streaming.
- Suppressed redundant component updates for unchanged parsed fields while maintaining raw preview integrity for bash and patch renderers.
- Added adaptive parsing logic to `ToolArgsRevealController` that distinguishes between renderers requiring continuous raw JSON streams and those consuming parsed arguments.
- Updated `EventController` to dynamically determine exposure requirements based on tool type and wire-format metadata.
- Optimized instruction sets for core agent tools including task, lsp, job, and irc.
- Standardized tool documentation structure by replacing parameter listings with structural instruction blocks.
- Mandated new communication and technical workflows for subagent results, symbol-aware code intelligence, and background task management.
- Refined messaging and coordination guidelines to prioritize inter-agent communication and direct operations.
- Added --par flag to execute benchmark runs concurrently with a default degree of 4.
- Added --service-tier flag to allow overriding the provider service tier per benchmark.
- Increased default benchmark run count from 1 to 10 to provide more robust averaging.
- Updated benchmarking logic to process requests in a concurrency-limited pool while preserving output order.
- Implemented pre-flight credential checks to prevent unnecessary worker spawning when authentication is missing.
llama.cpp /props.default_generation_settings.params.{max_tokens,n_predict} are per-request defaults the server applies when a client omits the field, not a hard model cap. Only the -1 unlimited sentinel is promoted to the runtime context window now; positive values fall back to the discovery default so client-side per-request overrides remain unconstrained.
Fixes#3781
Resolved selected-model refresh maxTokens against the effective context window, including live contextWindow overrides, so unlimited llama.cpp caps cannot exceed the configured context.
Fixes#3781
Gated discoverLlamaCppModelRuntimeMetadata's /props context fallback on the selected entry being present in /models, so refreshSelectedModelMetadata never patches a stale cached id with a different model's runtime metadata.
Fixes#3781
Mapped llama.cpp -1 generation limits from /props to the discovered runtime context window instead of the generic discovery default, including selected-model metadata refresh.
Fixes#3781
Fixed the bash interceptor rule so allowed /dev sink redirects are skipped while scanning for later real file redirects in the same command.
Fixes#3763
Fixed the bash interceptor's echo/printf redirect rule so device sinks under /dev/null, /dev/tty, /dev/stdout, and /dev/stderr remain executable while real file redirects are still blocked.
Fixes#3763
Used compact hashed isolation directory segments and the short m mount dir so long task ids are not copied into subagent working paths.
Kept worktree cleanup compatible with legacy merged task-isolation directories.
Fixes#3756
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Ollama and llama.cpp discovery now prefer runtime context settings over model training metadata, so compaction thresholds match the window local servers actually accept.
Fixes#3752
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.