urlHyperlinkAlways now short-circuits to plain text when the user has explicitly opted out via tui.hyperlinks=off, while still bypassing capability auto-detection for auto mode.
MCP OAuth fallback prompts now emit an auth-safe terminal hyperlink even when auto-detection disables normal URL hyperlinks, matching the provider login behavior while preserving the raw copy URL.\n\nFixes #2196
Same shape of bug the reviewer flagged for custom tools: forwarding
`LoadExtensionsResult` from parent to subagent reused Extension instances
whose factories closed over the parent's `ExtensionAPI` — cwd, eventBus,
and runtime all pointed at the parent. Any tool/handler/command that
referenced `api.exec()`, `api.events`, or `api.runtime` still acted on the
parent session/worktree from inside an isolated subagent.
Forward only the path list; each session rebuilds extensions through
`loadExtensions` so factories see the right `ExtensionAPI`.
- `extensibility/extensions/loader.ts`: extract `discoverExtensionPaths`
(FS scan only) from `discoverAndLoadExtensions`. The combined helper now
composes the two. New export added to the package barrel.
- `sdk.ts`:
- Add `discoverSessionExtensionPaths()` (the `disableExtensionDiscovery`-aware
path-only counterpart of `loadSessionExtensions`).
- Add `preloadedExtensionPaths?: string[]` to `CreateAgentSessionOptions`.
Three loader branches: `preloadedExtensions` (CLI same-process reuse,
still shallow-cloned), `preloadedExtensionPaths` (subagent: skip scan,
reload locally), or full discovery.
- Document `preloadedExtensions` as same-process-only; subagent
forwarding MUST use `preloadedExtensionPaths`.
- `tools/index.ts`: `ToolSession.extensionsResult` → `extensionPaths:
string[]` for the same reason.
- `task/executor.ts` and `task/index.ts`: forward `extensionPaths`. Drop
the forward for the isolated `runSubprocess` branch — worktree cwd ≠
parent cwd, so the subagent re-discovers extensions against its own
tree.
- New `test/sdk-extensions-per-session-binding.test.ts` pins the contract:
two `loadExtensions` calls on the same path with different `cwd` and
different `EventBus` instances yield distinct Extension + runtime
objects whose factories close over the per-call bindings.
- Updated `executor-pass-through` and `sdk-preloaded-extensions-isolation`
tests for the new option name and comment context.
Refs PR review on #2193
Reviewer flagged that forwarding `LoadedCustomTool[]` from a parent session
to a subagent reused tool instances whose factories had closed over the
parent's `CustomToolAPI` — `cwd`, `exec`, `pushPendingAction`, and `ui` all
pointed at the parent. In isolated tasks the tool would `exec` against the
parent worktree and queue pending actions on the parent session.
Forward only the path list; let each session rebuild tools through
`loadCustomTools` so factories see the right `CustomToolAPI`.
- `extensibility/custom-tools/loader.ts`: extract `discoverCustomToolPaths`
(FS scan only) from `discoverAndLoadCustomTools`; export
`ToolPathWithSource`. The combined helper is now `discoverCustomToolPaths`
+ `loadCustomTools`.
- `sdk.ts`: replace `preloadedCustomTools` (`LoadedCustomTool[]`) with
`preloadedCustomToolPaths` (`ToolPathWithSource[]`). The custom-tools
block runs `loadCustomTools` unconditionally; only the path scan is
skipped when the caller pre-discovered it.
- `tools/index.ts`: `ToolSession.loadedCustomTools` →
`ToolSession.customToolPaths` for the same reason.
- `task/executor.ts` and `task/index.ts`: forward `customToolPaths`.
Drop the forward for isolated subagents — the worktree shifts `cwd`, so
the subagent re-discovers tools against its own working tree.
- New `test/sdk-custom-tools-per-session-binding.test.ts` pins the contract:
two `loadCustomTools` calls on the same path with different `cwd` and
different `pushPendingAction` callbacks yield distinct tool instances
whose factories see the per-call bindings.
- Updated `executor-pass-through` and `sdk-preloaded-extensions-isolation`
tests for the new option name and added a `ToolPathWithSource` fixture.
Refs PR review on #2193
Each `runSubprocess` call re-ran `loadCapability<Rule>()`,
`loadSessionExtensions()`, and `discoverAndLoadCustomTools()` because
`ExecutorOptions` and the `createAgentSession()` call inside the executor
omitted three pass-through fields the parent had already paid for. The
already-correct paths (skills, context files, workspace tree, MCP manager)
showed the intended pattern.
- Cache `rules`, `extensionsResult`, and `loadedCustomTools` on the
parent's `ToolSession`.
- Add `rules` / `preloadedExtensions` / `preloadedCustomTools` to
`ExecutorOptions`; forward them from both `runSubprocess` call sites
in `task/index.ts` and into the executor's `createAgentSession()`.
- Add `preloadedCustomTools` to `CreateAgentSessionOptions` and skip
`discoverAndLoadCustomTools()` when it is supplied.
- Shallow-clone `extensionsResult.extensions` when reusing
`preloadedExtensions`, so the per-session autoresearch + custom-tools
inline wrappers never leak back into the caller's array.
Fixes#2190
The google-antigravity ranking strategy originally returned the most-pressured
counter as primary and the runner-up as secondary. AuthStorage compares the
secondary ranking metrics before primary, so credentials with a dangerous
bottleneck but a healthy runner-up counter could outrank credentials with
balanced headroom.
Return the bottleneck as secondary and the runner-up as primary, and update the
strategy contract test to lock the ordering invariant.
Fixes#2187
Cloud Code Assist's quota-exhaustion 429 carries 'You have exhausted your
capacity on this model. Your quota will reset after …'. The literal
'capacity' hit the MODEL_CAPACITY_EXHAUSTED branch in parseRateLimitReason
before the 'quota will reset' suffix had a say, downgrading the failure to
a 45-75s transient backoff; isUsageLimitError missed the same shape, so
agent-session.ts (line 8314) and auth-storage.ts (line 3457) never invoked
markUsageLimitReached. The session kept hammering the exhausted credential
while the retry layer bailed on the multi-hour retry-after, even when
sibling Antigravity OAuth accounts had headroom.
- Short-circuit parseRateLimitReason for 'quota will reset' / 'exhausted
your capacity' to QUOTA_EXHAUSTED before the MODEL_CAPACITY fallthrough.
- Extend USAGE_LIMIT_PATTERN so isUsageLimitError matches the Antigravity
phrasing too, enabling credential-rotation paths in agent-session,
auth-retry, stream, and the auth-gateway.
- Add antigravityRankingStrategy and register it under
DEFAULT_RANKING_STRATEGIES so multi-account selection consumes the
per-counter Antigravity usage report (already sorted ascending by
remainingFraction in fetchAntigravityUsage) instead of falling back to
round-robin.
- Cover the regression with parseRateLimitReason / isUsageLimitError tests
on the literal Antigravity message, plus a ranking-strategy contract
test asserting primary/secondary mapping and the 24h fallback window.
Fixes#2187
DeepSeek V4 reasoning models on api.deepseek.com emit no SSE bytes
while the model finishes its private chain-of-thought, which routinely
takes longer than the generic 100s first-event budget under load.
The OpenAI completions stream then aborts with
"OpenAI completions stream timed out while waiting for the first event"
and silently retries, doubling user-visible latency on almost every chat.
getOpenAICompletionsStreamIdleTimeoutFallbackMs now returns a 300s
floor for reasoning models when (provider === "deepseek") or
(baseUrl includes api.deepseek.com). first-event floors at idle, so
the watchdog gains 5 minutes — enough for reasoning warm-ups — without
changing steady-state streaming behavior. Mirrors the existing GLM
coding-plan widening.
Fixes#2177
The inline writethrough budget is 500ms (INLINE_DIAGNOSTICS_WAIT_TIMEOUT_MS);
the test published the deferred diagnostics at 900ms and asserted the inline
call returned in <800ms, leaving only ~300ms of headroom over the budget. CI
jitter pushed elapsed to 844ms (still correct deferral, just slow), failing the
over-tight bound. Publish at 2000ms and assert <1500ms so the deferral margin
is wide while still proving inline does not block on the slow publish.
Completes the injectable-fetch transport wiring (15.10.8) that the feature
left half-done, fixing the deterministic CI test failures:
- compaction.compact() rebuilt summaryOptions field-by-field but dropped
`fetch`, so the injected transport never reached
requestOpenAiRemoteCompaction / generateSummary's remote path. Thread it.
- Read-tool URL pipeline had no fetch seam: renderHtmlToText gained a
fetchOverride param but renderUrl/ToolSession never carried one, so the
jina/parallel reader backends always used global fetch. Add
ToolSession.fetch -> renderUrl -> renderHtmlToText (defaults to global).
- searchWithParallel mirrored extractWithParallel but missed the fetch
option; add it.
- Repair tests whose deleted hookFetch interceptors were never replaced
with a FetchImpl seam (fetch-kagi-toggle, web-search-parallel,
issue-970 discovery).
- Update issue-1746 POSIX case to the #2154 preserved-scrollback contract:
unknown-viewport streaming deferral is now platform-independent.
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
- Added OpenRouter detection in openai-completions parameter construction and skipped max-token emission for non-Kimi OpenRouter routes.
- Preserved existing 64k-cap clamping behavior for non-OpenRouter requests and model-specific limits.
- Expanded max-output-token tests to cover OpenRouter omission behavior and Kimi-over-OpenRouter token handling.
- Parsed a trailing `@slug` to set OpenRouter `provider.only` or Vercel Gateway routing per invocation.
- Resolved via `parseModelPattern`, composing with thinking levels and round-tripping through selectors.
- Split only when the base resolves to an aggregator, so ids containing `@` stay intact.
Stopped foreground streaming from treating an unknown native viewport as permission for destructive history rebuilds, so offscreen growth stays incremental after the viewport fills.\n\nFixes #2154
- Added `AssistantMessage.upstreamProvider` and emitted it as `pi.gen_ai.response.upstream_provider` telemetry.
- Clamped OpenAI-family output caps to `OPENAI_MAX_OUTPUT_TOKENS` and model limits to avoid oversized token requests.
- Stopped aborting the shared per-request AbortController, which latched and blocked stream reopens.
- Dropped the half-built degenerate tool call and replayed the request, bounded by a retry limit of 2.
- Surfaced the original error once retries are exhausted, without polluting the message.
- Dropped blank lines, capped at 8 lines, and width-truncated each line so a proxy 502's HTML body can't flood the transcript.
- Mirrored the pinned error banner's preview behavior; full text still kept in the persisted session.
- Added a `bash.enabled` boolean setting with a default of `true` in the settings schema.
- Updated tool generation to include the `bash` model tool only when `bash.enabled` is enabled.
- Extended createTools tests to assert `bash` is omitted when disabled and omitted from requested disabled tool lists.
Interrupting the model during its visible output produced an assistant turn
whose thinking block had already finished streaming (fully signed) followed
by partial text. `transformMessages` stripped the signature from every
thinking block of an aborted/errored turn, so the valid signature was
replayed empty (`signature: ""`) and signature-enforcing Anthropic
rejected the next request with HTTP 400 "Invalid `signature` in
`thinking` block" — including when routed through an LLM gateway baseUrl.
Anthropic delivers a block's signature at `content_block_stop` before the
next block's `content_block_start`, so a thinking block followed by
text/tool_use is necessarily complete and valid. Only the single mid-stream
block at the abort point can hold a partial signature. Strip just that
final block now; completed thinking blocks keep their replayable
signatures. Abandoned end_turn+tool_use turns are unchanged (all their
signatures remain stripped).
Co-authored-by: DeprecatedLuke <16418011+DeprecatedLuke@users.noreply.github.com>
Fixes#2144
The auth-gateway client now sends `${provider}/${id}` (008981112) and the
gateway registry keys on the qualified id first. Update the stale request-shape
assertion that still expected the bare `claude-sonnet-4-5`.
Pins issue #2123: OAuth requests to adaptive-thinking Claude Opus
(4.6+) attach a 'context_management.edits[clear_thinking_20251015]'
block alongside 'thinking'. When the eager-todo prelude (and other
plan-mode paths that force tool_choice to 'tool'/'any' on the first
user turn) routes through 'disableThinkingIfToolChoiceForced',
'params.thinking' is stripped; 15.10.4 left the orphan
'params.context_management' behind, producing a 400 with
'clear_thinking_20251015 strategy requires thinking to be enabled
or adaptive'. The 15.10.5 fix in 0890b2be6 already drops both
fields together — the new repro test asserts this invariant for
both the named-tool and 'any' forced shapes so the strategy can
never outlive its enabling thinking payload again.
Fixes#2123
MCPManager.connectServers used to fall through to an unbounded
Promise.allSettled over every still-pending server without cached tools,
so a single MCP server stuck waiting on the per-request MCP timeout
(OMP_MCP_TIMEOUT_MS, default 30 000 ms) gated the entire UI ready
signal — exactly the 30.282 s stall the reporter observed against
sbox-superdocs in #2100.
Drop the fallback wait. Pending-without-cache servers are left in flight
and their tools surface via the existing background #onToolsChanged ->
refreshMCPTools path the moment the connect completes; failures continue
to log through the background catch handler (gated on
allowBackgroundLogging) so users still see which server failed.
Adds a regression test that spawns an unresponsive stdio MCP fixture
and asserts connectServers returns inside the 250 ms STARTUP_TIMEOUT_MS
window (padded for CI jitter). The same test times out at 15 s without
the patch.
Fixes#2100