Waited briefly for process exit publication after clean stdout EOF so the process handler preserves the real exit code and stderr, while genuine reader errors still tear down immediately.
Cleared only the matching initialization failure for explicit reloads and added regressions for quick exits, reader errors, ordinary backoff, and immediate reload retries.
Fixes#7041
(cherry picked from commit a76522b759f14421202d4cc437ec611b78be20d1)
buildSystemPrompt hand-rolled systemPrompt?.map(...) and devin's request
builder called (context.systemPrompt ?? []).join(...); both crash when
Context.systemPrompt is a bare string, as legacy @earendil-works/pi-ai
extensions remapped onto the fork pass it. The failure surfaced as
stopReason "error" with "systemPrompt?.map is not a function".
Route both through the existing normalizeSystemPrompts() helper, which
already accepts readonly string[] | string, matching the other providers.
Fixes#7037
(cherry picked from commit d681de5daa7e3316f118e6fb4cf8805ec5318c32)
parseConnectEndStream repeats the default status phrase in the message
body, so the trailer reads "resource_exhausted: resource exhausted".
Using a global strip removes both occurrences, so the leftover
"exhausted" no longer trips the generic quota branch and reintroduces
the 30-minute credential block for an otherwise bare status.
Fixes#7032
(cherry picked from commit 3d46a042c649e3892d9c68abbc3789a7719f0667)
Stripped the resource_exhausted status token before classifying the
remaining provider message. Bare or opaque status errors still use the
short model-capacity backoff, while explicit quota, rate-limit, capacity,
or server details remain authoritative.
Added regression coverage for a Connect resource_exhausted trailer with
an explicit quota-exceeded body.
Fixes#7032
(cherry picked from commit ee04c065692e34ccabe1a07b45b5f5987297ff84)
Connect/gRPC end-streams carry the status name `resource_exhausted`
(underscore), but parseRateLimitReason only matched the space phrase
"resource exhausted" in its MODEL_CAPACITY branch. The underscore form
fell through to the generic includes("exhausted") catch-all and was
classified QUOTA_EXHAUSTED, producing a 30-min credential block and the
retry.maxDelayMs fail-fast at the session layer.
Match both forms via /resource.?exhausted/i, consistent with the
existing resource.?exhausted clause in USAGE_LIMIT_PATTERN, which is left
untouched so stream/session credential rotation is preserved.
Fixes#7032
(cherry picked from commit bc18cbb5b9bcf5c2bf41dc494581ae31061fcfcb)
The models config resource regression started two cold child processes inside one five-second test. Under parallel CI chunk load, the second child could still be waiting for pipe drain or process exit after the validator had completed.
Keep the process boundary as the owner of the lifecycle measurement and run the missing and custom phases in one child. The missing phase closes its storage and registry before the baseline snapshot, while the custom phase proves schema identity and retention without changing config loading or validator cleanup semantics.
Signed-off-by: Christian Stewart <christian@aperture.us>
(cherry picked from commit 81fa98491544c7fbcbf075922889e0a9e1a0a3b4)
Spawn-based lazy-loading tests assert exitCode===0 on a child process but
set no per-test timeout, so bun's 5s default kills the child under CPU
contention and the assertion reports a dead child rather than a regression.
Give each spawn test an explicit 60s per-test timeout, matching the
existing precedent in auth-gateway-anthropic-caching.test.ts.
Fixes#7018
(cherry picked from commit 00e5ec8855eb8ba31c1bdf571bd7c16f2405d679)
The soft request budget resolved to `SOFT_REQUEST_BUDGET[agent.name] ??
configured`, so the bundled entries for scout and sonic replaced the
configured value outright. Lowering `task.softRequestBudget` to tighten
the guard therefore did nothing for exactly the two agents that spawn
most often: a scout kept its 100-request budget no matter how small the
user set the knob. Only 0 (disable) and raising the value for
non-bundled agents had any effect.
Treat both numbers as upper bounds and take the smaller one. The bundled
entries stay ceilings, so a runaway scout is still stopped at 100 by
default and existing behavior is unchanged for anyone who has not
lowered the setting; a configured 0 still disables the guard entirely.
Resolution moves into `resolveSoftRequestBudget`, which also normalizes
negative and fractional inputs, so the rule is testable without standing
up a subprocess run.
This composes with `task.maxEffort` on a separate axis: effort caps how
hard each request thinks, this caps how many requests a run may spend.
(cherry picked from commit f0db29f8f725f11390b64ca9342300c482ff5c5d)
Catalog summaries of mounted xd:// devices are inlined verbatim into the
system prompt. External devices (MCP servers, plugins) supply that text, and
it was bounded only by character count: a summary of multi-byte script passed
roughly three times the intended budget, and control characters survived into
the prompt where they can forge structure.
Summaries now go through a single sanitize-and-bound step that strips C0/C1
control characters and bounds the result in UTF-8 bytes via the central
truncateHeadBytes helper, so a cut lands on a code point boundary and never
renders a partial code point. The built-in/external distinction is derived
once per entry, and that same boolean both selects the description cap and is
exposed as `dynamic`, so the cap and the flag cannot disagree. The prompt uses
the flag to state that dynamic summaries are untrusted metadata, and the mount
notice says the same for newly appeared devices.
(cherry picked from commit 5989da6235d820bc687779a791e655e6f1b2df0f)
GPT-5.6 Responses-Lite models receive tool_choice "auto" (the forced
hosted choice is invalid under the lite shape, #5771/#5772), so the model
may answer without invoking the hosted web_search tool. The codex search
parser accepted any non-empty answer, returning a stale completion with
zero sources as a successful search.
callCodexSearch now tracks response.web_search_call.* events (and
web_search_call output items) and throws CodexNoWebSearchError when none
occurred. The candidate chain treats that error as retryable, advancing
default lite models to a non-lite model that forces web_search, and
surfaces a clear failure when the model was explicitly configured.
Fixes#6988
(cherry picked from commit a276cd0b3df1d0d041faf0a63fabcbb884e36a91)