Waited briefly for process exit publication after clean stdout EOF so the process handler preserves the real exit code and stderr, while genuine reader errors still tear down immediately.
Cleared only the matching initialization failure for explicit reloads and added regressions for quick exits, reader errors, ordinary backoff, and immediate reload retries.
Fixes#7041
(cherry picked from commit a76522b759f14421202d4cc437ec611b78be20d1)
The soft request budget resolved to `SOFT_REQUEST_BUDGET[agent.name] ??
configured`, so the bundled entries for scout and sonic replaced the
configured value outright. Lowering `task.softRequestBudget` to tighten
the guard therefore did nothing for exactly the two agents that spawn
most often: a scout kept its 100-request budget no matter how small the
user set the knob. Only 0 (disable) and raising the value for
non-bundled agents had any effect.
Treat both numbers as upper bounds and take the smaller one. The bundled
entries stay ceilings, so a runaway scout is still stopped at 100 by
default and existing behavior is unchanged for anyone who has not
lowered the setting; a configured 0 still disables the guard entirely.
Resolution moves into `resolveSoftRequestBudget`, which also normalizes
negative and fractional inputs, so the rule is testable without standing
up a subprocess run.
This composes with `task.maxEffort` on a separate axis: effort caps how
hard each request thinks, this caps how many requests a run may spend.
(cherry picked from commit f0db29f8f725f11390b64ca9342300c482ff5c5d)
Catalog summaries of mounted xd:// devices are inlined verbatim into the
system prompt. External devices (MCP servers, plugins) supply that text, and
it was bounded only by character count: a summary of multi-byte script passed
roughly three times the intended budget, and control characters survived into
the prompt where they can forge structure.
Summaries now go through a single sanitize-and-bound step that strips C0/C1
control characters and bounds the result in UTF-8 bytes via the central
truncateHeadBytes helper, so a cut lands on a code point boundary and never
renders a partial code point. The built-in/external distinction is derived
once per entry, and that same boolean both selects the description cap and is
exposed as `dynamic`, so the cap and the flag cannot disagree. The prompt uses
the flag to state that dynamic summaries are untrusted metadata, and the mount
notice says the same for newly appeared devices.
(cherry picked from commit 5989da6235d820bc687779a791e655e6f1b2df0f)
GPT-5.6 Responses-Lite models receive tool_choice "auto" (the forced
hosted choice is invalid under the lite shape, #5771/#5772), so the model
may answer without invoking the hosted web_search tool. The codex search
parser accepted any non-empty answer, returning a stale completion with
zero sources as a successful search.
callCodexSearch now tracks response.web_search_call.* events (and
web_search_call output items) and throws CodexNoWebSearchError when none
occurred. The candidate chain treats that error as retryable, advancing
default lite models to a non-lite model that forces web_search, and
surfaces a clear failure when the model was explicitly configured.
Fixes#6988
(cherry picked from commit a276cd0b3df1d0d041faf0a63fabcbb884e36a91)
- reuse the centralized replay-unsafe predicate in Fast fallback
- exercise visible-text and replay-safe Fireworks branches
(cherry picked from commit f596392b891bb8ec82171a25c6651ed3cc8e0d61)
Snapshot this.ctx.sessionManager.getSessionId() for the tan clone's local://
mapping instead of session.sessionId. The two diverge after /fresh or a
provider session override, and the Windows short-root fallback keys
%TEMP%/omp-local/<id> off the session-manager id used by the parent's
large-paste writes and '/data/workspaces/can1357__oh-my-pi__6971/.omp-session/2026-07-29T06-08-45-283Z_019fac7d-5ee3-7000-a7aa-16fe9394fdc9/local' reads, so the mismatched id left attachments
unreachable.
Diverge the mocked session id from the manager id in the regression test so
it pins the session-manager id.
Fixes#6971
(cherry picked from commit 1efcd22326d76fdb8b50c0977a836c892e80ab76)
Keep subagent localProtocolOptions on their ToolSession instead of installing
them as the process-global LocalProtocolHandler override. No-context URL
consumers therefore retain the active top-level session's mapping while tan
and task subagents continue to resolve through their caller context.
Add SDK regression coverage proving subagent creation preserves an existing
global mapping.
Fixes#6971
(cherry picked from commit a02eef174b03036dc960c842a0901d22333ad9cd)
Capture the parent artifacts directory and session ID when /tan dispatches
instead of resolving them through the mutable interactive SessionManager.
This keeps background tan '/data/workspaces/can1357__oh-my-pi__6971/.omp-session/2026-07-29T06-08-45-283Z_019fac7d-5ee3-7000-a7aa-16fe9394fdc9/local' reads pinned to the dispatching transcript
after the user switches or resumes another session.
Extend the regression test to switch the mocked interactive session before
the background job starts and assert the original local mapping is retained.
Fixes#6971
(cherry picked from commit e05db428eabd5087e4b2b4a462f47927c0628e72)
TanCommandController.start nests the tan clone at
<parent-artifacts>/Tan-<id>.jsonl, so the clone's session manager derived
its own artifacts dir and hence local root <parent-artifacts>/Tan-<id>/local.
Its sdk.createAgentSession call omitted localProtocolOptions, unlike the
task-subagent path which inherits the parent's mapping, so parent-session
'/data/workspaces/can1357__oh-my-pi__6971/.omp-session/2026-07-29T06-08-45-283Z_019fac7d-5ee3-7000-a7aa-16fe9394fdc9/local' attachments (pasted files, generated references) were unreadable.
Thread the parent session manager's localProtocolOptions into the tan clone
so local:// resolves against <parent-artifacts>/local.
Fixes#6971
(cherry picked from commit 1ded46e182fc24f9f57d8e9a907aaad58f790783)
Root cause of the reported "edit tool silently reformats the whole
file" corruption: fs/write_text_file has no verbatim guarantee. When
an ACP client (e.g. Zed with format_on_save: on) reformats a buffer
on save, routeWriteThroughBridge reported the pre-write content as
successfully written, and Patcher.commit keyed the returned snapshot
tag on that same pre-write text instead of what actually landed on
disk. The next edit anchored on that tag then resolved hunks against
a baseline the file had already drifted away from, which is what
produced whole-file "corruption" from single-line hunks -- reproduced
live in this session against real Swift/JSON/TypeScript files with
Zed as the ACP client.
- routeWriteThroughBridge reads the file back after the bridge write
and returns the verified content plus a drift flag (best-effort:
ACP defines no ordering between the client acking the write and its
own async format-on-save settling, so this degrades gracefully to
the old stale-tag-on-next-read failure mode, never to corruption).
- HashlineFilesystem.writeText propagates that verified content in
view-space (the same space readText returns -- e.g. a notebook's
editable cell text, not its raw JSON), not storage-space, so tag
validation on the next edit compares like with like.
- Patcher.commit keys fileHash/header/snapshot on the verified
post-write content (normalized, so BOM/line-ending restoration never
produces a false "drift") when it diverges from what was sent, and
appends a warning naming the drift -- but deliberately leaves the
returned `after` (and therefore the model-visible diff) scoped to
the intended hunk. Diffing against the full drifted file would
balloon the tool response to span every reformatted line (measured
~6.8x inflation on a 245-line file with one touched line); the
warning is the correct O(1) channel for "your editor reformatted
this," not an O(file-size) diff.
- write.ts keys its own snapshot header on the verified bridge content
too (no diff-size concern there since write always replaces the
whole file).
Caught via code review (dispatched against the first pass of this
fix): a naive "just use the verified content everywhere" fix broke
.ipynb editing outright (write-space vs read-space content mismatch,
tag invalid on every notebook edit) and would have inflated every
drifted edit response by ~6.8x. Both are now covered by regression
tests that fail against the pre-fix code and pass against this one.
(cherry picked from commit 35ab80e43be5800b2f48728e4400eb9fd7f7f7d2)
The previous commits patched each rebuild path individually to avoid feeding a
modifyModels hook its own output. That left the invariant implicit and the
provider-scoped path applying only a subset of hooks, which is wrong for a hook
that inspects or suppresses another provider's models.
Keep #unprojectedModels as the canonical pre-projection catalog and derive
#models from it at every mutation point, so projections are always a pure
function of the unprojected base:
- #composeUnprojectedStaticModels builds the catalog; #composeStaticModels
projects it. A scoped lookup with modifiers registered composes and projects
the whole catalog before narrowing, matching getAll() followed by a filter.
Providers without modifiers keep the cheap filtered path.
- Discovery completion, registerProvider, and runtime transport overrides
update the unprojected snapshot and reproject, instead of mutating an
already-projected array.
- Runtime metadata patches apply to the unprojected model, then reproject, so
a later registration cannot discard them.
- Provider lookup snapshots are invalidated wherever the projection changes.
Hooks no longer take a providerFilter: a modifier is a whole-catalog transform
and every rebuild now runs the full ordered set exactly once.
(cherry picked from commit e5d2e9eac7c371cc196e9b362f77d3a5d7bdf507)
registerProvider composed nextModels from the already-projected #models,
stripping only the incoming provider, then reran every stored modifier over
it. Loaders drain registrations one at a time, so the previously registered
provider's projection was fed back into its own hook — an append-style hook
compounded on each subsequent registration.
Apply only the incoming provider's hook. Every other provider's projection is
already present exactly once, and full rebuilds still go through
#composeStaticModels.
(cherry picked from commit b16db642b08cc223b062c0c556f00d46fda2520a)
Review follow-up on two defects in the original change:
- The throwing-hook fallback wrote to #lastDiscoveryWarnings, which is only
ever read to dedup a logger.warn inside #warnProviderDiscoveryFailure. No
log line was emitted, so a broken extension degraded invisibly, and the
shared key could mask a later discovery failure for the same provider. Log
via logger.warn with its own dedup map.
- #refreshRuntimeDiscoveries starts from the already-projected #models, and
the overlay merge only replaces matching provider+id pairs, so a hook's
projection-only entries survived and were fed back into it. An append-style
hook duplicated its output on every refresh. Drop each modifier provider
before the merge so it re-seeds from the unprojected overlays.
(cherry picked from commit b6f841e080d4882a08b8d713de009461b6acc6fe)
`registerProvider` applies `oauth.modifyModels` once and assigns the result
straight to `#models`, but only the pre-projection definitions are persisted
in `#runtimeModelOverlays`. Any subsequent static reload rebuilds `#models`
from those overlays and silently drops the projection.
The model selector reloads on every open (`refresh("offline")`), so an
extension provider that projects a credential-aware catalog shows its
correct models everywhere except the picker — the one place users look.
`refreshProvider()` and online discovery completion had the same hole.
Persist the hook per provider and re-apply it wherever `#models` is
recomposed, honouring the `providerFilter` used by scoped lookups. A hook
that throws now degrades to that provider's unprojected catalog instead of
failing the whole composition, so one broken extension cannot empty the
registry.
(cherry picked from commit 33b7c72f225b4253b68bb71ecb3a9186151b18b3)
The previous reset dropped #pendingXdevMountDelta alongside the announced
baseline. Unlike /new and different-session switchSession, branch() does
not rebuild the base system prompt afterward, and because the device is
already in mountedNames no later refresh re-queues an add delta. Dropping
the undelivered delta therefore left the branched transcript unaware of a
still-mounted discoverable device.
Only the announced baseline is reset now; pending adds (still-live mounts
awaiting delivery) survive and announce on the next prompt in the new
transcript. Redundant on /new (the rebuilt prompt lists them too) but
harmless, and correct for branch.
Fixes#6921
(cherry picked from commit 5866440f27c15a320657e25fbfa0d2d65aad1fa2)
The announced-mount baseline persisted across /new, switchSession, and
branch, which replace agent.state.messages but only clear session-scoped
tool state. A device announced in the old transcript stayed in the cache,
so reconnecting it into the fresh history was filtered as already known
and never announced, leaving the new conversation unaware of the device.
Reset the announced baseline (and any undelivered pending delta) from
#clearSessionScopedToolState, so the next notice re-seeds from the new
transcript and a reconnecting device announces again.
Fixes#6921
(cherry picked from commit d06dde02b9de4aacacd7aea5ee51edc7e524e3fb)
Replayed the stable added and removed inventory sections from legacy
xdev-mount-notice content when structured details are absent. This keeps
the first post-upgrade resume from re-announcing devices that persisted
history already introduced.
Covered both structured and legacy resume histories, including removed
devices and inline docs that must not be interpreted as inventory.
Fixes#6921
(cherry picked from commit 634a4c2de75f99e219f408c56bed83c04fe1290a)