Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.
`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
Codex found the shared attribution predicate recognising only tool calls,
text and signed thinking. A native image response often arrives with no
text and no tool call at all, so an image-only turn was read as producing
nothing: attribution stayed on whichever model spoke before it, and the
empty-stop rule could classify a successful generation as empty.
Everything the assistant can emit now counts except two: unsigned thinking,
which is not provider-authenticated and was already excluded, and
Anthropic's `fallback` marker, which records that a request was routed
elsewhere rather than carrying output. Redacted thinking and server-tool
blocks are real work by the same argument as the image.
The `toolUse` arm keeps its stricter rule — an orphaned toolUse stop needs
a tool_use block to anchor a later tool_result, and an image cannot.
Also switches the new test to the namespace import AGENTS.md requires for
node builtins.
Review found the branch does not compile. `bun check` runs biome before
the per-workspace type check and joins them with `&&`, so a pre-existing
format error on `main` short-circuited the run: `tsgo --noEmit` never
executed, and `bun test` type-strips, so the suite stayed green over eight
type errors. `ServingModel` was used in `turn-recovery.ts` without being
imported, and three `subscribe` closures in the retry-fallback suite lost
the outer narrowing of `session`. Verified now against the workspace check
directly rather than the aggregate.
Codex also found `#fallbackRouted` outliving its session. It said how the
CURRENT model was reached, but nothing reset it when a transcript was
switched or resumed, so a freshly loaded session with no served attribution
yet described its own model with the previous session's routing. It is now
anchored on the session id exactly like the attribution beside it — the two
facts are earned together and expire together — and a session switch
reports the current model without claiming to know how it got there.
The replay-unsafe host stub gained `getSessionId`, which the real
`SessionManager` has always had and the anchoring now calls.
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.
Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.
The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.
`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
Two cleanups from review.
The offline walk matched a served turn to its `model_change` by also
accepting a `:level` thinking suffix, but no writer produces one: every
`appendModelChange` call site records a bare `provider/id` or a
`formatModelStringWithRouting` selector, and only the latter decorates —
with `@upstream`, never a level. Speculative matching in the part of the
change that already reconstructs the most from raw strings. Removed, with
its test narrowed to the format writers actually emit.
Attribution was also retained for the life of the recovery component, so
switching sessions in place could report a model from a different
transcript by name. It is now tagged with the session file it was earned
in and ignored when that no longer matches — the same self-invalidating
anchor `#ensurePersistedMessageKeys` uses, so no mutation call site has
to remember to clear anything. Anchoring on the file rather than the leaf
keeps attribution across compaction and appends, which stay within the
transcript that earned it.
Review of the previous commit found the same defect in four more places,
each a variant of one mistake: the unproven-model gate was expressed as
"a retry fallback is armed and has not served", so it only protected the
one swap path that arms a chain.
- The cooldown-expiry restore to the primary deletes the fallback record
before swapping, so the gate vanished with it and an unproven, just-
restored primary was credited immediately.
- A session's first-ever fallback arm constructs the record only after
the swap, so there was nothing to mark and the guard was a no-op.
- The Fireworks Fast degrade swaps models without arming a chain at all,
so a live row showed the degraded base model as the configured primary.
- The fresh-versus-resumed heuristic read the message list, which
compaction collapses, so a resumed-and-compacted session seeded its
startup candidate as already proven.
Attribution now names the last model that settled a turn here, full stop.
A switch of any kind — into a candidate, back to a restored primary, or a
capability degrade — moves it only once the new model answers. Before
anything has served there is no earlier work to miscredit, so the
configured model is both the only answer and a safe one.
Fallback routing is tracked as its own flag rather than inferred from the
chain record, because the degrade path routes without a chain, and it is
set before each swap so the synchronous `model_changed` fan-out cannot
observe a model mid-relabel. The resumed-session heuristic is gone: it
was only needed to decide whether an unserved candidate could be trusted,
and nothing is trusted before it serves.
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.
Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.
Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.
Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.
One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.
Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.
A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.
`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
- 28 symbols across discovery, mcp header policy, agent-hub projection and
rendering, the agent registry, shell tokenizing and changelog comparison
were exported but referenced only inside their own module; they are now
module-private, shrinking the deep-import surface.
- Kept AGENT_PLUGIN_MANIFEST_SCHEMA, AGENT_PLUGIN_MCP_SCHEMA,
parseAgentPluginManifest, clearAgentPluginRootCache and mergeMCPHeaders
exported: each is a seam for tests that defend real parsing or header
precedence behavior.
- Nothing reachable from an explicit exports entry or public barrel changed.
- src/lsp/index.ts is the explicit ./lsp package entry, yet held 2821 lines of
warmup, config caching, diagnostics, external build-command workspace
diagnostics, the writethrough batching subsystem and the LspTool class.
- Those are now servers, diagnostics, workspace-diagnostics, writethrough and
tool modules; index.ts is 22 lines and re-exports the same public surface.
- configCache and writethroughBatches remain single instances and every tuned
diagnostics timing constant moved verbatim.
- Moved the module-level machinery that sat in front of the ModelRegistry
class into model-config-values, model-patch, custom-models and
model-provider-discovery; model-registry.ts drops 646 lines.
- commandValueCache and its negative-cache TTL stay single instances, so the
execSync storm the cache exists to prevent cannot return.
- The setCodexAttestationProvider import-time side effect stays in
model-registry.ts. The class itself was left alone: its private state is
shared across the methods, so splitting it is not a straight move.
- theme.ts mixed symbol presets, JSON schema, color math, the Theme class,
loading, global state, appearance handling and TUI adapters in 3171 lines.
- Symbols, schema, color, theme-class, loader and tui-adapters are now
siblings; theme.ts keeps global state, the watcher, appearance handling and
HTML export at 745 lines, with all 44 exports intact.
- Left appearance and export-colors in place: both read private mutable
auto-theme state, so extracting them would have required new exported
internals or DI rather than a straight move.
- Separated deterministic replacement generation, placeholder derivation,
placeholder-range scanning and message-tree transforms out of the 2647-line
module; obfuscator.ts now holds the types and SecretObfuscator.
- ephemeralPlaceholderKey stays a single instance and both global regexes stay
beside the code that resets their lastIndex, so placeholder stability and
the security argument in the moved comments are preserved verbatim.
- Repointed every importer at the real modules rather than leaving a re-export
shim; the public ./secrets barrel exports the same 15 names as before.
- gh.ts held wire types, search, Actions run-watch, PR checkout/push/create,
PR diff parsing and view fetch/format in 3958 lines.
- Split into gh-types, gh-search, gh-run-watch, gh-pr-checkout, gh-pr-diff,
gh-view and a gh-common module holding the shared primitives and the single
process-lifetime default-repo memo pair; gh.ts is now 246 lines.
- All 22 exports stay on gh.ts because tools/index.ts star-exports ./gh, so
the issue:// and pr:// protocol handlers needed no edits.
- ReadTool mixed plain-file reading with archive, sqlite, pdf-image, summary,
selector, formatting and renderer concerns in one 3763-line module.
- Each now owns a sibling module; read.ts drops to 2020 lines and keeps its
public exports, including the readToolRenderer re-export required because
tools/index.ts star-exports ./read through the explicit ./tools entry.
- The pdfImageExtractions map and summaryParseCaches WeakMap stay single
instances; execute() was deliberately left intact.
- builtin-registry.ts held a 2341-line array of every command spec with its
handler inlined, plus the autocomplete builders and the TUI dispatcher.
- Specs moved verbatim into modes, collaboration, session, lifecycle,
marketplace and control modules; completions moved to builtin-completions.
- builtin-registry.ts is now a 174-line composition/dispatch module and keeps
all 16 exports. Command order is unchanged, which matters because it drives
autocomplete ordering and BUILTIN_SLASH_COMMAND_DEFS.
- Python, Ruby and Julia each carried their own copy of the same session
maps, acquire/reset/replace/dispose lifecycle and executeOnSession; one
generic registry now owns it, parameterized by a per-language descriptor.
- Julia additionally re-implemented seven executor-base helpers locally; those
copies are gone and executor-base gained sparse managed-env and timeout
resolver hooks so Julia's differing behavior survives unchanged.
- All twelve exported entry points keep their names and signatures.
- zune-jpeg 0.5.15 (image 0.25's JPEG decoder) cannot compile with its
non-default log feature off: zune-core's no-log warn! stub is not
expression-safe. A feature-activation-only workspace dep on
zune-jpeg { features = ["log"] } fixes the cold build; log stays 0.4.33.
- model-registry-default-config's local ModelSnapshot type gains the optional
streamIdleTimeoutMs the Bedrock watchdog compat now emits.
- A provider-supplied retry-after now bypasses the transient rate/concurrency
heuristic window instead of being overridden by it (regression from the
subscription-cap retry change).
- Updated event-controller/ui-helpers test doubles for provenance-gated
renderer selection (hasBuiltInTool), aggregated retryErrors on
auto_retry_end, and Bedrock override compat gaining streamIdleTimeoutMs.
Preserved zero-width readiness and wait matches across the daemon wire protocol, and isolated malformed completion events from unrelated pending RPCs.
Fixes#7908
Added the upstream unregisterProvider lifecycle to queued and initialized extension runtimes. Provider removal now clears runtime model/auth state before replacement, while failed factories restore the prior registration queue.
Fixes#7914
Published per-cwd discovery snapshots to existing task tools and refreshed them from TUI, ACP, and Agent Control Center reload paths.
Added regressions for existing and future task tools across TUI and ACP reloads.
Fixes#7940
A cooldown-expiry model revert runs at a turn boundary. The user-prompt
path reverts then re-checks accumulated context against the restored
model via runPrePromptCompactionIfNeeded, but the automatic
agent.continue() path (#scheduleAgentContinue) reverted and issued the
next request with no such check. When a transient failure had fallen
back to a larger-window model and the conversation then grew past the
original model's window, restoring the primary once its cooldown expired
sent a predictably oversized request to the smaller model.
maybeRestoreRetryFallbackPrimary now reports whether it actually
switched, and the auto-continue path runs the same post-revert
context-fit maintenance (compaction/promotion) the prompt path already
runs, but only when a revert occurred.
Fixes#7952
- Runtime error now just says to download vscode-js-debug from its GitHub repo;
tarball recipe, extract path, env var, and Mason detail stay in docs/tools/debug.md.
The handoff catch in session-handoff.ts and the /handoff handler in
command-controller.ts mapped any error named AbortError to "Handoff
cancelled" regardless of whether the handoff signal was actually
aborted. Providers throw name-AbortError errors on non-user conditions
(stalls, idle timeouts, nested resolution failures), so a genuine
generation failure surfaced as a user cancellation and hid the cause.
Only report "Handoff cancelled" when handoffSignal.aborted is set;
re-throw the real error otherwise. The controller now trusts the
normalized "Handoff cancelled" message and drops its own AbortError
check so re-thrown provider failures render as "Handoff failed: ...".
Fixes#7903