- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
Allowed advisor streams to treat Google STOP responses without visible content as successful silence while preserving the default retry behavior for interactive agents.
Added provider and advisor-path regressions covering retry counts, system instructions, and the advise declaration.
Fixes#8223
Every session dispose broadcast the shutdown abort reason, so an explicit
hard kill of a subagent (release with tombstone, then live.dispose()) tagged
its nested children as shutdown and rediscovered them as parked instead of
terminal.
- Gate ASYNC_JOB_MANAGER_SHUTDOWN_REASON on #ownedAsyncJobManager so only the
top-level owning session's dispose (genuine process shutdown) uses it.
- Subagent disposes propagate a generic cancellation, keeping nested children
terminal.
- Cover the subagent generic-cancel path alongside the owning-session shutdown.
Fixes#8216
Session teardown pre-cancels owner jobs via #cancelOwnAsyncJobs before
manager.dispose(), so the shutdown abort reason must ride along that
cancelAll or an owned subagent job sees a generic caller signal and is
tombstoned instead of parked.
- Forward an abort reason through AsyncJobManager.cancelAll.
- Pass ASYNC_JOB_MANAGER_SHUTDOWN_REASON from #disposeOwnedAsyncJobs.
- Cover owned-job shutdown tagging through a real AgentSession dispose.
Fixes#8216
Why:
The module-global array memo strongly retains the most recently converted
transcript and output after its session is disposed.
Changes:
- Store exact-repeat and append-growth state in a WeakMap per input array.
- Preserve generation invalidation and per-message weak caching.
Evidence:
- Seven disposal runs collected both arrays after 50 forced-GC passes and
reduced median heap delta by 87.54%.
Refs #8119
Why:
SessionManager.open() parses the complete journal for its header and then
setSessionFile() parses the same journal again.
Changes:
- Reuse the entries already loaded by open() through a private setup path.
- Keep the public setSessionFile() contract unchanged.
Evidence:
- A 19.98 MB, 12,000-entry journal improved from 48.011 ms to
25.406 ms median across 15 runs with identical restored state.
Refs #8117
- seal() now runs before the final dispose close() and bumps the disk
epoch: work an event handler enqueues while dispose awaits the closing
tail is superseded, and an already-running fenced or authoritative
rewrite fails its commit guard at the rename fence instead of
publishing over a file a revival reopened.
- The authoritative repair path resets the disk tail itself, escaping the
close() serialization, and would atomically publish the emptied entry
list; it now no-ops once sealed and commit guards also check the seal.
- Execution-time gates cover the queued title persist and fenced rewrite
callbacks; setSessionName/appendCustomEntry attempted by a handler that
outlives dispose are dropped and covered by the seal regression, and a
failed atomic batch across the seal can no longer truncate the file.
- AgentLifecycleManager.park() resolves as soon as dispose() returns, so
ensureLive() may reopen the same JSONL through a new manager while a
timed-out event handler still holds the old one; a late append reopened
a second writer on that file and the deferred finalize then closed it.
- A post-release rewrite was worse: it persisted the emptied entry list,
truncating the transcript on disk.
- releaseRetainedEntries() now seals the manager: appends, title changes,
and rewrites become dropped no-ops and the append writer is closed, so
the deferred dispose pass is in-memory only and can never touch the file.
- File-backed regression: dispose on the drain deadline, revive the JSONL
immediately, unpark the late handler, and prove the file is byte-stable
and the revival writer owns it exclusively.
- The dispose drain deadline does not cancel in-flight event handlers: one
parked in a slow extension hook resumed after close/release, reopened the
append writer for its late persist, and repopulated the released state.
- Track whether the drain settled; on deadline, redo the final close +
release once the pipeline genuinely settles (hook runtime is bounded by
the extension runner).
- Expose drainTimeoutMs on AgentSessionDisposeOptions for bounded teardown
paths and deterministic coverage of the deadline branch.
- Added account-scoped policy error detection to correctly identify Codex cyber-policy rejections.
- Updated credential storage and retry logic to route denied accounts through sibling rotation instead of bypassing it.
- Ensured coding-agent sessions exhaust all sibling accounts before falling back on cyber denials.
- Added comprehensive test coverage for credential rotation and retry behavior on policy errors.
agent-core dispatches the session's event subscriber fire-and-forget (agent.ts #emit), so a message_end/agent_end handler can still be awaiting extension/subscriber/maintenance work — and its sessionManager/agent.state append — after agent.waitForIdle() resolves. The earlier settle waited only on the core run, so a late handler could append the finished message/entries back into the disposed session and re-pin the transcript.
Track every #handleAgentEvent dispatch in #inFlightEventHandlers and drain it (alongside agent.waitForIdle) inside the bounded settle before reset/clear/release. Added a regression test using a real extension whose message_end hook blocks before persistence, asserting dispose does not release memory until the in-flight handler settles.
dispose() only *signalled* the agent loop via abort(); it never awaited the run, so a mid-turn dispose (Ctrl-C/timeout/hard-killed subagent) could let the loop unwind after the release ran — its response/SSE interceptors re-recording wire frames into rawSseDebugBuffer and its terminal message re-appending to agent.state.messages, repopulating the disposed session with exactly the retained state the release drops.
Detach the response/SSE interceptors and await a bounded agent.waitForIdle() before the reset/clear so it lands on a quiescent session. Added a deterministic regression test that gates the active turn and asserts dispose blocks on it before clearing.
Disposed sessions remained reachable through lifecycle reviver closures. Agent.reset() cleared the live message array but left AppendOnlyContextManager attached, retaining its normalized provider transcript and stable prompt/tool prefix.
Detach the append-only manager during terminal disposal and extend the memory-release regression test to cover that second transcript copy.
Keep-alive subagents are handed to AgentLifecycleManager.adopt, which stores
their reviver closure in the process-global #adopted map. The closure is
defined inside runSubagent's scope, which also captures the live AgentSession
(extension-runner callbacks, buildSubagentSessionOptions), so its lexical
environment pins the whole session graph. park() disposes and detaches the
session but leaves the adoption record indefinitely, and #doDispose never
dropped the in-memory transcript, session-manager entries, or the raw-SSE
debug buffer (whose trimmed records retain slice() views of full wire frames),
so every completed subagent's heavy state leaked for the process lifetime.
dispose() is terminal and every revival path reopens from disk, so #doDispose
now sheds retained conversation memory via agent.reset(),
RawSseDebugBuffer.clear(), and SessionManager.releaseRetainedEntries(). The
adoption record can still reference the session, but only as a husk.
Fixes#8003
Harness-initiated session aborts previously cancelled compaction before the handoff reason was recorded. The handoff catch then saw only an aborted signal and replaced the harness reason with "Handoff cancelled".
Abort the handoff first with the session reason, forward caller-signal reasons, and reserve "Handoff cancelled" for direct or unreasoned cancellation. Add a regression test for an in-flight handoff aborted through AgentSession.abort.
Fixes#7993
The #7904 fix stopped masking provider errors as "Handoff cancelled", but
an empty or whitespace-only generation still fell through: whitespace-only
text passed the `!handoffText` guard and produced a bogus handoff, while
empty text returned undefined which the interactive /handoff caller mapped
to "Handoff cancelled" with no detail and no log entry.
Treat empty/whitespace-only output as a real failure: a user-initiated
handoff throws "Handoff generation produced no content" (surfaced as
"Handoff failed: ...") and logs it; auto-handoff keeps returning undefined
so maintenance falls back to context-full compaction. Also log genuine
handoff failures in the command controller so they persist for debugging.
Fixes#7993
Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.
`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
Codex found the shared attribution predicate recognising only tool calls,
text and signed thinking. A native image response often arrives with no
text and no tool call at all, so an image-only turn was read as producing
nothing: attribution stayed on whichever model spoke before it, and the
empty-stop rule could classify a successful generation as empty.
Everything the assistant can emit now counts except two: unsigned thinking,
which is not provider-authenticated and was already excluded, and
Anthropic's `fallback` marker, which records that a request was routed
elsewhere rather than carrying output. Redacted thinking and server-tool
blocks are real work by the same argument as the image.
The `toolUse` arm keeps its stricter rule — an orphaned toolUse stop needs
a tool_use block to anchor a later tool_result, and an image cannot.
Also switches the new test to the namespace import AGENTS.md requires for
node builtins.
Review found the branch does not compile. `bun check` runs biome before
the per-workspace type check and joins them with `&&`, so a pre-existing
format error on `main` short-circuited the run: `tsgo --noEmit` never
executed, and `bun test` type-strips, so the suite stayed green over eight
type errors. `ServingModel` was used in `turn-recovery.ts` without being
imported, and three `subscribe` closures in the retry-fallback suite lost
the outer narrowing of `session`. Verified now against the workspace check
directly rather than the aggregate.
Codex also found `#fallbackRouted` outliving its session. It said how the
CURRENT model was reached, but nothing reset it when a transcript was
switched or resumed, so a freshly loaded session with no served attribution
yet described its own model with the previous session's routing. It is now
anchored on the session id exactly like the attribution beside it — the two
facts are earned together and expire together — and a session switch
reports the current model without claiming to know how it got there.
The replay-unsafe host stub gained `getSessionId`, which the real
`SessionManager` has always had and the anchoring now calls.
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.
Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.
The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.
`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
Two cleanups from review.
The offline walk matched a served turn to its `model_change` by also
accepting a `:level` thinking suffix, but no writer produces one: every
`appendModelChange` call site records a bare `provider/id` or a
`formatModelStringWithRouting` selector, and only the latter decorates —
with `@upstream`, never a level. Speculative matching in the part of the
change that already reconstructs the most from raw strings. Removed, with
its test narrowed to the format writers actually emit.
Attribution was also retained for the life of the recovery component, so
switching sessions in place could report a model from a different
transcript by name. It is now tagged with the session file it was earned
in and ignored when that no longer matches — the same self-invalidating
anchor `#ensurePersistedMessageKeys` uses, so no mutation call site has
to remember to clear anything. Anchoring on the file rather than the leaf
keeps attribution across compaction and appends, which stay within the
transcript that earned it.
Review of the previous commit found the same defect in four more places,
each a variant of one mistake: the unproven-model gate was expressed as
"a retry fallback is armed and has not served", so it only protected the
one swap path that arms a chain.
- The cooldown-expiry restore to the primary deletes the fallback record
before swapping, so the gate vanished with it and an unproven, just-
restored primary was credited immediately.
- A session's first-ever fallback arm constructs the record only after
the swap, so there was nothing to mark and the guard was a no-op.
- The Fireworks Fast degrade swaps models without arming a chain at all,
so a live row showed the degraded base model as the configured primary.
- The fresh-versus-resumed heuristic read the message list, which
compaction collapses, so a resumed-and-compacted session seeded its
startup candidate as already proven.
Attribution now names the last model that settled a turn here, full stop.
A switch of any kind — into a candidate, back to a restored primary, or a
capability degrade — moves it only once the new model answers. Before
anything has served there is no earlier work to miscredit, so the
configured model is both the only answer and a safe one.
Fallback routing is tracked as its own flag rather than inferred from the
chain record, because the degrade path routes without a chain, and it is
set before each swap so the synchronous `model_changed` fan-out cannot
observe a model mid-relabel. The resumed-session heuristic is gone: it
was only needed to decide whether an unserved candidate could be trusted,
and nothing is trusted before it serves.
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.
Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.
Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.
Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.
One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.
Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.
A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.
`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
- Separated deterministic replacement generation, placeholder derivation,
placeholder-range scanning and message-tree transforms out of the 2647-line
module; obfuscator.ts now holds the types and SecretObfuscator.
- ephemeralPlaceholderKey stays a single instance and both global regexes stay
beside the code that resets their lastIndex, so placeholder stability and
the security argument in the moved comments are preserved verbatim.
- Repointed every importer at the real modules rather than leaving a re-export
shim; the public ./secrets barrel exports the same 15 names as before.
- A provider-supplied retry-after now bypasses the transient rate/concurrency
heuristic window instead of being overridden by it (regression from the
subscription-cap retry change).
- Updated event-controller/ui-helpers test doubles for provenance-gated
renderer selection (hasBuiltInTool), aggregated retryErrors on
auto_retry_end, and Bedrock override compat gaining streamIdleTimeoutMs.
A cooldown-expiry model revert runs at a turn boundary. The user-prompt
path reverts then re-checks accumulated context against the restored
model via runPrePromptCompactionIfNeeded, but the automatic
agent.continue() path (#scheduleAgentContinue) reverted and issued the
next request with no such check. When a transient failure had fallen
back to a larger-window model and the conversation then grew past the
original model's window, restoring the primary once its cooldown expired
sent a predictably oversized request to the smaller model.
maybeRestoreRetryFallbackPrimary now reports whether it actually
switched, and the auto-continue path runs the same post-revert
context-fit maintenance (compaction/promotion) the prompt path already
runs, but only when a revert occurred.
Fixes#7952
Anthropic's request classifier can refuse a turn after the model has already
streamed a tool call. `isRetryableError` bailed on `#hasReplayUnsafeOutput`
one line before `isClassifierRefusal` could make the turn retryable, so the
refusal ended the turn terminally and the configured `retry.fallbackChains`
entry was never consulted, even though a different model family would very
likely have served the request. The same refusal with no tool calls already
cascaded, and the advisor path recovers from the identical refusal via model
fallback, so only the main session path died.
`#refusalReplaySafe` lifts the veto for exactly that shape: a classifier
refusal whose only replay-unsafe blocks are tool calls, where every emitted
call id has a result after the assistant message in `agent.state.messages`
and every such result is synthetic with `executed === false` (the agent
loop's positive record that `tool.execute()` never ran). Images, Anthropic
server tools, and committed non-whitespace text still veto, since replaying
would duplicate rendered output. Any uncertainty keeps the veto: assistant
message absent from state, a call with no result, a non-synthetic result, or
an `executed` that is not exactly `false`.
`isHardErrorFallbackEligible` is unchanged; it declines refusals on purpose
because the chain consult for refusals runs on the retryable path with
`pinFallback`.
The handoff catch in session-handoff.ts and the /handoff handler in
command-controller.ts mapped any error named AbortError to "Handoff
cancelled" regardless of whether the handoff signal was actually
aborted. Providers throw name-AbortError errors on non-user conditions
(stalls, idle timeouts, nested resolution failures), so a genuine
generation failure surfaced as a user cancellation and hid the cause.
Only report "Handoff cancelled" when handoffSignal.aborted is set;
re-throw the real error otherwise. The controller now trusts the
normalized "Handoff cancelled" message and drops its own AbortError
check so re-thrown provider failures render as "Handoff failed: ...".
Fixes#7903