- Added a 60-character minimum length threshold for demoting interrupted thinking into hidden continuity context in agent-session.ts.
- Updated LLM message conversion in messages.ts to strip incomplete thinking from user-interrupted assistant turns regardless of continuity note presence.
- Added test coverage in agent-session-interrupted-thinking.test.ts verifying behavior for reasoning lengths below and at the threshold.
- Added the `--external-thinking` CLI flag alongside model capability checks to gate external thinking tool availability.
- Updated Anthropic and Google transports to honor `forceReasoningOff` for native thinking-off controls.
- Renamed the `thoughts` property and parameter to `notes` across think fixtures, tools, and tests.
- Updated system prompt instructions and test suites to verify transport-specific thinking and tool activation.
- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
Every session dispose broadcast the shutdown abort reason, so an explicit
hard kill of a subagent (release with tombstone, then live.dispose()) tagged
its nested children as shutdown and rediscovered them as parked instead of
terminal.
- Gate ASYNC_JOB_MANAGER_SHUTDOWN_REASON on #ownedAsyncJobManager so only the
top-level owning session's dispose (genuine process shutdown) uses it.
- Subagent disposes propagate a generic cancellation, keeping nested children
terminal.
- Cover the subagent generic-cancel path alongside the owning-session shutdown.
Fixes#8216
Session teardown pre-cancels owner jobs via #cancelOwnAsyncJobs before
manager.dispose(), so the shutdown abort reason must ride along that
cancelAll or an owned subagent job sees a generic caller signal and is
tombstoned instead of parked.
- Forward an abort reason through AsyncJobManager.cancelAll.
- Pass ASYNC_JOB_MANAGER_SHUTDOWN_REASON from #disposeOwnedAsyncJobs.
- Cover owned-job shutdown tagging through a real AgentSession dispose.
Fixes#8216
System and mode prompts referenced tools that may be absent from the
session catalog, forcing the model to satisfy requirements it cannot
execute (generalizes #8139's browser-verification mismatch):
- system-prompt.md: browser verification now keys on the actual UI
surface and available tools (browser/computer/TUI/CLI), with an
explicit behavioral/smoke-test fallback when no runtime tool exists;
todo workflow guidance and the AST-section grep hint are gated on
tool presence; the auto-QA report_issue block additionally requires
the write tool.
- project-prompt.md: workspace-tree drill-in and additional-roots tool
hints only name tools present in the session.
- plan-mode-active.md: ask-tool directives get a prose fallback and
scout-via-task dispatch is gated on the task tool (render site now
passes askAvailable/taskAvailable).
- orchestrate-notice.md: the tool budget, verify gates, todo tracking,
and inline-edit guidance are gated on tool presence; the notice is
skipped entirely when the task tool is inactive (render fn takes the
active tool list).
Adds regression tests for each gate; changelog entry.
retry_fallback_applied and retry_fallback_succeeded were emitted to the
TUI and RPC subscribers but never reached extensions: AgentSession's
#emitExtensionEvent had no branch mapping either event to
ExtensionRunner.emit, and ExtensionEvent / ExtensionAPI.on lacked the
types and overloads, so registration was also rejected at compile time.
Add the typed events, on() overloads, union members, and the two
forwarding branches so extensions observe the same { from, to, role } /
{ model, role } payloads as TUI and RPC. Matches the contract already
documented in docs/non-compaction-retry-policy.md.
Fixes#8079
- Passed the known advisor role through retry fallback resolution so model and wildcard keys retain precedence while ambiguous role matches cannot select another role.
- Covered shared-model roles with distinct thinking levels and asserted advisor fallback lifecycle ownership.
Fixes#8075
Retry-fallback candidate selection filtered on suppression, effort
ceiling, model resolution, and API key, but never compared a
candidate's context window with the live context. A large-window
primary hitting a retryable error could switch onto a smaller-window
fallback and immediately send a predictably oversized request that the
provider rejects, stalling the run. This is the forward counterpart of
the #7952 cooldown-expiry revert fix.
Generalize the existing retry-fit budget check into
SessionMaintenance.contextFitsModel(model) and consult it from both
retry-fallback selection loops (#tryRetryModelFallback and the
usage-aware loop): skip any candidate whose usable window cannot hold
the current context and advance to the first configured candidate that
fits. The check is independent of compaction.enabled since an oversized
request overflows regardless.
Fixes#8065
- seal() now runs before the final dispose close() and bumps the disk
epoch: work an event handler enqueues while dispose awaits the closing
tail is superseded, and an already-running fenced or authoritative
rewrite fails its commit guard at the rename fence instead of
publishing over a file a revival reopened.
- The authoritative repair path resets the disk tail itself, escaping the
close() serialization, and would atomically publish the emptied entry
list; it now no-ops once sealed and commit guards also check the seal.
- Execution-time gates cover the queued title persist and fenced rewrite
callbacks; setSessionName/appendCustomEntry attempted by a handler that
outlives dispose are dropped and covered by the seal regression, and a
failed atomic batch across the seal can no longer truncate the file.
- AgentLifecycleManager.park() resolves as soon as dispose() returns, so
ensureLive() may reopen the same JSONL through a new manager while a
timed-out event handler still holds the old one; a late append reopened
a second writer on that file and the deferred finalize then closed it.
- A post-release rewrite was worse: it persisted the emptied entry list,
truncating the transcript on disk.
- releaseRetainedEntries() now seals the manager: appends, title changes,
and rewrites become dropped no-ops and the append writer is closed, so
the deferred dispose pass is in-memory only and can never touch the file.
- File-backed regression: dispose on the drain deadline, revive the JSONL
immediately, unpark the late handler, and prove the file is byte-stable
and the revival writer owns it exclusively.
- The dispose drain deadline does not cancel in-flight event handlers: one
parked in a slow extension hook resumed after close/release, reopened the
append writer for its late persist, and repopulated the released state.
- Track whether the drain settled; on deadline, redo the final close +
release once the pipeline genuinely settles (hook runtime is bounded by
the extension runner).
- Expose drainTimeoutMs on AgentSessionDisposeOptions for bounded teardown
paths and deterministic coverage of the deadline branch.
agent-core dispatches the session's event subscriber fire-and-forget (agent.ts #emit), so a message_end/agent_end handler can still be awaiting extension/subscriber/maintenance work — and its sessionManager/agent.state append — after agent.waitForIdle() resolves. The earlier settle waited only on the core run, so a late handler could append the finished message/entries back into the disposed session and re-pin the transcript.
Track every #handleAgentEvent dispatch in #inFlightEventHandlers and drain it (alongside agent.waitForIdle) inside the bounded settle before reset/clear/release. Added a regression test using a real extension whose message_end hook blocks before persistence, asserting dispose does not release memory until the in-flight handler settles.
dispose() only *signalled* the agent loop via abort(); it never awaited the run, so a mid-turn dispose (Ctrl-C/timeout/hard-killed subagent) could let the loop unwind after the release ran — its response/SSE interceptors re-recording wire frames into rawSseDebugBuffer and its terminal message re-appending to agent.state.messages, repopulating the disposed session with exactly the retained state the release drops.
Detach the response/SSE interceptors and await a bounded agent.waitForIdle() before the reset/clear so it lands on a quiescent session. Added a deterministic regression test that gates the active turn and asserts dispose blocks on it before clearing.
Disposed sessions remained reachable through lifecycle reviver closures. Agent.reset() cleared the live message array but left AppendOnlyContextManager attached, retaining its normalized provider transcript and stable prompt/tool prefix.
Detach the append-only manager during terminal disposal and extend the memory-release regression test to cover that second transcript copy.
Keep-alive subagents are handed to AgentLifecycleManager.adopt, which stores
their reviver closure in the process-global #adopted map. The closure is
defined inside runSubagent's scope, which also captures the live AgentSession
(extension-runner callbacks, buildSubagentSessionOptions), so its lexical
environment pins the whole session graph. park() disposes and detaches the
session but leaves the adoption record indefinitely, and #doDispose never
dropped the in-memory transcript, session-manager entries, or the raw-SSE
debug buffer (whose trimmed records retain slice() views of full wire frames),
so every completed subagent's heavy state leaked for the process lifetime.
dispose() is terminal and every revival path reopens from disk, so #doDispose
now sheds retained conversation memory via agent.reset(),
RawSseDebugBuffer.clear(), and SessionManager.releaseRetainedEntries(). The
adoption record can still reference the session, but only as a husk.
Fixes#8003
Harness-initiated session aborts previously cancelled compaction before the handoff reason was recorded. The handoff catch then saw only an aborted signal and replaced the harness reason with "Handoff cancelled".
Abort the handoff first with the session reason, forward caller-signal reasons, and reserve "Handoff cancelled" for direct or unreasoned cancellation. Add a regression test for an in-flight handoff aborted through AgentSession.abort.
Fixes#7993
Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.
`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.
Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.
The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.
`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.
Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.
Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.
Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.
One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.
Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.
A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.
`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
- Separated deterministic replacement generation, placeholder derivation,
placeholder-range scanning and message-tree transforms out of the 2647-line
module; obfuscator.ts now holds the types and SecretObfuscator.
- ephemeralPlaceholderKey stays a single instance and both global regexes stay
beside the code that resets their lastIndex, so placeholder stability and
the security argument in the moved comments are preserved verbatim.
- Repointed every importer at the real modules rather than leaving a re-export
shim; the public ./secrets barrel exports the same 15 names as before.
A cooldown-expiry model revert runs at a turn boundary. The user-prompt
path reverts then re-checks accumulated context against the restored
model via runPrePromptCompactionIfNeeded, but the automatic
agent.continue() path (#scheduleAgentContinue) reverted and issued the
next request with no such check. When a transient failure had fallen
back to a larger-window model and the conversation then grew past the
original model's window, restoring the primary once its cooldown expired
sent a predictably oversized request to the smaller model.
maybeRestoreRetryFallbackPrimary now reports whether it actually
switched, and the auto-continue path runs the same post-revert
context-fit maintenance (compaction/promotion) the prompt path already
runs, but only when a revert occurred.
Fixes#7952
The per-turn systemPrompt returned by before_agent_start was applied only to the agent state, so any base-prompt rebuild firing in the prompt window (context-overflow compaction/promotion, memory promotion, MCP/RPC tool refresh, hindsight MM-TTL refresh) re-pushed the rebuilt base via setSystemPrompt and silently dropped the override before the request.
SessionTools now tracks the active per-turn override and re-applies it on every base rebuild during the turn, clearing it when the turn ends.
Fixes#7755