- Added the `--external-thinking` CLI flag alongside model capability checks to gate external thinking tool availability.
- Updated Anthropic and Google transports to honor `forceReasoningOff` for native thinking-off controls.
- Renamed the `thoughts` property and parameter to `notes` across think fixtures, tools, and tests.
- Updated system prompt instructions and test suites to verify transport-specific thinking and tool activation.
- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
/handoff mints a fresh session via newSession(), producing a new
artifactsDir and an empty local/ root. The handoff document routinely
references plans and scratch files under '/data/workspaces/can1357__oh-my-pi__8261/.omp-session/2026-08-11T16-39-09-489Z_019ff1b1-31b1-7000-81f5-c540f4ebf43d/local/,' so every reference
became a dangling pointer in the new session. The plan approve-and-execute
path already copies artifacts across the boundary; handoff did not.
Extracted the plan-approve copy helper into a shared copyLocalArtifacts()
in local-protocol.ts and invoke it across the handoff session switch
(best-effort, since the switch is already committed).
Fixes#8261
Allowed advisor streams to treat Google STOP responses without visible content as successful silence while preserving the default retry behavior for interactive agents.
Added provider and advisor-path regressions covering retry counts, system instructions, and the advise declaration.
Fixes#8223
Every session dispose broadcast the shutdown abort reason, so an explicit
hard kill of a subagent (release with tombstone, then live.dispose()) tagged
its nested children as shutdown and rediscovered them as parked instead of
terminal.
- Gate ASYNC_JOB_MANAGER_SHUTDOWN_REASON on #ownedAsyncJobManager so only the
top-level owning session's dispose (genuine process shutdown) uses it.
- Subagent disposes propagate a generic cancellation, keeping nested children
terminal.
- Cover the subagent generic-cancel path alongside the owning-session shutdown.
Fixes#8216
Session teardown pre-cancels owner jobs via #cancelOwnAsyncJobs before
manager.dispose(), so the shutdown abort reason must ride along that
cancelAll or an owned subagent job sees a generic caller signal and is
tombstoned instead of parked.
- Forward an abort reason through AsyncJobManager.cancelAll.
- Pass ASYNC_JOB_MANAGER_SHUTDOWN_REASON from #disposeOwnedAsyncJobs.
- Cover owned-job shutdown tagging through a real AgentSession dispose.
Fixes#8216
Why:
The module-global array memo strongly retains the most recently converted
transcript and output after its session is disposed.
Changes:
- Store exact-repeat and append-growth state in a WeakMap per input array.
- Preserve generation invalidation and per-message weak caching.
Evidence:
- Seven disposal runs collected both arrays after 50 forced-GC passes and
reduced median heap delta by 87.54%.
Refs #8119
Why:
SessionManager.open() parses the complete journal for its header and then
setSessionFile() parses the same journal again.
Changes:
- Reuse the entries already loaded by open() through a private setup path.
- Keep the public setSessionFile() contract unchanged.
Evidence:
- A 19.98 MB, 12,000-entry journal improved from 48.011 ms to
25.406 ms median across 15 runs with identical restored state.
Refs #8117
runRootCommand called discoverAuthStorage without a try/catch, so a
configured-but-unreachable broker with no cached snapshot re-threw
AuthBrokerError as a raw uncaught exception at startup, unlike the other
startup paths that print a clean stderr message and exit non-zero.
Wrap the startup auth discovery: broker failures now report an actionable
message naming the broker URL and the recovery options (start it with
`omp auth-broker serve`, or reset `auth.broker.url`/`auth.broker.token`)
and exit 1. Unrelated errors still propagate. The broker still replaces
the local store when configured; no silent fallback to local credentials.
Fixes#8096
retry_fallback_applied and retry_fallback_succeeded were emitted to the
TUI and RPC subscribers but never reached extensions: AgentSession's
#emitExtensionEvent had no branch mapping either event to
ExtensionRunner.emit, and ExtensionEvent / ExtensionAPI.on lacked the
types and overloads, so registration was also rejected at compile time.
Add the typed events, on() overloads, union members, and the two
forwarding branches so extensions observe the same { from, to, role } /
{ model, role } payloads as TUI and RPC. Matches the contract already
documented in docs/non-compaction-retry-policy.md.
Fixes#8079
- Passed the known advisor role through retry fallback resolution so model and wildcard keys retain precedence while ambiguous role matches cannot select another role.
- Covered shared-model roles with distinct thinking levels and asserted advisor fallback lifecycle ownership.
Fixes#8075
Retry-fallback candidate selection filtered on suppression, effort
ceiling, model resolution, and API key, but never compared a
candidate's context window with the live context. A large-window
primary hitting a retryable error could switch onto a smaller-window
fallback and immediately send a predictably oversized request that the
provider rejects, stalling the run. This is the forward counterpart of
the #7952 cooldown-expiry revert fix.
Generalize the existing retry-fit budget check into
SessionMaintenance.contextFitsModel(model) and consult it from both
retry-fallback selection loops (#tryRetryModelFallback and the
usage-aware loop): skip any candidate whose usable window cannot hold
the current context and advance to the first configured candidate that
fits. The check is independent of compaction.enabled since an oversized
request overflows regardless.
Fixes#8065
- seal() now runs before the final dispose close() and bumps the disk
epoch: work an event handler enqueues while dispose awaits the closing
tail is superseded, and an already-running fenced or authoritative
rewrite fails its commit guard at the rename fence instead of
publishing over a file a revival reopened.
- The authoritative repair path resets the disk tail itself, escaping the
close() serialization, and would atomically publish the emptied entry
list; it now no-ops once sealed and commit guards also check the seal.
- Execution-time gates cover the queued title persist and fenced rewrite
callbacks; setSessionName/appendCustomEntry attempted by a handler that
outlives dispose are dropped and covered by the seal regression, and a
failed atomic batch across the seal can no longer truncate the file.
- AgentLifecycleManager.park() resolves as soon as dispose() returns, so
ensureLive() may reopen the same JSONL through a new manager while a
timed-out event handler still holds the old one; a late append reopened
a second writer on that file and the deferred finalize then closed it.
- A post-release rewrite was worse: it persisted the emptied entry list,
truncating the transcript on disk.
- releaseRetainedEntries() now seals the manager: appends, title changes,
and rewrites become dropped no-ops and the append writer is closed, so
the deferred dispose pass is in-memory only and can never touch the file.
- File-backed regression: dispose on the drain deadline, revive the JSONL
immediately, unpark the late handler, and prove the file is byte-stable
and the revival writer owns it exclusively.
- The dispose drain deadline does not cancel in-flight event handlers: one
parked in a slow extension hook resumed after close/release, reopened the
append writer for its late persist, and repopulated the released state.
- Track whether the drain settled; on deadline, redo the final close +
release once the pipeline genuinely settles (hook runtime is bounded by
the extension runner).
- Expose drainTimeoutMs on AgentSessionDisposeOptions for bounded teardown
paths and deterministic coverage of the deadline branch.
- Added account-scoped policy error detection to correctly identify Codex cyber-policy rejections.
- Updated credential storage and retry logic to route denied accounts through sibling rotation instead of bypassing it.
- Ensured coding-agent sessions exhaust all sibling accounts before falling back on cyber denials.
- Added comprehensive test coverage for credential rotation and retry behavior on policy errors.
agent-core dispatches the session's event subscriber fire-and-forget (agent.ts #emit), so a message_end/agent_end handler can still be awaiting extension/subscriber/maintenance work — and its sessionManager/agent.state append — after agent.waitForIdle() resolves. The earlier settle waited only on the core run, so a late handler could append the finished message/entries back into the disposed session and re-pin the transcript.
Track every #handleAgentEvent dispatch in #inFlightEventHandlers and drain it (alongside agent.waitForIdle) inside the bounded settle before reset/clear/release. Added a regression test using a real extension whose message_end hook blocks before persistence, asserting dispose does not release memory until the in-flight handler settles.
dispose() only *signalled* the agent loop via abort(); it never awaited the run, so a mid-turn dispose (Ctrl-C/timeout/hard-killed subagent) could let the loop unwind after the release ran — its response/SSE interceptors re-recording wire frames into rawSseDebugBuffer and its terminal message re-appending to agent.state.messages, repopulating the disposed session with exactly the retained state the release drops.
Detach the response/SSE interceptors and await a bounded agent.waitForIdle() before the reset/clear so it lands on a quiescent session. Added a deterministic regression test that gates the active turn and asserts dispose blocks on it before clearing.
Disposed sessions remained reachable through lifecycle reviver closures. Agent.reset() cleared the live message array but left AppendOnlyContextManager attached, retaining its normalized provider transcript and stable prompt/tool prefix.
Detach the append-only manager during terminal disposal and extend the memory-release regression test to cover that second transcript copy.
Keep-alive subagents are handed to AgentLifecycleManager.adopt, which stores
their reviver closure in the process-global #adopted map. The closure is
defined inside runSubagent's scope, which also captures the live AgentSession
(extension-runner callbacks, buildSubagentSessionOptions), so its lexical
environment pins the whole session graph. park() disposes and detaches the
session but leaves the adoption record indefinitely, and #doDispose never
dropped the in-memory transcript, session-manager entries, or the raw-SSE
debug buffer (whose trimmed records retain slice() views of full wire frames),
so every completed subagent's heavy state leaked for the process lifetime.
dispose() is terminal and every revival path reopens from disk, so #doDispose
now sheds retained conversation memory via agent.reset(),
RawSseDebugBuffer.clear(), and SessionManager.releaseRetainedEntries(). The
adoption record can still reference the session, but only as a husk.
Fixes#8003
Harness-initiated session aborts previously cancelled compaction before the handoff reason was recorded. The handoff catch then saw only an aborted signal and replaced the harness reason with "Handoff cancelled".
Abort the handoff first with the session reason, forward caller-signal reasons, and reserve "Handoff cancelled" for direct or unreasoned cancellation. Add a regression test for an in-flight handoff aborted through AgentSession.abort.
Fixes#7993
The #7904 fix stopped masking provider errors as "Handoff cancelled", but
an empty or whitespace-only generation still fell through: whitespace-only
text passed the `!handoffText` guard and produced a bogus handoff, while
empty text returned undefined which the interactive /handoff caller mapped
to "Handoff cancelled" with no detail and no log entry.
Treat empty/whitespace-only output as a real failure: a user-initiated
handoff throws "Handoff generation produced no content" (surfaced as
"Handoff failed: ...") and logs it; auto-handoff keeps returning undefined
so maintenance falls back to context-full compaction. Also log genuine
handoff failures in the command controller so they persist for debugging.
Fixes#7993
Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.
`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
Codex found the shared attribution predicate recognising only tool calls,
text and signed thinking. A native image response often arrives with no
text and no tool call at all, so an image-only turn was read as producing
nothing: attribution stayed on whichever model spoke before it, and the
empty-stop rule could classify a successful generation as empty.
Everything the assistant can emit now counts except two: unsigned thinking,
which is not provider-authenticated and was already excluded, and
Anthropic's `fallback` marker, which records that a request was routed
elsewhere rather than carrying output. Redacted thinking and server-tool
blocks are real work by the same argument as the image.
The `toolUse` arm keeps its stricter rule — an orphaned toolUse stop needs
a tool_use block to anchor a later tool_result, and an image cannot.
Also switches the new test to the namespace import AGENTS.md requires for
node builtins.