Commit Graph
1878 Commits
Author SHA1 Message Date
can1357 e8782235e2 Merge PR #8097: fix(cli): surface actionable error when auth broker is unreachable (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 932bb6d244 Merge PR #8080: fix(session): forward retry fallback events to extensions (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 a5434d7341 Merge PR #8076: fix(advisor): preserve advisor fallback role ownership (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 abc1897fc3 fix(agent): fit retry fallback against retry prompt 2026-08-13 01:14:46 +02:00
can1357 f73baac583 Merge PR #8068: fix(agent): fit-check retry fallback before switching models (@roboomp) 2026-08-13 01:14:46 +02:00
can1357 19c0afcc0d feat: implemented external thinking flags and transport reasoning controls
- Added the `--external-thinking` CLI flag alongside model capability checks to gate external thinking tool availability.
- Updated Anthropic and Google transports to honor `forceReasoningOff` for native thinking-off controls.
- Renamed the `thoughts` property and parameter to `notes` across think fixtures, tools, and tests.
- Updated system prompt instructions and test suites to verify transport-specific thinking and tool activation.
2026-08-12 02:23:17 +02:00
can1357 fe837038e9 Merge PR #8262: fix(coding-agent): copy local artifacts across handoff session boundary (@roboomp) 2026-08-12 01:53:44 +02:00
can1357 10fd42289c feat: introduced external thinking support and private scratchpad think tool
- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
2026-08-11 20:39:57 +02:00
roboomp dbc199d2db fix(coding-agent): copy local artifacts across handoff session boundary
/handoff mints a fresh session via newSession(), producing a new
artifactsDir and an empty local/ root. The handoff document routinely
references plans and scratch files under '/data/workspaces/can1357__oh-my-pi__8261/.omp-session/2026-08-11T16-39-09-489Z_019ff1b1-31b1-7000-81f5-c540f4ebf43d/local/,' so every reference
became a dangling pointer in the new session. The plan approve-and-execute
path already copies artifacts across the boundary; handoff did not.

Extracted the plan-approve copy helper into a shared copyLocalArtifacts()
in local-protocol.ts and invoke it across the handoff session switch
(best-effort, since the switch is already committed).

Fixes #8261
2026-08-11 16:45:15 +00:00
can1357 47b282ff9b Merge PR #8069: fix(extensions): register lifecycle tools (@mrexodia) 2026-08-11 15:24:18 +02:00
can1357 ca6e13fd47 Merge PR #8218: fix(agent): park subagents on shutdown (@roboomp) 2026-08-11 15:18:17 +02:00
can1357 3b66177abd Merge PR #8226: fix(advisor): accept silent Gemini reviews (@roboomp) 2026-08-11 15:15:44 +02:00
can1357 2670dc345c Merge PR #7930: fix(coding-agent): let a classifier refusal cascade when its tool calls never ran (@mvid)
# Conflicts:
#	packages/coding-agent/src/session/turn-recovery.ts
2026-08-11 15:09:34 +02:00
can1357 7e8be71ead Merge PR #7247: fix(advisor): prevent false full replays and preserve cache growth (@cuipengfei) 2026-08-11 15:07:43 +02:00
can1357 55d1831dd3 Merge PR #8152: perf(session): weakly cache conversion histories (@MikeeI) 2026-08-11 15:06:15 +02:00
can1357 5ba362777a Merge PR #8149: perf(session): reuse parsed journal during resume (@MikeeI) 2026-08-11 15:06:15 +02:00
can1357 6592b799b3 Merge PR #7980: fix(session): attribute a run to the model that produced its output (@enieuwy) 2026-08-11 15:06:12 +02:00
roboomp 86b8f510c8 fix(advisor): accepted silent Gemini reviews
Allowed advisor streams to treat Google STOP responses without visible content as successful silence while preserving the default retry behavior for interactive agents.

Added provider and advisor-path regressions covering retry counts, system instructions, and the advise declaration.

Fixes #8223
2026-08-11 08:22:54 +00:00
roboomp 546016596e fix(agent): scope shutdown reason to the owning session
Every session dispose broadcast the shutdown abort reason, so an explicit
hard kill of a subagent (release with tombstone, then live.dispose()) tagged
its nested children as shutdown and rediscovered them as parked instead of
terminal.

- Gate ASYNC_JOB_MANAGER_SHUTDOWN_REASON on #ownedAsyncJobManager so only the
  top-level owning session's dispose (genuine process shutdown) uses it.
- Subagent disposes propagate a generic cancellation, keeping nested children
  terminal.
- Cover the subagent generic-cancel path alongside the owning-session shutdown.

Fixes #8216
2026-08-11 06:40:44 +00:00
roboomp 8614979281 fix(agent): tag owned jobs with shutdown reason before dispose
Session teardown pre-cancels owner jobs via #cancelOwnAsyncJobs before
manager.dispose(), so the shutdown abort reason must ride along that
cancelAll or an owned subagent job sees a generic caller signal and is
tombstoned instead of parked.

- Forward an abort reason through AsyncJobManager.cancelAll.
- Pass ASYNC_JOB_MANAGER_SHUTDOWN_REASON from #disposeOwnedAsyncJobs.
- Cover owned-job shutdown tagging through a real AgentSession dispose.

Fixes #8216
2026-08-11 06:32:46 +00:00
Bonobo 3321be5055 perf(session): weakly cache conversion histories
Why:
The module-global array memo strongly retains the most recently converted
transcript and output after its session is disposed.

Changes:
- Store exact-repeat and append-growth state in a WeakMap per input array.
- Preserve generation invalidation and per-message weak caching.

Evidence:
- Seven disposal runs collected both arrays after 50 forced-GC passes and
  reduced median heap delta by 87.54%.

Refs #8119
2026-08-10 04:09:44 +02:00
Bonobo 00fc4d5376 perf(session): reuse parsed journal during resume
Why:
SessionManager.open() parses the complete journal for its header and then
setSessionFile() parses the same journal again.

Changes:
- Reuse the entries already loaded by open() through a private setup path.
- Keep the public setSessionFile() contract unchanged.

Evidence:
- A 19.98 MB, 12,000-entry journal improved from 48.011 ms to
  25.406 ms median across 15 runs with identical restored state.

Refs #8117
2026-08-10 04:08:09 +02:00
Duncan Ogilvie 47f7fb0f0d fix(mcp): activate only retained manager tools 2026-08-10 02:24:30 +02:00
Duncan Ogilvie 158a70ad54 fix(extensions): preserve MCP ownership on refresh 2026-08-10 02:09:23 +02:00
Duncan Ogilvie 77916b6c99 fix(extensions): restore activation rollback context 2026-08-10 01:54:36 +02:00
Duncan Ogilvie 0cf54f428c fix(extensions): preserve mutation queue ownership 2026-08-10 01:40:18 +02:00
Duncan Ogilvie bbca0153eb fix(extensions): serialize dynamic tool refreshes 2026-08-10 01:28:43 +02:00
Duncan Ogilvie d1f73de44c fix(extensions): serialize registry mutations 2026-08-10 01:17:23 +02:00
Duncan Ogilvie 4d0c346f87 fix(extensions): abort timed-out tool activations 2026-08-10 00:54:42 +02:00
Duncan Ogilvie 2c55e20365 fix(extensions): expose prompt refresh option 2026-08-09 22:13:11 +02:00
Duncan Ogilvie 05f17fa76f fix(extensions): make lifecycle registration atomic 2026-08-09 22:07:46 +02:00
roboomp a67da14e4e fix(cli): surface actionable error when auth broker is unreachable
runRootCommand called discoverAuthStorage without a try/catch, so a
configured-but-unreachable broker with no cached snapshot re-threw
AuthBrokerError as a raw uncaught exception at startup, unlike the other
startup paths that print a clean stderr message and exit non-zero.

Wrap the startup auth discovery: broker failures now report an actionable
message naming the broker URL and the recovery options (start it with
`omp auth-broker serve`, or reset `auth.broker.url`/`auth.broker.token`)
and exit 1. Unrelated errors still propagate. The broker still replaces
the local store when configured; no silent fallback to local credentials.

Fixes #8096
2026-08-09 18:41:05 +00:00
roboomp ffd9d5c8ae fix(session): forward retry fallback events to extensions
retry_fallback_applied and retry_fallback_succeeded were emitted to the
TUI and RPC subscribers but never reached extensions: AgentSession's
#emitExtensionEvent had no branch mapping either event to
ExtensionRunner.emit, and ExtensionEvent / ExtensionAPI.on lacked the
types and overloads, so registration was also rejected at compile time.

Add the typed events, on() overloads, union members, and the two
forwarding branches so extensions observe the same { from, to, role } /
{ model, role } payloads as TUI and RPC. Matches the contract already
documented in docs/non-compaction-retry-policy.md.

Fixes #8079
2026-08-09 14:20:24 +00:00
roboomp 121bcb3663 fix(advisor): preserved advisor role fallback ownership
- Passed the known advisor role through retry fallback resolution so model and wildcard keys retain precedence while ambiguous role matches cannot select another role.

- Covered shared-model roles with distinct thinking levels and asserted advisor fallback lifecycle ownership.

Fixes #8075
2026-08-09 13:48:46 +00:00
roboomp 7b6548f182 fix(agent): fit-check retry fallback before switching models
Retry-fallback candidate selection filtered on suppression, effort
ceiling, model resolution, and API key, but never compared a
candidate's context window with the live context. A large-window
primary hitting a retryable error could switch onto a smaller-window
fallback and immediately send a predictably oversized request that the
provider rejects, stalling the run. This is the forward counterpart of
the #7952 cooldown-expiry revert fix.

Generalize the existing retry-fit budget check into
SessionMaintenance.contextFitsModel(model) and consult it from both
retry-fallback selection loops (#tryRetryModelFallback and the
usage-aware loop): skip any candidate whose usable window cannot hold
the current context and advance to the first configured candidate that
fits. The check is independent of compaction.enabled since an oversized
request overflows regardless.

Fixes #8065
2026-08-09 11:03:25 +00:00
can1357 d85dde52c4 fix(session): fence in-flight disk work behind the terminal seal
- seal() now runs before the final dispose close() and bumps the disk
  epoch: work an event handler enqueues while dispose awaits the closing
  tail is superseded, and an already-running fenced or authoritative
  rewrite fails its commit guard at the rename fence instead of
  publishing over a file a revival reopened.
- The authoritative repair path resets the disk tail itself, escaping the
  close() serialization, and would atomically publish the emptied entry
  list; it now no-ops once sealed and commit guards also check the seal.
- Execution-time gates cover the queued title persist and fenced rewrite
  callbacks; setSessionName/appendCustomEntry attempted by a handler that
  outlives dispose are dropped and covered by the seal regression, and a
  failed atomic batch across the seal can no longer truncate the file.
2026-08-08 20:49:28 +02:00
can1357 63aa8cf6f8 fix(session): seal the manager at terminal release against revival races
- AgentLifecycleManager.park() resolves as soon as dispose() returns, so
  ensureLive() may reopen the same JSONL through a new manager while a
  timed-out event handler still holds the old one; a late append reopened
  a second writer on that file and the deferred finalize then closed it.
- A post-release rewrite was worse: it persisted the emptied entry list,
  truncating the transcript on disk.
- releaseRetainedEntries() now seals the manager: appends, title changes,
  and rewrites become dropped no-ops and the append writer is closed, so
  the deferred dispose pass is in-memory only and can never touch the file.
- File-backed regression: dispose on the drain deadline, revive the JSONL
  immediately, unpark the late handler, and prove the file is byte-stable
  and the revival writer owns it exclusively.
2026-08-08 20:04:09 +02:00
can1357 31d7655477 fix(agent): re-finalize dispose after the drain deadline
- The dispose drain deadline does not cancel in-flight event handlers: one
  parked in a slow extension hook resumed after close/release, reopened the
  append writer for its late persist, and repopulated the released state.
- Track whether the drain settled; on deadline, redo the final close +
  release once the pipeline genuinely settles (hook runtime is bounded by
  the extension runner).
- Expose drainTimeoutMs on AgentSessionDisposeOptions for bounded teardown
  paths and deterministic coverage of the deadline branch.
2026-08-08 19:54:33 +02:00
can1357 f87a8ecc19 Merge PR #7994: fix(session): surface empty handoff generation as failure not cancel (@roboomp) 2026-08-08 19:38:32 +02:00
can1357 b2d0ade665 fix(agent): finish session disposal after event drain 2026-08-08 19:38:32 +02:00
can1357 bca8d5281b Merge PR #8004: fix(agent): release parked subagent session memory on dispose (@roboomp) 2026-08-08 19:38:32 +02:00
can1357 7ca140f66e fix(ai): routed policy-rejected accounts through sibling rotation
- Added account-scoped policy error detection to correctly identify Codex cyber-policy rejections.
- Updated credential storage and retry logic to route denied accounts through sibling rotation instead of bypassing it.
- Ensured coding-agent sessions exhaust all sibling accounts before falling back on cyber denials.
- Added comprehensive test coverage for credential rotation and retry behavior on policy errors.
2026-08-08 19:29:49 +02:00
roboomp 88021b90ea fix(agent): drained in-flight session event handlers on dispose
agent-core dispatches the session's event subscriber fire-and-forget (agent.ts #emit), so a message_end/agent_end handler can still be awaiting extension/subscriber/maintenance work — and its sessionManager/agent.state append — after agent.waitForIdle() resolves. The earlier settle waited only on the core run, so a late handler could append the finished message/entries back into the disposed session and re-pin the transcript.

Track every #handleAgentEvent dispatch in #inFlightEventHandlers and drain it (alongside agent.waitForIdle) inside the bounded settle before reset/clear/release. Added a regression test using a real extension whose message_end hook blocks before persistence, asserting dispose does not release memory until the in-flight handler settles.
2026-08-08 10:26:41 +00:00
roboomp a7a32f35dd fix(agent): settled active turn before clearing session memory
dispose() only *signalled* the agent loop via abort(); it never awaited the run, so a mid-turn dispose (Ctrl-C/timeout/hard-killed subagent) could let the loop unwind after the release ran — its response/SSE interceptors re-recording wire frames into rawSseDebugBuffer and its terminal message re-appending to agent.state.messages, repopulating the disposed session with exactly the retained state the release drops.

Detach the response/SSE interceptors and await a bounded agent.waitForIdle() before the reset/clear so it lands on a quiescent session. Added a deterministic regression test that gates the active turn and asserts dispose blocks on it before clearing.
2026-08-08 10:06:55 +00:00
roboomp 4ca6c376b4 fix(agent): detached append-only context on dispose
Disposed sessions remained reachable through lifecycle reviver closures. Agent.reset() cleared the live message array but left AppendOnlyContextManager attached, retaining its normalized provider transcript and stable prompt/tool prefix.

Detach the append-only manager during terminal disposal and extend the memory-release regression test to cover that second transcript copy.
2026-08-08 09:52:15 +00:00
roboomp ed6300b35e fix(agent): release parked subagent session memory on dispose
Keep-alive subagents are handed to AgentLifecycleManager.adopt, which stores
their reviver closure in the process-global #adopted map. The closure is
defined inside runSubagent's scope, which also captures the live AgentSession
(extension-runner callbacks, buildSubagentSessionOptions), so its lexical
environment pins the whole session graph. park() disposes and detaches the
session but leaves the adoption record indefinitely, and #doDispose never
dropped the in-memory transcript, session-manager entries, or the raw-SSE
debug buffer (whose trimmed records retain slice() views of full wire frames),
so every completed subagent's heavy state leaked for the process lifetime.

dispose() is terminal and every revival path reopens from disk, so #doDispose
now sheds retained conversation memory via agent.reset(),
RawSseDebugBuffer.clear(), and SessionManager.releaseRetainedEntries(). The
adoption record can still reference the session, but only as a husk.

Fixes #8003
2026-08-08 09:42:35 +00:00
roboomp 7914e7c451 fix(session): preserved harness handoff abort reasons
Harness-initiated session aborts previously cancelled compaction before the handoff reason was recorded. The handoff catch then saw only an aborted signal and replaced the harness reason with "Handoff cancelled".

Abort the handoff first with the session reason, forward caller-signal reasons, and reserve "Handoff cancelled" for direct or unreasoned cancellation. Add a regression test for an in-flight handoff aborted through AgentSession.abort.

Fixes #7993
2026-08-08 09:02:28 +00:00
roboomp 92e574cb02 fix(session): surface empty handoff generation as failure not cancel
The #7904 fix stopped masking provider errors as "Handoff cancelled", but
an empty or whitespace-only generation still fell through: whitespace-only
text passed the `!handoffText` guard and produced a bogus handoff, while
empty text returned undefined which the interactive /handoff caller mapped
to "Handoff cancelled" with no detail and no log entry.

Treat empty/whitespace-only output as a real failure: a user-initiated
handoff throws "Handoff generation produced no content" (surfaced as
"Handoff failed: ...") and logs it; auto-handoff keeps returning undefined
so maintenance falls back to context-full compaction. Also log genuine
handoff failures in the command controller so they persist for debugging.

Fixes #7993
2026-08-08 08:21:16 +00:00
enieuwy a5be3ea0df fix(session): carry attribution across a fork
Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.

`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
2026-08-08 15:57:31 +08:00
enieuwy ae5965241e fix(session): count every real output block as a served turn
Codex found the shared attribution predicate recognising only tool calls,
text and signed thinking. A native image response often arrives with no
text and no tool call at all, so an image-only turn was read as producing
nothing: attribution stayed on whichever model spoke before it, and the
empty-stop rule could classify a successful generation as empty.

Everything the assistant can emit now counts except two: unsigned thinking,
which is not provider-authenticated and was already excluded, and
Anthropic's `fallback` marker, which records that a request was routed
elsewhere rather than carrying output. Redacted thinking and server-tool
blocks are real work by the same argument as the image.

The `toolUse` arm keeps its stricter rule — an orphaned toolUse stop needs
a tool_use block to anchor a later tool_result, and an image cannot.

Also switches the new test to the namespace import AGENTS.md requires for
node builtins.
2026-08-08 13:16:42 +08:00