Commit Graph

1384 Commits

Author SHA1 Message Date
can1357 e8eed95130 feat(coding-agent): introduced fine grained per agent advisor configuration
- Replaced blanket subagent advisor global settings with fine-grained per-agent configuration and frontmatter support.
- Added dashboard keybindings and inline override editors for managing agent advisor patterns.
- Implemented settings migration logic to convert legacy global options into per-agent settings.
- Updated session persistence and execution layers to restore and enforce per-agent advisor behaviors.
2026-08-13 04:59:50 +02:00
can1357 3e28a1afb4 feat(coding-agent): implemented threshold for demoting interrupted thinking
- Added a 60-character minimum length threshold for demoting interrupted thinking into hidden continuity context in agent-session.ts.
- Updated LLM message conversion in messages.ts to strip incomplete thinking from user-interrupted assistant turns regardless of continuity note presence.
- Added test coverage in agent-session-interrupted-thinking.test.ts verifying behavior for reasoning lengths below and at the threshold.
2026-08-13 02:16:08 +02:00
can1357 61e12988fe Merge PR #8283: fix(coding-agent): skip idle mid-turn persistence waits (@ethancawse) 2026-08-13 02:00:50 +02:00
can1357 2dddba9fe4 Merge PR #8199: fix(coding-agent): gate prompt tool mentions on session tool availability (@szavadsky)
# Conflicts:
#	packages/coding-agent/src/prompts/system/orchestrate-notice.md
#	packages/coding-agent/src/prompts/system/plan-mode-active.md
#	packages/coding-agent/src/prompts/system/project-prompt.md
#	packages/coding-agent/src/prompts/system/system-prompt.md
2026-08-13 01:18:42 +02:00
can1357 eafefaccf7 Merge PR #8337: fix(usage): use authoritative OpenCode Go quotas (@will-bogusz) 2026-08-13 01:14:53 +02:00
can1357 932bb6d244 Merge PR #8080: fix(session): forward retry fallback events to extensions (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 a5434d7341 Merge PR #8076: fix(advisor): preserve advisor fallback role ownership (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 abc1897fc3 fix(agent): fit retry fallback against retry prompt 2026-08-13 01:14:46 +02:00
can1357 f73baac583 Merge PR #8068: fix(agent): fit-check retry fallback before switching models (@roboomp) 2026-08-13 01:14:46 +02:00
Will Bogusz 9eff02d36c fix(usage): use authoritative OpenCode Go quotas 2026-08-12 13:26:01 +02:00
Ethan Cawse 7b0d12ee63 fix(coding-agent): detach noncloneable message metadata 2026-08-11 23:20:41 -04:00
Ethan Cawse 7ab6b554f6 fix(coding-agent): isolate message-end notifications 2026-08-11 23:03:11 -04:00
Ethan Cawse a457173750 fix(coding-agent): preserve persistence after listener failures 2026-08-11 23:03:10 -04:00
can1357 19c0afcc0d feat: implemented external thinking flags and transport reasoning controls
- Added the `--external-thinking` CLI flag alongside model capability checks to gate external thinking tool availability.
- Updated Anthropic and Google transports to honor `forceReasoningOff` for native thinking-off controls.
- Renamed the `thoughts` property and parameter to `notes` across think fixtures, tools, and tests.
- Updated system prompt instructions and test suites to verify transport-specific thinking and tool activation.
2026-08-12 02:23:17 +02:00
can1357 10fd42289c feat: introduced external thinking support and private scratchpad think tool
- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
2026-08-11 20:39:57 +02:00
can1357 47b282ff9b Merge PR #8069: fix(extensions): register lifecycle tools (@mrexodia) 2026-08-11 15:24:18 +02:00
can1357 ca6e13fd47 Merge PR #8218: fix(agent): park subagents on shutdown (@roboomp) 2026-08-11 15:18:17 +02:00
can1357 7e8be71ead Merge PR #7247: fix(advisor): prevent false full replays and preserve cache growth (@cuipengfei) 2026-08-11 15:07:43 +02:00
can1357 6592b799b3 Merge PR #7980: fix(session): attribute a run to the model that produced its output (@enieuwy) 2026-08-11 15:06:12 +02:00
roboomp 546016596e fix(agent): scope shutdown reason to the owning session
Every session dispose broadcast the shutdown abort reason, so an explicit
hard kill of a subagent (release with tombstone, then live.dispose()) tagged
its nested children as shutdown and rediscovered them as parked instead of
terminal.

- Gate ASYNC_JOB_MANAGER_SHUTDOWN_REASON on #ownedAsyncJobManager so only the
  top-level owning session's dispose (genuine process shutdown) uses it.
- Subagent disposes propagate a generic cancellation, keeping nested children
  terminal.
- Cover the subagent generic-cancel path alongside the owning-session shutdown.

Fixes #8216
2026-08-11 06:40:44 +00:00
roboomp 8614979281 fix(agent): tag owned jobs with shutdown reason before dispose
Session teardown pre-cancels owner jobs via #cancelOwnAsyncJobs before
manager.dispose(), so the shutdown abort reason must ride along that
cancelAll or an owned subagent job sees a generic caller signal and is
tombstoned instead of parked.

- Forward an abort reason through AsyncJobManager.cancelAll.
- Pass ASYNC_JOB_MANAGER_SHUTDOWN_REASON from #disposeOwnedAsyncJobs.
- Cover owned-job shutdown tagging through a real AgentSession dispose.

Fixes #8216
2026-08-11 06:32:46 +00:00
Slava Zavadsky b6a3862ebc fix(coding-agent): gate prompt tool mentions on session tool availability
System and mode prompts referenced tools that may be absent from the
session catalog, forcing the model to satisfy requirements it cannot
execute (generalizes #8139's browser-verification mismatch):

- system-prompt.md: browser verification now keys on the actual UI
  surface and available tools (browser/computer/TUI/CLI), with an
  explicit behavioral/smoke-test fallback when no runtime tool exists;
  todo workflow guidance and the AST-section grep hint are gated on
  tool presence; the auto-QA report_issue block additionally requires
  the write tool.
- project-prompt.md: workspace-tree drill-in and additional-roots tool
  hints only name tools present in the session.
- plan-mode-active.md: ask-tool directives get a prose fallback and
  scout-via-task dispatch is gated on the task tool (render site now
  passes askAvailable/taskAvailable).
- orchestrate-notice.md: the tool budget, verify gates, todo tracking,
  and inline-edit guidance are gated on tool presence; the notice is
  skipped entirely when the task tool is inactive (render fn takes the
  active tool list).

Adds regression tests for each gate; changelog entry.
2026-08-10 20:24:34 -04:00
Duncan Ogilvie 158a70ad54 fix(extensions): preserve MCP ownership on refresh 2026-08-10 02:09:23 +02:00
Duncan Ogilvie d1f73de44c fix(extensions): serialize registry mutations 2026-08-10 01:17:23 +02:00
Duncan Ogilvie 4d0c346f87 fix(extensions): abort timed-out tool activations 2026-08-10 00:54:42 +02:00
Duncan Ogilvie 2c55e20365 fix(extensions): expose prompt refresh option 2026-08-09 22:13:11 +02:00
Duncan Ogilvie 05f17fa76f fix(extensions): make lifecycle registration atomic 2026-08-09 22:07:46 +02:00
roboomp ffd9d5c8ae fix(session): forward retry fallback events to extensions
retry_fallback_applied and retry_fallback_succeeded were emitted to the
TUI and RPC subscribers but never reached extensions: AgentSession's
#emitExtensionEvent had no branch mapping either event to
ExtensionRunner.emit, and ExtensionEvent / ExtensionAPI.on lacked the
types and overloads, so registration was also rejected at compile time.

Add the typed events, on() overloads, union members, and the two
forwarding branches so extensions observe the same { from, to, role } /
{ model, role } payloads as TUI and RPC. Matches the contract already
documented in docs/non-compaction-retry-policy.md.

Fixes #8079
2026-08-09 14:20:24 +00:00
roboomp 121bcb3663 fix(advisor): preserved advisor role fallback ownership
- Passed the known advisor role through retry fallback resolution so model and wildcard keys retain precedence while ambiguous role matches cannot select another role.

- Covered shared-model roles with distinct thinking levels and asserted advisor fallback lifecycle ownership.

Fixes #8075
2026-08-09 13:48:46 +00:00
roboomp 7b6548f182 fix(agent): fit-check retry fallback before switching models
Retry-fallback candidate selection filtered on suppression, effort
ceiling, model resolution, and API key, but never compared a
candidate's context window with the live context. A large-window
primary hitting a retryable error could switch onto a smaller-window
fallback and immediately send a predictably oversized request that the
provider rejects, stalling the run. This is the forward counterpart of
the #7952 cooldown-expiry revert fix.

Generalize the existing retry-fit budget check into
SessionMaintenance.contextFitsModel(model) and consult it from both
retry-fallback selection loops (#tryRetryModelFallback and the
usage-aware loop): skip any candidate whose usable window cannot hold
the current context and advance to the first configured candidate that
fits. The check is independent of compaction.enabled since an oversized
request overflows regardless.

Fixes #8065
2026-08-09 11:03:25 +00:00
can1357 d85dde52c4 fix(session): fence in-flight disk work behind the terminal seal
- seal() now runs before the final dispose close() and bumps the disk
  epoch: work an event handler enqueues while dispose awaits the closing
  tail is superseded, and an already-running fenced or authoritative
  rewrite fails its commit guard at the rename fence instead of
  publishing over a file a revival reopened.
- The authoritative repair path resets the disk tail itself, escaping the
  close() serialization, and would atomically publish the emptied entry
  list; it now no-ops once sealed and commit guards also check the seal.
- Execution-time gates cover the queued title persist and fenced rewrite
  callbacks; setSessionName/appendCustomEntry attempted by a handler that
  outlives dispose are dropped and covered by the seal regression, and a
  failed atomic batch across the seal can no longer truncate the file.
2026-08-08 20:49:28 +02:00
can1357 63aa8cf6f8 fix(session): seal the manager at terminal release against revival races
- AgentLifecycleManager.park() resolves as soon as dispose() returns, so
  ensureLive() may reopen the same JSONL through a new manager while a
  timed-out event handler still holds the old one; a late append reopened
  a second writer on that file and the deferred finalize then closed it.
- A post-release rewrite was worse: it persisted the emptied entry list,
  truncating the transcript on disk.
- releaseRetainedEntries() now seals the manager: appends, title changes,
  and rewrites become dropped no-ops and the append writer is closed, so
  the deferred dispose pass is in-memory only and can never touch the file.
- File-backed regression: dispose on the drain deadline, revive the JSONL
  immediately, unpark the late handler, and prove the file is byte-stable
  and the revival writer owns it exclusively.
2026-08-08 20:04:09 +02:00
can1357 31d7655477 fix(agent): re-finalize dispose after the drain deadline
- The dispose drain deadline does not cancel in-flight event handlers: one
  parked in a slow extension hook resumed after close/release, reopened the
  append writer for its late persist, and repopulated the released state.
- Track whether the drain settled; on deadline, redo the final close +
  release once the pipeline genuinely settles (hook runtime is bounded by
  the extension runner).
- Expose drainTimeoutMs on AgentSessionDisposeOptions for bounded teardown
  paths and deterministic coverage of the deadline branch.
2026-08-08 19:54:33 +02:00
can1357 f87a8ecc19 Merge PR #7994: fix(session): surface empty handoff generation as failure not cancel (@roboomp) 2026-08-08 19:38:32 +02:00
can1357 b2d0ade665 fix(agent): finish session disposal after event drain 2026-08-08 19:38:32 +02:00
roboomp 88021b90ea fix(agent): drained in-flight session event handlers on dispose
agent-core dispatches the session's event subscriber fire-and-forget (agent.ts #emit), so a message_end/agent_end handler can still be awaiting extension/subscriber/maintenance work — and its sessionManager/agent.state append — after agent.waitForIdle() resolves. The earlier settle waited only on the core run, so a late handler could append the finished message/entries back into the disposed session and re-pin the transcript.

Track every #handleAgentEvent dispatch in #inFlightEventHandlers and drain it (alongside agent.waitForIdle) inside the bounded settle before reset/clear/release. Added a regression test using a real extension whose message_end hook blocks before persistence, asserting dispose does not release memory until the in-flight handler settles.
2026-08-08 10:26:41 +00:00
roboomp a7a32f35dd fix(agent): settled active turn before clearing session memory
dispose() only *signalled* the agent loop via abort(); it never awaited the run, so a mid-turn dispose (Ctrl-C/timeout/hard-killed subagent) could let the loop unwind after the release ran — its response/SSE interceptors re-recording wire frames into rawSseDebugBuffer and its terminal message re-appending to agent.state.messages, repopulating the disposed session with exactly the retained state the release drops.

Detach the response/SSE interceptors and await a bounded agent.waitForIdle() before the reset/clear so it lands on a quiescent session. Added a deterministic regression test that gates the active turn and asserts dispose blocks on it before clearing.
2026-08-08 10:06:55 +00:00
roboomp 4ca6c376b4 fix(agent): detached append-only context on dispose
Disposed sessions remained reachable through lifecycle reviver closures. Agent.reset() cleared the live message array but left AppendOnlyContextManager attached, retaining its normalized provider transcript and stable prompt/tool prefix.

Detach the append-only manager during terminal disposal and extend the memory-release regression test to cover that second transcript copy.
2026-08-08 09:52:15 +00:00
roboomp ed6300b35e fix(agent): release parked subagent session memory on dispose
Keep-alive subagents are handed to AgentLifecycleManager.adopt, which stores
their reviver closure in the process-global #adopted map. The closure is
defined inside runSubagent's scope, which also captures the live AgentSession
(extension-runner callbacks, buildSubagentSessionOptions), so its lexical
environment pins the whole session graph. park() disposes and detaches the
session but leaves the adoption record indefinitely, and #doDispose never
dropped the in-memory transcript, session-manager entries, or the raw-SSE
debug buffer (whose trimmed records retain slice() views of full wire frames),
so every completed subagent's heavy state leaked for the process lifetime.

dispose() is terminal and every revival path reopens from disk, so #doDispose
now sheds retained conversation memory via agent.reset(),
RawSseDebugBuffer.clear(), and SessionManager.releaseRetainedEntries(). The
adoption record can still reference the session, but only as a husk.

Fixes #8003
2026-08-08 09:42:35 +00:00
roboomp 7914e7c451 fix(session): preserved harness handoff abort reasons
Harness-initiated session aborts previously cancelled compaction before the handoff reason was recorded. The handoff catch then saw only an aborted signal and replaced the harness reason with "Handoff cancelled".

Abort the handoff first with the session reason, forward caller-signal reasons, and reserve "Handoff cancelled" for direct or unreasoned cancellation. Add a regression test for an in-flight handoff aborted through AgentSession.abort.

Fixes #7993
2026-08-08 09:02:28 +00:00
enieuwy a5be3ea0df fix(session): carry attribution across a fork
Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.

`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
2026-08-08 15:57:31 +08:00
enieuwy a1e60c3450 refactor(session): fold the pending-fallback surface into servingModel
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.

Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.

The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.

`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
2026-08-08 13:16:41 +08:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
can1357 0697e7f688 refactor(coding-agent): split secret obfuscator into domain modules
- Separated deterministic replacement generation, placeholder derivation,
  placeholder-range scanning and message-tree transforms out of the 2647-line
  module; obfuscator.ts now holds the types and SecretObfuscator.
- ephemeralPlaceholderKey stays a single instance and both global regexes stay
  beside the code that resets their lastIndex, so placeholder stability and
  the security argument in the moved comments are preserved verbatim.
- Repointed every importer at the real modules rather than leaving a re-export
  shim; the public ./secrets barrel exports the same 15 names as before.
2026-08-08 06:32:01 +02:00
roboomp 7d4b2e7998 fix(agent): re-check context on cooldown-expiry revert in auto-continue path
A cooldown-expiry model revert runs at a turn boundary. The user-prompt
path reverts then re-checks accumulated context against the restored
model via runPrePromptCompactionIfNeeded, but the automatic
agent.continue() path (#scheduleAgentContinue) reverted and issued the
next request with no such check. When a transient failure had fallen
back to a larger-window model and the conversation then grew past the
original model's window, restoring the primary once its cooldown expired
sent a predictably oversized request to the smaller model.

maybeRestoreRetryFallbackPrimary now reports whether it actually
switched, and the auto-continue path runs the same post-revert
context-fit maintenance (compaction/promotion) the prompt path already
runs, but only when a revert occurred.

Fixes #7952
2026-08-07 23:38:24 +02:00
can1357 a8a8188e4b Merge PR #7756: fix(coding-agent): preserve before_agent_start prompt override across base rebuilds (@roboomp) 2026-08-07 13:39:53 +02:00
can1357 5f758778b8 Merge PR #7785: fix(coding-agent): make prewalk lifecycle one-shot (@eggpeat) 2026-08-07 13:39:51 +02:00
can1357 68dece477f Merge PR #7772: fix(session): handle subscription-cap retry exhaustion (@roboomp) 2026-08-07 13:37:54 +02:00
Brent 2efe8896ea fix(coding-agent): make prewalk lifecycle one-shot 2026-08-06 19:57:10 +00:00
roboomp f6c5a43a1f fix(session): handled subscription-cap retry exhaustion
- Classified subscription and plan rate caps as credential-rotatable usage limits while preserving transient per-minute throttles.

- Applied reason-specific backoff to transient rate limits and collapsed exhausted retry attempts behind one budget-labeled terminal error.

- Covered classification, delay selection, persisted transcript aggregation, and retry event propagation.

Fixes #7767
2026-08-06 01:31:22 +00:00