Commit Graph
17 Commits
Author SHA1 Message Date
can1357 6592b799b3 Merge PR #7980: fix(session): attribute a run to the model that produced its output (@enieuwy) 2026-08-11 15:06:12 +02:00
can1357 261eb965ac fix(task): preserved direct role alias fallback 2026-08-11 15:06:11 +02:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
enieuwy 77ee3f2e7e fix(task): key subagent fallback chains off the pre-expansion model role
A single-model subagent is pinned to a `subagent:<id>` role whose
`retry.fallbackChains` entry shadows every configured role chain, so the
chain it inherits decides where the child retries. Inheritance resolved
the role by re-deriving it from the child's `modelPatterns` — but every
spawn path expands the role alias into `modelOverride` before calling
`runSubprocess` (`modelPatterns = normalizeModelPatterns(modelOverride ??
agent.model)`), so `@task` never reached the derivation and it returned
`undefined` every time. Every task subagent inherited `chains.default`.

With `modelRoles.task: anthropic/claude-sonnet-5`, `task` chained to
sonnet alone, and `default` chained to sonnet plus a second provider, a
transient stall on sonnet routed the child onto the default chain's
second model — one the operator had deliberately kept out of the `task`
chain — and a quota error there killed a 28-minute run.

#7694 fixed only the shape where an unexpanded alias reaches the
executor, which no production caller produces; its tests supplied a bare
`agent.model: ["@smol"]` with no `modelOverride`. The incident above
happened on v17.2.10, which contains that fix.

Route inheritance off the role identity the spawn path already computes
and passes as `modelRole`. Since that leaves the pattern-derived operand
unreachable, drop it and the parameter it was the only user of.

The vibe worker path had the same defect independently: `#resolveWorker`
expanded `@task`/`@smol` for the bundled `task`/`sonic` workers and kept
no role, so vibe children inherited `default` no matter what the
executor did. It now carries `modelRole` on `ResolvedVibeWorker` and
`VibeRecord` through both the spawn and rehydrate sites.

To stop the two halves drifting apart again — the mistake that caused
this bug — `resolveAgentModelSelection` returns the expanded `patterns`
and the pre-expansion `role` from one call, and both spawn paths take
both from it. `resolveAgentModelSource` is removed: its only use was
being fed to `resolveExplicitModelRole`, and keeping it invites the same
split derivation. `resolveAgentModelPatterns` stays for the UI callers
that legitimately want patterns alone.

Tests cover the producible shapes: the incident's chain layout (role
chain equal to the primary, default chain a superset), role identity
surviving expansion for every alias-routed bundled agent, and the
patterns/role pairing itself. #7694's two tests are re-anchored to a
shape a real caller produces.
2026-08-07 21:16:27 +08:00
enieuwy b279f06849 fix(task): inherit the aliased role's retry fallback chain for subagents
A single-model subagent has no fallbacks of its own, so it inherits one
and is pinned to a `subagent:<id>` role. That pin is inserted first in
`retry.fallbackChains` so no other role can capture its routing — which
also means it shadows every configured role chain at runtime.

Inheritance was hardcoded to `chains.default`, so a subagent spawned
through a role alias (the bundled scout's `model: "@smol"`) retried on
the default role's chain instead of its own. With `smol` chained to
composer/grok/luna and `default` chained to gpt-5.6-sol, every scout
fell back onto sol.

Resolve the inherited chain from the role identity still present in the
raw pattern (`@smol` -> `smol`), falling back to `default` when that
role configures no chain. An explicitly empty role chain still means
"no fallbacks", mirroring `expandDefaultRetryFallbackChains`. Explicit
model selectors keep inheriting `default`: they carry no role identity,
and a role assigned the same model must not capture the child's routing.
2026-08-05 18:13:42 +08:00
can1357 b3e0bde7b7 refactor(coding-agent): adjusted block resolution formatting and update tests
- Update hashline block resolution formatting to correctly incorporate anchor lines within operation labels.
- Fix and update test assertions and mock contexts across coding agent tests.
2026-07-31 20:55:37 +02:00
pr-evalandcan1357 424458e99d chore(secrets): dropped drive-by changes unrelated to secret placeholders
Reverted branch-side edits to spawn-policy prompts/tests, settings tab
groups, mermaid cache typing, prewalk todo gating, and packages/ai test
churn back to merge-base content; trimmed their changelog entries. These
repaired stale CI against an older main and are stale or conflicting
against current main.
2026-07-23 17:56:29 +02:00
can1357 c0c1622012 Merge PR #4636: feat(secrets): add friendly names to secret placeholders (@Mathews-Tom)
# Conflicts:
#	packages/ai/test/pi-native-client.test.ts
#	packages/coding-agent/src/advisor/runtime.ts
#	packages/coding-agent/src/prompts/tools/eval.md
2026-07-23 17:56:24 +02:00
Jeff Scott Ward 6dcc485389 fix(task): inherit default subagent fallback 2026-07-17 17:43:04 -04:00
can1357 d25d9300ed chore: fixed remaining stale tests after prompt and session api changes
- coordination advisory now points at hub, not irc
- eval description renders 'Allowed:' list instead of removed default-agent prose
- task description no longer renders spawn-policy text; assert agent-list filtering instead
- issue-2750 mock session gained getEnabledToolNames (renamed session api)
2026-07-15 19:50:46 +02:00
Mathews-Tom 63ece0218a test(coding-agent): repair stale CI contracts 2026-07-15 20:56:42 +05:30
roboomp 4409e6cfb5 fix(task): preserved deferred fallback chains
Install subagent retry fallback chains after deferred model patterns resolve so runtime-only candidates keep their ordered fallbacks.\n\nFixes #4421
2026-07-03 09:41:25 +00:00
roboomp 17163a2b9f fix(task): preserved auth fallback for deferred models
Preserve the subagent parent-model auth fallback when an explicit selector resolves only after child runtime provider loading.\n\nFixes #4421
2026-07-03 09:29:11 +00:00
roboomp ddd6ae1182 fix(task): preserved deferred subagent model selectors
Forward unresolved explicit subagent model selectors into child session startup so modelRoles.task cannot disappear during executor preflight and fall through to an unrelated provider default.\n\nFixes #4421
2026-07-03 09:16:30 +00:00
roboomp 3cdb867d28 fix(agent): preserved routed subagent fallbacks
Kept OpenRouter and Vercel upstream routing suffixes in subagent retry fallback selectors so same-base routed candidates stay distinct.

Resolved retry fallback candidates from the raw selector before model switching so routed fallback models keep their requested upstream route.
2026-06-16 07:47:48 +00:00
roboomp 3742e07791 fix(agent): prioritized subagent fallback chains
Placed subagent-scoped fallback chains ahead of inherited retry chains so overlapping primary selectors prefer the explicit subagent model order.

Expanded the regression to pin fallback chain insertion order when a global chain shares the same primary selector.
2026-06-16 07:09:14 +00:00
roboomp 8c9910318d fix(agent): retried subagent model fallbacks
Installed subagent-scoped retry fallback chains from ordered task model candidates so provider failures can advance to the next configured worker model.

Updated task result tracking to surface the fallback-applied final model and added focused regression coverage for the executor wiring.

Fixes #2750
2026-06-16 06:48:55 +00:00