102 Commits

Author SHA1 Message Date
can1357 8c35861d96 feat(coding-agent): implemented ordered compaction fallback and settings
- Replaced legacy `compaction.strategy` and `remoteEnabled` settings with `compaction.methodOrder` across session maintenance, schema, and tests.
- Added automatic fallback mechanism to try subsequent compaction methods upon failure or unsupported model capabilities.
- Added mouse drag-and-drop reordering support and click handlers to multi-select settings submenus.
- Updated documentation and test suites to reflect ordered compaction strategy preferences and fallback chains.
2026-08-20 02:59:49 +02:00
roboomp 412d0e3b42 fix(coding-agent): keep thinking-loop retries on the same model
A ThinkingLoop abort is the loop guard asking for a same-model resample
(it injects a thinking-loop-redirect notice that only makes sense on the
model that looped), not a provider failure. #handleRetryableError routed
it through the generic retryable-error branch, so on attempt 1 it called
noteRetryFallbackCooldown + #tryRetryModelFallback and could switch to
another family from retry.fallbackChains while parking the original
selector on a 5-minute cooldown. A healthy Grok 4.6 planning turn got
replaced by whatever the chain listed next.

Carve ThinkingLoop out of the model-fallback branch and out of the
Fireworks Fast->base degrade so the loop guard always re-samples the same
model; the retry budget still bounds a genuinely stuck stream.

Fixes #8760
2026-08-16 21:16:35 +00:00
roboomp 77336d880c fix(coding-agent): blocked model-bound Anthropic fallback
Skipped same-provider cross-model fallback candidates when the latest assistant turn contains signed or redacted Anthropic thinking.

Kept same-model retry available for transient failures and added regression coverage.

Fixes #8558
2026-08-14 14:40:20 +00:00
roboomp c3b1032c3c fix(coding-agent): stopped invalid Anthropic thinking fallback
Prevented deterministic 400 errors for immutable thinking blocks from entering same-model retries or configured model fallback.

Added a signed-thinking session regression covering the terminal error and retry UI lifecycle.

Fixes #8558
2026-08-14 14:31:18 +00:00
usr-bin-roygbiv fe33232298 fix(gemini): preserve retry and rendering boundaries 2026-08-14 02:48:18 +00:00
usr-bin-roygbiv b1ce77c109 fix(gemini): recover malformed and thought-only turns 2026-08-14 01:35:21 +00:00
can1357 28997c0d46 refactor(coding-agent): restructured transcript rendering during initialization
- Stage transcript initialization inside a detached TranscriptContainer to keep existing messages visible during incremental rendering.
- Add fallback state restoration in InteractiveMode.renderInitialMessages when chat rendering is aborted or fails.
- Update render-initial-messages tests to assert that old transcripts remain visible until replacements are fully committed.
2026-08-13 18:55:36 +02:00
can1357 23aff5ae3f test(tui): model direct HerdR in-place resize in stress oracle
The HerdR flicker fix routed direct HerdR panes onto the in-place
multiplexer resize path, but the randomized render-stress oracle still
modeled them as ED3 replaying direct terminals. The width-epoch ledger
now applies to any in-place-resize scenario (multiplexer + direct HerdR)
via a dedicated resizeRepaintsInPlace trait, keeping HerdR's direct
scrollback semantics intact. Adds a deterministic replay regression
(seed 0xcafed00d, 24 iterations) that fails without the oracle update.
2026-08-13 18:39:29 +02:00
can1357 6b4823181b test: cleaned test suites and documented filtering guidelines
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
2026-08-13 08:28:42 +02:00
roboomp 04a0b44d6a fix(tests): align test fixtures with low/high/max ladder for deepseek-v4-pro (#8405)
Updates test assertions and fixtures to reflect that V4 Pro now exposes the full [low, high, max] effort ladder, switches the incompatible-fallback test to the openrouter deepseek-v4-pro entry, and adds a shared browser lease in the evaluation suite to avoid relaunching Chromium per test under load.
2026-08-13 08:28:42 +02:00
can1357 932bb6d244 Merge PR #8080: fix(session): forward retry fallback events to extensions (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 a5434d7341 Merge PR #8076: fix(advisor): preserve advisor fallback role ownership (@roboomp) 2026-08-13 01:14:47 +02:00
can1357 abc1897fc3 fix(agent): fit retry fallback against retry prompt 2026-08-13 01:14:46 +02:00
can1357 f73baac583 Merge PR #8068: fix(agent): fit-check retry fallback before switching models (@roboomp) 2026-08-13 01:14:46 +02:00
roboomp ffd9d5c8ae fix(session): forward retry fallback events to extensions
retry_fallback_applied and retry_fallback_succeeded were emitted to the
TUI and RPC subscribers but never reached extensions: AgentSession's
#emitExtensionEvent had no branch mapping either event to
ExtensionRunner.emit, and ExtensionEvent / ExtensionAPI.on lacked the
types and overloads, so registration was also rejected at compile time.

Add the typed events, on() overloads, union members, and the two
forwarding branches so extensions observe the same { from, to, role } /
{ model, role } payloads as TUI and RPC. Matches the contract already
documented in docs/non-compaction-retry-policy.md.

Fixes #8079
2026-08-09 14:20:24 +00:00
roboomp 121bcb3663 fix(advisor): preserved advisor role fallback ownership
- Passed the known advisor role through retry fallback resolution so model and wildcard keys retain precedence while ambiguous role matches cannot select another role.

- Covered shared-model roles with distinct thinking levels and asserted advisor fallback lifecycle ownership.

Fixes #8075
2026-08-09 13:48:46 +00:00
roboomp 7b6548f182 fix(agent): fit-check retry fallback before switching models
Retry-fallback candidate selection filtered on suppression, effort
ceiling, model resolution, and API key, but never compared a
candidate's context window with the live context. A large-window
primary hitting a retryable error could switch onto a smaller-window
fallback and immediately send a predictably oversized request that the
provider rejects, stalling the run. This is the forward counterpart of
the #7952 cooldown-expiry revert fix.

Generalize the existing retry-fit budget check into
SessionMaintenance.contextFitsModel(model) and consult it from both
retry-fallback selection loops (#tryRetryModelFallback and the
usage-aware loop): skip any candidate whose usable window cannot hold
the current context and advance to the first configured candidate that
fits. The check is independent of compaction.enabled since an oversized
request overflows regardless.

Fixes #8065
2026-08-09 11:03:25 +00:00
enieuwy a5be3ea0df fix(session): carry attribution across a fork
Codex found the session-id anchor too blunt. It assumes a new id means an
unrelated transcript, which holds for `/new` and for resuming something
else — but `fork()` mints a fresh id while cloning the transcript and
keeping the same recovery state running. Attribution and routing both
expired there, so immediately after `/fork` an unproven fallback
bootstrapped as the current model with `isFallback: false`: the run was
re-credited to a model that never produced any of it, and mislabelled as
the configured primary. Exactly the bug the anchor exists to prevent,
reopened for the one switch that is a continuation.

`AgentSession.fork()` now re-tags both onto the new id after the fork
succeeds, moving only state that belonged to the pre-fork id so an id left
behind by an earlier switch stays expired.
2026-08-08 15:57:31 +08:00
enieuwy 837b1ac0e2 fix(session): import ServingModel, anchor fallback routing to its session
Review found the branch does not compile. `bun check` runs biome before
the per-workspace type check and joins them with `&&`, so a pre-existing
format error on `main` short-circuited the run: `tsgo --noEmit` never
executed, and `bun test` type-strips, so the suite stayed green over eight
type errors. `ServingModel` was used in `turn-recovery.ts` without being
imported, and three `subscribe` closures in the retry-fallback suite lost
the outer narrowing of `session`. Verified now against the workspace check
directly rather than the aggregate.

Codex also found `#fallbackRouted` outliving its session. It said how the
CURRENT model was reached, but nothing reset it when a transcript was
switched or resumed, so a freshly loaded session with no served attribution
yet described its own model with the previous session's routing. It is now
anchored on the session id exactly like the attribution beside it — the two
facts are earned together and expire together — and a session switch
reports the current model without claiming to know how it got there.

The replay-unsafe host stub gained `getSessionId`, which the real
`SessionManager` has always had and the anchoring now calls.
2026-08-08 13:16:42 +08:00
enieuwy a1e60c3450 refactor(session): fold the pending-fallback surface into servingModel
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.

Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.

The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.

`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
2026-08-08 13:16:41 +08:00
enieuwy a06a7f5ef1 refactor(session): anchor attribution to its transcript, drop a dead format branch
Two cleanups from review.

The offline walk matched a served turn to its `model_change` by also
accepting a `:level` thinking suffix, but no writer produces one: every
`appendModelChange` call site records a bare `provider/id` or a
`formatModelStringWithRouting` selector, and only the latter decorates —
with `@upstream`, never a level. Speculative matching in the part of the
change that already reconstructs the most from raw strings. Removed, with
its test narrowed to the format writers actually emit.

Attribution was also retained for the life of the recovery component, so
switching sessions in place could report a model from a different
transcript by name. It is now tagged with the session file it was earned
in and ignored when that no longer matches — the same self-invalidating
anchor `#ensurePersistedMessageKeys` uses, so no mutation call site has
to remember to clear anything. Anchoring on the file rather than the leaf
keeps attribution across compaction and appends, which stay within the
transcript that earned it.
2026-08-08 13:16:11 +08:00
enieuwy a93a53185b fix(session): key attribution off what served, not off the fallback chain
Review of the previous commit found the same defect in four more places,
each a variant of one mistake: the unproven-model gate was expressed as
"a retry fallback is armed and has not served", so it only protected the
one swap path that arms a chain.

- The cooldown-expiry restore to the primary deletes the fallback record
  before swapping, so the gate vanished with it and an unproven, just-
  restored primary was credited immediately.
- A session's first-ever fallback arm constructs the record only after
  the swap, so there was nothing to mark and the guard was a no-op.
- The Fireworks Fast degrade swaps models without arming a chain at all,
  so a live row showed the degraded base model as the configured primary.
- The fresh-versus-resumed heuristic read the message list, which
  compaction collapses, so a resumed-and-compacted session seeded its
  startup candidate as already proven.

Attribution now names the last model that settled a turn here, full stop.
A switch of any kind — into a candidate, back to a restored primary, or a
capability degrade — moves it only once the new model answers. Before
anything has served there is no earlier work to miscredit, so the
configured model is both the only answer and a safe one.

Fallback routing is tracked as its own flag rather than inferred from the
chain record, because the degrade path routes without a chain, and it is
set before each swap so the synchronous `model_changed` fan-out cannot
observe a model mid-relabel. The resumed-session heuristic is gone: it
was only needed to decide whether an unserved candidate could be trusted,
and nothing is trusted before it serves.
2026-08-08 13:16:11 +08:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
can1357 38b61ae342 fix(session): honored explicit retry-after over reason backoff
- A provider-supplied retry-after now bypasses the transient rate/concurrency
  heuristic window instead of being overridden by it (regression from the
  subscription-cap retry change).
- Updated event-controller/ui-helpers test doubles for provenance-gated
  renderer selection (hasBuiltInTool), aggregated retryErrors on
  auto_retry_end, and Bedrock override compat gaining streamIdleTimeoutMs.
2026-08-07 23:38:25 +02:00
roboomp 7d4b2e7998 fix(agent): re-check context on cooldown-expiry revert in auto-continue path
A cooldown-expiry model revert runs at a turn boundary. The user-prompt
path reverts then re-checks accumulated context against the restored
model via runPrePromptCompactionIfNeeded, but the automatic
agent.continue() path (#scheduleAgentContinue) reverted and issued the
next request with no such check. When a transient failure had fallen
back to a larger-window model and the conversation then grew past the
original model's window, restoring the primary once its cooldown expired
sent a predictably oversized request to the smaller model.

maybeRestoreRetryFallbackPrimary now reports whether it actually
switched, and the auto-continue path runs the same post-revert
context-fit maintenance (compaction/promotion) the prompt path already
runs, but only when a revert occurred.

Fixes #7952
2026-08-07 23:38:24 +02:00
can1357 e9888367d1 refactor: migrated packages to internal utility modules and removed external dependencies
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
2026-08-05 13:39:09 +02:00
Brent 2da96d77a7 fix(coding-agent): close fallback lifecycle races 2026-08-03 17:51:32 +00:00
Brent 353bbc034b fix(coding-agent): harden usage fallback review races 2026-08-03 17:15:57 +00:00
Brent db97103c32 fix(coding-agent): complete usage-aware fallback integration 2026-08-03 16:34:35 +00:00
can1357 832d4f1506 Merge PR #6987: fix(coding-agent): treat streamed visible text as replay-unsafe in turn recovery (@metaphorics) 2026-07-29 23:08:51 +02:00
metaphorics 93819dbec4 fix(coding-agent): gate refusal retries on replay safety
(cherry picked from commit 531b25beffa2096c76adff6692f3697e649f0e95)
2026-07-29 23:08:51 +02:00
metaphorics 9733d13827 fix(coding-agent): make replay safety authoritative
(cherry picked from commit 118ecabb22f406afbd56864ff0fbafc8f4f1fa4c)
2026-07-29 23:08:50 +02:00
metaphorics b5602ddfc1 fix(coding-agent): treat streamed visible text as replay-unsafe in turn recovery
(cherry picked from commit 3aeac1a4afa756e9c7d579246d8f334cd0731cda)
2026-07-29 23:08:50 +02:00
Paolo Mazzitti 3c7233af44 fix(coding-agent): scope advisor cost to active session 2026-07-29 07:18:48 +00:00
can1357 8c5dc16344 fix(task): enforced per-spawn effort ceiling across retry fallbacks
- task.maxEffort only clamped the initial thinking level; a retry
  fallback candidate could clamp back up to its model floor and run a
  low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
  ModelControls (constructor, setThinkingLevel, auto classifier,
  restore) and in applyRetryFallbackCandidate; fallback candidates whose
  floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
  attribution added.
- Review follow-up for PR #6794.
2026-07-27 16:09:28 +02:00
can1357 12a2b13e93 fix(session): unwedged auto-retry after assistant-tail removal miss
A context rebuild that recreated the failed turn's message object made the
identity-keyed active-context removal miss, so the scheduled retry
continuation rejected the terminal assistant error message locally
("Cannot continue from message role: assistant") before any provider
request. auto_retry_end never fired, retryPromise stayed pending, and the
in-flight prompt() plus the TUI retry indicator hung until a manual
follow-up.

The retry path now strips a still-failed assistant tail positionally after
the backoff (generation-guarded, never in preserveFailedTurn mode), and a
continuation that still fails locally closes the retry saga with a failed
auto_retry_end via the new scheduleAgentContinue onError hook.

Fixes #5382
2026-07-27 04:59:31 +02:00
Brent 8a9630c031 fix(coding-agent): reselect fallback account by health 2026-07-24 02:23:51 +02:00
Brent 0bfb0bd1e3 feat(coding-agent): add usage-aware model fallback 2026-07-24 02:23:51 +02:00
can1357 3829bff31a Merge PR #4637: feat(notifications): add error turn notifications (@Mathews-Tom)
# Conflicts:
#	packages/coding-agent/src/modes/controllers/event-controller.ts
#	packages/coding-agent/src/prompts/tools/eval.md
#	packages/coding-agent/src/session/agent-session.ts
#	packages/coding-agent/src/task/index.ts
2026-07-23 17:49:06 +02:00
can1357 db3a6a1407 Merge PR #6318: fix(tui): show fallback models in Agent Hub (@roboomp) 2026-07-23 11:37:13 +02:00
can1357 e4de9509c7 test: adapt catalog-pinned tests to regenerated model catalog
- Repointed tests off removed gpt-5.2/5.3 codex variants and devin models (e06ac0b787): context-promotion and TTSR tests pin gpt-5.5 -> gpt-5.6-sol via per-test modelOverrides since no bundled codex model has a runtime-effective promotion target anymore; replay-boundary and history-payload suites use gpt-5.5; advisor quota fallback uses devin/swe-1-6-slow with suffix-less selectors per #4579 devin-agent semantics; gateway-reference pins kilo/giga-potato and now asserts the reference carries effortRouting so the cross-provider no-inherit contract stays meaningful.
- Verified cross-provider gateway references do not inherit wire routing after variant collapse (identity/reference.ts:145-154) - stale test, no product bug.
2026-07-22 22:26:10 +02:00
roboomp f9a56ff94f fix(tui): showed fallback models in agent hub
- Exposed the active retry fallback selector from live agent sessions.
- Rendered fallback rows with an explicit marker and resolved provider/model.
- Added an end-to-end fallback-to-Agent-Hub regression assertion.

Fixes #6316
2026-07-22 19:08:32 +00:00
roboomp b257a6dcbf fix(session): preserved startup fallback ownership
- Carried startup-selected fallback role and primary selector into AgentSession.
- Continued remaining role fallback entries after the startup fallback fails.
- Added regression coverage for chained startup failover.

Fixes #6283
2026-07-22 11:11:30 +00:00
can1357 c655db4e3c fix(session): added #prunedTerminalRefusal field to store the
- Added `#prunedTerminalRefusal` field to store the pruned refusal for post-settle consumers.
- Modified `getLastAssistantMessage()` to return the pruned refusal before active-context lookup.
- Reset `#prunedTerminalRefusal` on `agent_start` so a fresh run supersedes the settled refusal.
- Updated test mock helpers to include `getLastAssistantMessage` for consistency.
2026-07-18 17:51:48 +02:00
can1357 eac51b6a04 Merge: darkphilosophy/feat/advisor-per-agent-toggle
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:

- #failing latch: waitForCatchup resolves immediately while an advisor is
  mid-failure; parked waiters wake the moment a turn fails, before any
  async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
  dedup snapshot and never propagates into the primary's turn-end callback
  (per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
  runtime latches quotaExhausted, requeues the batch, and notifies —
  cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
  before any assistant turn is recorded.
2026-07-17 07:37:29 +02:00
can1357 6f42a4375f merge PR #5748 via eval/pr-5748: fix(advisor): apply configured fallback chains 2026-07-17 04:45:46 +02:00
can1357 1424cae066 feat(coding-agent): added id-prefixed wildcard support to retry fallback chains
- Implemented parsing of id-prefixed wildcard keys and entries, allowing provider-specific prefixes in retry fallback configuration.
- Added logic to re-prefix failing model IDs and to match id-prefixed keys, with validation of provider existence.
- Updated settings schema description and changelog, and added tests covering the new behavior.
2026-07-17 03:49:43 +02:00
roboomp 203b959056 fix(advisor): restored primary after fallback cooldown
- Retained the advisor's original selector and thinking level while progressing through fallback candidates.
- Restored the configured primary before later advisor turns once its selector cooldown expired.
- Covered quota fallback restoration under the default cooldown-expiry policy.
2026-07-16 20:08:51 +00:00
roboomp 777e5e0982 fix(advisor): applied configured fallback chains
- Switched advisor turns to the next configured model after provider quota or rate-limit failures.
- Emitted fallback applied and succeeded lifecycle events without reporting advisor unavailability after recovery.
- Added an end-to-end advisor quota fallback regression test.

Fixes #5740
2026-07-16 19:56:24 +00:00
Mathews-Tom ea3882f8a3 fix(coding-agent): close terminal retry lifecycle 2026-07-15 22:53:47 +05:30