Treat content-less advisor stop completions as failed turns so the advisor retry/drop path handles silent provider responses instead of accepting them as successful reviews.
Fixes#5212
Blockers from final review:
- runtime.ts:457 TS2741: #collectAndMaintainBatch now returns wip in its
result type; #drain destructures it and passes it to the retry-requeue
unshift — a WIP batch that fails retry no longer silently loses its
[in progress] heading.
- agent-session.ts import order: annotateForStaleness moved before
formatAdvisorBatchContent (biome enforces case-insensitive alpha order).
Cap edge case (advisor nit + review CONCERN): final round of the
coalescing for-loop now breaks BEFORE the late-item splice so any
deltas that arrived during round MAX_COALESCE_ROUNDS-1's
maintainContext call stay in #pending for the next drain iteration
instead of being merged into an unbudgeted batch. Doc comment updated
to match ('left for the next drain iteration' is now accurate).
Cap test: bounded version that only pushes new turns for the first 3
maintainContext calls so the drain while-loop terminates cleanly after
a second iteration. The previous unbounded version created an infinite
drain loop (each maintainContext call unconditionally pushed another
turn) and timed out.
Blocker: wip field added to PendingDelta and threaded into the reprime
path. #collectAndMaintainBatch now captures the most-recent WIP state
from each delta batch and forwards it to #renderDelta in both the
normal and reprime branches, so a willContinue:true turn never loses
its [in progress] heading through a reprime.
Safety cap: MAX_COALESCE_ROUNDS=3 constant defined and used in the
coalescing for-loop, preventing indefinite dispatch stall under
pathological fast-primary + slow-maintainContext conditions.
Testability: annotateForStaleness extracted as an exported pure function
in advise-tool.ts and used in AgentSession#routeAdvice. Three unit
tests added in advisor.test.ts covering the no-staleness, staleness,
and note-preservation contracts. This addresses the ReviewSession
regression-test concern without requiring a full AgentSession harness.
Reprime turn-tally coverage: new test 'backlog stays accurate when a
delta arrives during the reprime-triggering maintainContext' asserts
runtime.backlog === 0 after all three turns, catching a deleted
turns += reduce(...) line.
Fragile double-await tests converted: sends-batch-when-maintenance-fails
and expands-plan-mode-context now use Promise.withResolvers signals
instead of counted await Promise.resolve() hops.
Docs: onTurnEnd JSDoc added; hasFreshBacklog comment broadened to cover
all drain-busy phases (not just agent.prompt).
Three related changes that address the pattern of the advisor flagging things
the primary already fixed:
Fix 1 — coalesce late-arriving deltas before agent.prompt (runtime.ts)
Refactored #drain into a reusable #collectAndMaintainBatch helper that loops
until the pending queue is stable (no new deltas arrive during a maintenance
check) before calling agent.prompt. Previously, any turn queued during the
maintainContext await was deferred a full extra model-call cycle; now it is
merged into the current batch after re-checking the token budget for the
expanded payload. Every await in the loop has an epoch guard so a
reset/dispose mid-await cannot leak a stale batch. finalTurns always counts
all merged turns so #backlog decrements correctly.
Fix 2 — hasFreshBacklog + delivery-time staleness annotation (runtime.ts, agent-session.ts)
Added AdvisorRuntime.hasFreshBacklog getter (true when #pending.length > 0
while agent.prompt is running — i.e., newer primary turns arrived after the
reviewed window). #routeAdvice checks it at delivery time and appends a
lightweight caveat to the note so the primary agent knows to verify before
acting. Uses #pending.length not #backlog, which is always > 0 mid-call.
Fix 3 — willContinue WIP marker in rendered delta + system prompt (runtime.ts, agent-session.ts, system.md)
onTurnEnd now accepts { willContinue } and passes it through to #renderDelta,
which tags the heading '[in progress — more steps follow]' for intermediate
turns. The agent-session.ts call site passes context.willContinue. The advisor
system prompt instructs the model to withhold critique on WIP updates.
Also fixed pre-existing inline casts in #renderDelta and #dedupContextMessage
that suppressed the type checker instead of using the narrowing already provided
by the role discriminant.
All 75 advisor tests pass; pre-existing type errors in cursor.ts are unrelated.
Included advisor tool-result text in the quarantine source check so legitimate findings from granted read/grep tools are not treated as model-generated contamination.
Kept assistant text out of the source set to avoid laundering prior advisor hallucinations.
Fixes#5181
Scanned allowed advise tool notes for output-only destructive directives before the tool can route them to the primary agent.
Kept provenance checks against the watched session update so legitimate warnings about user-provided dangerous text still pass.
Fixes#5181
Prevented late interrupting advisor findings from waking the primary after a terminal text answer when no queued work remains.
Added regression coverage for the advisor-confirmation path so duplicate primary turns are caught.
Fixes#4840
Cleared provider-native replay payloads and stop details when Advisor output is quarantined so persisted transcripts only contain the sanitized error.
Added regression coverage for OpenAI Responses-style providerPayload leakage.
Fixes#5181
Quarantined Advisor assistant turns that request tools outside the granted tool pool before they can enter the Advisor context.
Reset the Advisor runtime after quarantine so the next update re-primes from the primary transcript instead of replaying contaminated private context.
Fixes#5181
Separated advisor provider session identity from local advisor labels so Codex requests carry stable UUIDv7 values while transcripts keep their advisor-specific names.
Fixes#5040
- Added `onTurnError` hook to `AdvisorRuntime` to handle failed turns before retries.
- Integrated credential blocking in `AgentSession` to prevent retrying usage-limited accounts when advisor turns fail.
- Included account key in `codex-auto-reset` debug logs to improve skip reason visibility.
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
- Rewrote docs/advisor-watchdog.md 'Tools and isolation' to describe the
read-only default plus the WATCHDOG.yml tools: grant surface (edit, write,
bash, eval, browser, ...), and called out that grants do not bypass the
session's approval mode (always-ask / write / yolo).
- Added a WATCHDOG.yml section documenting the advisor roster file (fields,
legacy tool aliases, discovery locations) with an example that grants a
fixer advisor edit + bash.
- Reworked the intro (title + first paragraphs) and the trailing peer
sentence so they no longer promise a hard read-only observer.
- Updated the advisor system prompt to describe using whichever tools this
session grants instead of asserting read-only access.
- Fixed the AdvisorConfig docstring in advisor/config.ts to match the
runtime (any built-in name; default read/grep/glob).
Fixes#4044
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Removed the architectural restriction limiting advisors to read-only tools.
- Updated advisor configuration to permit any built-in tool, including `edit`, `write`, and `bash`.
- Defaulted advisor toolsets to `read`, `grep`, and `glob`, while maintaining strict session isolation for each advisor.
- Introduced comprehensive support for multiple concurrent, independently-configured advisors via `WATCHDOG.yml` files.
- Implemented a full-screen TUI overlay for managing advisor rosters, models, tools, and instructions.
- Added session-wide advisor initialization, telemetry aggregation, and named transcript isolation.
- Enhanced advisor security and observability with secret redaction in tool results and secure XML attribute encoding.
Agent.#runLoop appends the user batch and a synthetic stopReason: "error" assistant turn to state.messages before resolving prompt() with state.error set. AdvisorRuntime now snapshots state.messages.length before each prompt, restores it on failure (via a new AdvisorAgent.rollbackTo hook that also resets the advisor's append-only sync cursor), and clears state.error so retries replay a clean baseline and the drop-after-3 path never leaks orphan failed turns into the next successful run's context.
Fixes#3635
Agent.#runLoop catches provider/stream failures internally and resolves prompt() cleanly with the message recorded on state.error, so AdvisorRuntime treated the OpenRouter 404/no-endpoints turn as a success and never reached notifyFailure. Inspect state.error after each prompt and throw so the retry/notify path runs on real provider failures.
Fixes#3635
Surfaced non-recovering advisor prompt failures through session notices so provider errors like OpenRouter ZDR endpoint rejection are visible in the main session.
Fixes#3635
- Added formatAdvisorContextPrompt to render project context files into the advisor's system prompt.
- Updated AgentSession to accept and inject advisorContextPrompt into the session system prompt.
- Registered project context files for the advisor to ensure the reviewer evaluates the agent against standing project instructions like AGENTS.md.
- Update `umans-provider` test to remove references to deprecated GLM 5.1 model.
- Rename search tool reference to `grep` in `advisor` test.
- Improve test stability in TUI components by explicitly draining `setImmediate` queues before flushing terminal state.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Relocated json-parse (repairJson/parseJsonWithRepair/parseStreamingJson/
parseStreamingJsonThrottled) into @oh-my-pi/pi-utils; repointed all
ai/agent/coding-agent import sites to @oh-my-pi/pi-utils.
- readSseJson now recovers a truncated or lightly malformed final SSE event
through the shared streaming parser, removing the bespoke scanJsonState/
getClosingSuffix/isJsonTruncated scanner from stream.ts.
- Dropped the @oh-my-pi/pi-ai/utils/json-parse legacy bundled-plugin registry
entry; the symbols are exposed via the @oh-my-pi/pi-utils barrel.
- Moved json-parse tests into packages/utils/test.
The advisor system prompt told the watcher model "at most one advise per
update" and "NEVER send the same advice twice", but nothing enforced
either rule. Issue #3520 captured a session where the advisor emitted
309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No
issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker">
injections in the primary transcript and destabilizing the watched agent
after the task was already complete.
New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and:
- Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every
"Stop.", "*Stop*", "STOP!" variant keys to the same canonical form.
- Drops a small allowlist of content-free self-talk filler (stop, done,
complete, no issue continue, lgtm, nothing to add, no further input,
carry on, ...) - silence is the correct expression of "no concerns".
- Dedupes by exact normalized text across the session, FIFO-bounded at
4096 entries.
- Rate-limits to one accepted advise per advisor model prompt cycle. The
runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(),
so the new batch starts with a fresh budget. Suppressed calls don't
consume the budget - a noise call never displaces a real concern.
Reset on advisor reset (compaction, session switch, /new) so a re-primed
reviewer can re-raise old concerns against the rewritten transcript.
Suppression is invisible to the advisor model: AdviseTool still returns
"Recorded." for a dropped call. Surfacing "suppressed" risks the model
rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to
bypass the dedupe.
Fixes#3520
Tracked the highest delivered severity per note so a nit can later land as a concern or blocker without being silently dropped.
De-escalation back to nit/concern stays treated as a duplicate so the model cannot flap severities to bypass dedupe.
Fixes#3511
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.
Added coverage that repeated advice is allowed again after the dedupe state resets.
Fixes#3511
Deduplicated advisor notes inside the advise tool so a model cannot enqueue the same advisory repeatedly in one session.
Added focused regression coverage for duplicate advisory suppression.
Fixes#3511
Require advisor claims about tool arguments to cite transcript or inspected tool evidence instead of inventing hidden argument shapes. Add regression coverage for the prompt contract.\n\nFixes #3483
Walked custom/hook `details` recursively through the obfuscator so nested renderer fields (e.g. async-result `jobs[].label`) cannot leak configured secrets into the advisor prompt.
Fixes#3237
Rewrote file-mention path and content through the configured obfuscator before the advisor delta is formatted, matching the primary provider's hide-secrets behavior.
Fixes#3237