- Rewrote docs/advisor-watchdog.md 'Tools and isolation' to describe the
read-only default plus the WATCHDOG.yml tools: grant surface (edit, write,
bash, eval, browser, ...), and called out that grants do not bypass the
session's approval mode (always-ask / write / yolo).
- Added a WATCHDOG.yml section documenting the advisor roster file (fields,
legacy tool aliases, discovery locations) with an example that grants a
fixer advisor edit + bash.
- Reworked the intro (title + first paragraphs) and the trailing peer
sentence so they no longer promise a hard read-only observer.
- Updated the advisor system prompt to describe using whichever tools this
session grants instead of asserting read-only access.
- Fixed the AdvisorConfig docstring in advisor/config.ts to match the
runtime (any built-in name; default read/grep/glob).
Fixes#4044
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Removed the architectural restriction limiting advisors to read-only tools.
- Updated advisor configuration to permit any built-in tool, including `edit`, `write`, and `bash`.
- Defaulted advisor toolsets to `read`, `grep`, and `glob`, while maintaining strict session isolation for each advisor.
- Introduced comprehensive support for multiple concurrent, independently-configured advisors via `WATCHDOG.yml` files.
- Implemented a full-screen TUI overlay for managing advisor rosters, models, tools, and instructions.
- Added session-wide advisor initialization, telemetry aggregation, and named transcript isolation.
- Enhanced advisor security and observability with secret redaction in tool results and secure XML attribute encoding.
Agent.#runLoop appends the user batch and a synthetic stopReason: "error" assistant turn to state.messages before resolving prompt() with state.error set. AdvisorRuntime now snapshots state.messages.length before each prompt, restores it on failure (via a new AdvisorAgent.rollbackTo hook that also resets the advisor's append-only sync cursor), and clears state.error so retries replay a clean baseline and the drop-after-3 path never leaks orphan failed turns into the next successful run's context.
Fixes#3635
Agent.#runLoop catches provider/stream failures internally and resolves prompt() cleanly with the message recorded on state.error, so AdvisorRuntime treated the OpenRouter 404/no-endpoints turn as a success and never reached notifyFailure. Inspect state.error after each prompt and throw so the retry/notify path runs on real provider failures.
Fixes#3635
Surfaced non-recovering advisor prompt failures through session notices so provider errors like OpenRouter ZDR endpoint rejection are visible in the main session.
Fixes#3635
- Added formatAdvisorContextPrompt to render project context files into the advisor's system prompt.
- Updated AgentSession to accept and inject advisorContextPrompt into the session system prompt.
- Registered project context files for the advisor to ensure the reviewer evaluates the agent against standing project instructions like AGENTS.md.
- Update `umans-provider` test to remove references to deprecated GLM 5.1 model.
- Rename search tool reference to `grep` in `advisor` test.
- Improve test stability in TUI components by explicitly draining `setImmediate` queues before flushing terminal state.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Relocated json-parse (repairJson/parseJsonWithRepair/parseStreamingJson/
parseStreamingJsonThrottled) into @oh-my-pi/pi-utils; repointed all
ai/agent/coding-agent import sites to @oh-my-pi/pi-utils.
- readSseJson now recovers a truncated or lightly malformed final SSE event
through the shared streaming parser, removing the bespoke scanJsonState/
getClosingSuffix/isJsonTruncated scanner from stream.ts.
- Dropped the @oh-my-pi/pi-ai/utils/json-parse legacy bundled-plugin registry
entry; the symbols are exposed via the @oh-my-pi/pi-utils barrel.
- Moved json-parse tests into packages/utils/test.
The advisor system prompt told the watcher model "at most one advise per
update" and "NEVER send the same advice twice", but nothing enforced
either rule. Issue #3520 captured a session where the advisor emitted
309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No
issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker">
injections in the primary transcript and destabilizing the watched agent
after the task was already complete.
New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and:
- Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every
"Stop.", "*Stop*", "STOP!" variant keys to the same canonical form.
- Drops a small allowlist of content-free self-talk filler (stop, done,
complete, no issue continue, lgtm, nothing to add, no further input,
carry on, ...) - silence is the correct expression of "no concerns".
- Dedupes by exact normalized text across the session, FIFO-bounded at
4096 entries.
- Rate-limits to one accepted advise per advisor model prompt cycle. The
runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(),
so the new batch starts with a fresh budget. Suppressed calls don't
consume the budget - a noise call never displaces a real concern.
Reset on advisor reset (compaction, session switch, /new) so a re-primed
reviewer can re-raise old concerns against the rewritten transcript.
Suppression is invisible to the advisor model: AdviseTool still returns
"Recorded." for a dropped call. Surfacing "suppressed" risks the model
rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to
bypass the dedupe.
Fixes#3520
Tracked the highest delivered severity per note so a nit can later land as a concern or blocker without being silently dropped.
De-escalation back to nit/concern stays treated as a duplicate so the model cannot flap severities to bypass dedupe.
Fixes#3511
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.
Added coverage that repeated advice is allowed again after the dedupe state resets.
Fixes#3511
Deduplicated advisor notes inside the advise tool so a model cannot enqueue the same advisory repeatedly in one session.
Added focused regression coverage for duplicate advisory suppression.
Fixes#3511
Require advisor claims about tool arguments to cite transcript or inspected tool evidence instead of inventing hidden argument shapes. Add regression coverage for the prompt contract.\n\nFixes #3483
Walked custom/hook `details` recursively through the obfuscator so nested renderer fields (e.g. async-result `jobs[].label`) cannot leak configured secrets into the advisor prompt.
Fixes#3237
Rewrote file-mention path and content through the configured obfuscator before the advisor delta is formatted, matching the primary provider's hide-secrets behavior.
Fixes#3237
- Implemented message deduplication in `AdvisorRuntime` to collapse verbatim re-injected primary context (plan rules/approved plans) between turns.
- Reduced token consumption by replacing identical primary context segments with a status marker.
- Updated history formatter to allow expansion of load-bearing context types specifically, while retaining one-line summaries for other custom messages.
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
- Add an epoch counter to `AdvisorRuntime` to discard in-flight advisor batches when a reset or disposal occurs.
- Introduce `resetAdvisorSessionState` to clear advisor-specific queues, latches, and pending cards, ensuring pre-reset state does not interfere with new conversations.
- Extend `YieldQueue.clear` to support conditional clearing by entry kind.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Tracked completed primary turns in `AgentSession` and started an immune-turn window after each interrupting advisor steer.
- Routed follow-on `concern`/`blocker` notes to the aside channel while the immune window is active, while preserving prior auto-resume-suppressed handling.
- Added the `advisor.immuneTurns` setting and tests for immune-turn detection and delivery-channel decisions.
- Added resolveAdvisorDeliveryChannel in advisor tooling to map each note to aside, steer, or preserve using severity, auto-resume suppression, core-streaming, and abort state.
- Updated AgentSession advice enqueuing to route concern/blocker notes through that resolver, preserving them only when the interrupted turn is idle or tearing down and steering them during active resumed turns.
- Added regression tests for resolveAdvisorDeliveryChannel covering nit versus interrupting severities across streaming, aborting, and suppression combinations.
- Rewrote `formatSessionDumpText` in `session-dump-format.ts` to emit the pre-16.x full dump: system-prompt prelude, model/thinking config, tool inventory with parameters, and the transcript as markdown role headings (`## User`, `## Assistant`, `### Tool Call`/`### Tool Result`), reusing `renderDelimitedThinking` for `<thinking>` blocks.
- Dropped the compact default and the `[raw]` flag from `/dump`: removed the `isRaw` parameter from `handleDumpCommand` in `command-controller.ts`, `interactive-mode.ts`, and `types.ts`, and removed the `inlineHint: "[raw]"`/`compact` plumbing in `builtin-registry.ts`.
- Updated the `formatSessionAsText` doc comment in `agent-session.ts` to describe the verbose dump shape.
- Removed the obsolete `formatSessionDumpText raw thinking` suite from `advisor.test.ts` and refreshed `session-dump-format.test.ts` to assert the verbose dump output.
- Recorded the revert in the coding-agent changelog and trimmed `/dump` from the compact transcript tool-intent-prefix entry.
- Introduced advisory note output as `<advisory>` tags with optional severity and guidance.
- Updated session transcript formatting to `### Session update` and inline watched role labels.
- Added shared `escapeXmlText` utility and escaped XML-sensitive text in advisor outputs.
- Added one-shot success run token metrics and one-shot statistics reporting.
Parsed literal thinking envelopes independently before transcript rendering so interleaved thinking blocks do not collapse into one malformed wrapper. Added advisor raw dump regression coverage for sibling literal thinking blocks.\n\nFixes #2700
- Added an includeToolIntent option to session history formatting to control intent comments.
- Updated tool call rendering to prefix call lines with a leading comment when the tool argument contains a non-empty intent field.
- Enabled intent rendering in advisor history updates and added tests for intent-on and intent-off output paths.
- Added awaited `onTurnEnd` and `setOnTurnEnd` wiring for turn-end callbacks.
- Added `advisor.syncBacklog` settings (off/1/3/5) and documented 30-second catch-up caps.
- Fixed advisor runtime backlog handling with failure counters, waiters, and retry requeue.
- Updated agent sessions to enqueue advisor updates on turn end and removed direct `turn_end` branch logic.
- Added discovery of local, user, and ancestor `WATCHDOG.md` files via `discoverWatchdogFiles`.
- Appended discovered watchdog prompts to advisor system prompts during session setup.
- Added protocol startup defaults that force `advisor.enabled` and `advisor.subagents` false.
- Handled `maintainContext` failures and drained pending updates before token estimation.
- Removed the collapsed-mode truncation branch that forced advisor message cards to restrict bodies to two lines.
- Added a regression test covering long collapsed advisor notes to ensure they wrap at narrow widths instead of being cut short.
- Updated the package changelog to document the collapsed advisor note wrapping fix.
- Added advisor context maintenance hook and token estimation before prompting for auto-upkeep.
- Added re-prime replay handling to reset advisor context and recover deferred prompts.
- Implemented session-level context compaction with model promotion and snapcompact-first fallback summarization.
- Surfaced advisor settings in the model tab and updated advisor system guidance text.
- Created AdvisorRuntime and AdviseTool to drive a read-only advisor agent that delivers severity-tagged advice (nit, concern, blocker) with interruption policy and transcript delta rendering.
- Added /advisor slash command with on/off/status/dump subcommands to control advisor lifecycle and inspect advisor metrics (model, messages, tokens, cost).
- Added advisor.enabled and advisor.subagents settings to enable passive advisor review on main agent and spawned task/eval subagents.
- Implemented advisor message rendering with severity-color badges (blocker=error, concern=warning, nit=muted) in chat log and status line indicator (++ badge).
- Extended yield-queue and session-history-format to support advisor batching and optional thinking block inclusion.