Commit Graph
46 Commits
Author SHA1 Message Date
roboomp fbd9227a60 docs(advisor): described full-agent tool grants and WATCHDOG.yml roster
- Rewrote docs/advisor-watchdog.md 'Tools and isolation' to describe the
  read-only default plus the WATCHDOG.yml tools: grant surface (edit, write,
  bash, eval, browser, ...), and called out that grants do not bypass the
  session's approval mode (always-ask / write / yolo).
- Added a WATCHDOG.yml section documenting the advisor roster file (fields,
  legacy tool aliases, discovery locations) with an example that grants a
  fixer advisor edit + bash.
- Reworked the intro (title + first paragraphs) and the trailing peer
  sentence so they no longer promise a hard read-only observer.
- Updated the advisor system prompt to describe using whichever tools this
  session grants instead of asserting read-only access.
- Fixed the AdvisorConfig docstring in advisor/config.ts to match the
  runtime (any built-in name; default read/grep/glob).

Fixes #4044
2026-07-01 05:21:09 +00:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 0501addbd2 feat(coding-agent): allowed advisors to use mutating tools
- Removed the architectural restriction limiting advisors to read-only tools.
- Updated advisor configuration to permit any built-in tool, including `edit`, `write`, and `bash`.
- Defaulted advisor toolsets to `read`, `grep`, and `glob`, while maintaining strict session isolation for each advisor.
2026-06-28 16:07:46 +02:00
can1357 fbad280b57 feat: implemented multi-advisor concurrent runtime with tui management
- Introduced comprehensive support for multiple concurrent, independently-configured advisors via `WATCHDOG.yml` files.
- Implemented a full-screen TUI overlay for managing advisor rosters, models, tools, and instructions.
- Added session-wide advisor initialization, telemetry aggregation, and named transcript isolation.
- Enhanced advisor security and observability with secret redaction in tool results and secure XML attribute encoding.
2026-06-28 12:55:09 +02:00
roboompandcan1357 540eb73bcf fix(advisor): rolled back failed turns before retry or drop
Agent.#runLoop appends the user batch and a synthetic stopReason: "error" assistant turn to state.messages before resolving prompt() with state.error set. AdvisorRuntime now snapshots state.messages.length before each prompt, restores it on failure (via a new AdvisorAgent.rollbackTo hook that also resets the advisor's append-only sync cursor), and clears state.error so retries replay a clean baseline and the drop-after-3 path never leaks orphan failed turns into the next successful run's context.

Fixes #3635
2026-06-27 10:46:08 +02:00
roboompandcan1357 db377d417c fix(advisor): surfaced state.error from a cleanly resolved prompt
Agent.#runLoop catches provider/stream failures internally and resolves prompt() cleanly with the message recorded on state.error, so AdvisorRuntime treated the OpenRouter 404/no-endpoints turn as a success and never reached notifyFailure. Inspect state.error after each prompt and throw so the retry/notify path runs on real provider failures.

Fixes #3635
2026-06-27 08:20:46 +02:00
roboomp 22b840399c style: bun run fix 2026-06-27 06:07:06 +00:00
roboomp 38550db31d fix(advisor): notified when advisor model failed
Surfaced non-recovering advisor prompt failures through session notices so provider errors like OpenRouter ZDR endpoint rejection are visible in the main session.

Fixes #3635
2026-06-27 06:06:42 +00:00
can1357 a6ac86fe7e feat(coding-agent): enabled project context injection for advisor prompts
- Added formatAdvisorContextPrompt to render project context files into the advisor's system prompt.
- Updated AgentSession to accept and inject advisorContextPrompt into the session system prompt.
- Registered project context files for the advisor to ensure the reviewer evaluates the agent against standing project instructions like AGENTS.md.
2026-06-27 06:34:42 +02:00
can1357 26b3a22186 test: updated test suites and fix rendering test flakiness
- Update `umans-provider` test to remove references to deprecated GLM 5.1 model.
- Rename search tool reference to `grep` in `advisor` test.
- Improve test stability in TUI components by explicitly draining `setImmediate` queues before flushing terminal state.
2026-06-27 04:28:39 +02:00
can1357 ec03d3366e refactor(prompt): share prompt path normalization 2026-06-27 01:40:13 +02:00
can1357 5abe19eda1 Merge PR #3156: fix(advisor): surface nested repo context (@oldschoola)
# Conflicts:
#	packages/coding-agent/src/modes/components/status-line/component.ts
#	packages/coding-agent/src/sdk.ts
#	packages/coding-agent/src/system-prompt.ts
2026-06-27 01:40:12 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 1bd4f4b8c8 refactor(utils): moved shared JSON parser from pi-ai to pi-utils
- Relocated json-parse (repairJson/parseJsonWithRepair/parseStreamingJson/
  parseStreamingJsonThrottled) into @oh-my-pi/pi-utils; repointed all
  ai/agent/coding-agent import sites to @oh-my-pi/pi-utils.
- readSseJson now recovers a truncated or lightly malformed final SSE event
  through the shared streaming parser, removing the bespoke scanJsonState/
  getClosingSuffix/isJsonTruncated scanner from stream.ts.
- Dropped the @oh-my-pi/pi-ai/utils/json-parse legacy bundled-plugin registry
  entry; the symbols are exposed via the @oh-my-pi/pi-utils barrel.
- Moved json-parse tests into packages/utils/test.
2026-06-27 00:16:22 +02:00
can1357 4e3a0a0f46 Merge PR #3523: fix(advisor): filter content-free advisor notes at enqueue boundary (@roboomp)
# Conflicts:
#	packages/coding-agent/src/session/agent-session.ts
2026-06-26 23:42:40 +02:00
can1357 d81669a520 test(advisor): cover hidden argument transcript gap 2026-06-26 23:27:50 +02:00
can1357 d7a4f9e064 Merge PR #3488: fix(advisor): ground hidden-argument warnings (@roboomp) 2026-06-26 23:27:39 +02:00
roboomp ce4f5b8615 fix(advisor): enforced one-advise-per-update + dedupe + noise filter at the enqueueAdvice boundary
The advisor system prompt told the watcher model "at most one advise per
update" and "NEVER send the same advice twice", but nothing enforced
either rule. Issue #3520 captured a session where the advisor emitted
309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No
issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker">
injections in the primary transcript and destabilizing the watched agent
after the task was already complete.

New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and:

- Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every
  "Stop.", "*Stop*", "STOP!" variant keys to the same canonical form.
- Drops a small allowlist of content-free self-talk filler (stop, done,
  complete, no issue continue, lgtm, nothing to add, no further input,
  carry on, ...) - silence is the correct expression of "no concerns".
- Dedupes by exact normalized text across the session, FIFO-bounded at
  4096 entries.
- Rate-limits to one accepted advise per advisor model prompt cycle. The
  runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(),
  so the new batch starts with a fresh budget. Suppressed calls don't
  consume the budget - a noise call never displaces a real concern.

Reset on advisor reset (compaction, session switch, /new) so a re-primed
reviewer can re-raise old concerns against the rewritten transcript.

Suppression is invisible to the advisor model: AdviseTool still returns
"Recorded." for a dropped call. Surfacing "suppressed" risks the model
rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to
bypass the dedupe.

Fixes #3520
2026-06-26 02:48:27 +00:00
roboomp f62ee17080 fix(advisor): forwarded severity escalations past dedupe
Tracked the highest delivered severity per note so a nit can later land as a concern or blocker without being silently dropped.

De-escalation back to nit/concern stays treated as a duplicate so the model cannot flap severities to bypass dedupe.

Fixes #3511
2026-06-26 00:34:04 +00:00
roboomp cce5133bfa fix(advisor): reset advisory dedupe state
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.

Added coverage that repeated advice is allowed again after the dedupe state resets.

Fixes #3511
2026-06-26 00:31:51 +00:00
roboomp 6d58901cc3 fix(advisor): suppressed repeated advisories
Deduplicated advisor notes inside the advise tool so a model cannot enqueue the same advisory repeatedly in one session.

Added focused regression coverage for duplicate advisory suppression.

Fixes #3511
2026-06-26 00:26:01 +00:00
roboomp 07f43bfb86 style: bun run fix 2026-06-25 15:40:32 +00:00
roboomp e6ec029fbc fix(advisor): grounded hidden argument warnings
Require advisor claims about tool arguments to cite transcript or inspected tool evidence instead of inventing hidden argument shapes. Add regression coverage for the prompt contract.\n\nFixes #3483
2026-06-25 15:40:18 +00:00
roboomp 426ca4bf03 fix(coding-agent): redacted nested advisor custom details
Walked custom/hook `details` recursively through the obfuscator so nested renderer fields (e.g. async-result `jobs[].label`) cannot leak configured secrets into the advisor prompt.

Fixes #3237
2026-06-22 07:24:51 +00:00
roboomp df71f9bc18 fix(coding-agent): redacted advisor file mentions
Rewrote file-mention path and content through the configured obfuscator before the advisor delta is formatted, matching the primary provider's hide-secrets behavior.

Fixes #3237
2026-06-22 07:20:11 +00:00
roboomp 2d65f29d90 fix(coding-agent): redacted advisor context before formatting
Redacted structured advisor delta messages before markdown/XML formatting so secrets containing XML-significant characters still match configured hide-secrets entries.

Added regression coverage for expanded primary-context redaction.

Fixes #3237
2026-06-22 07:14:05 +00:00
roboomp a7ee69ffff style: bun run fix 2026-06-22 07:05:46 +00:00
roboomp 7dadeab331 fix(coding-agent): hid secrets from advisor prompts
Stopped restoring placeholders inside opaque assistant thinking blocks and threaded the configured secret obfuscator into advisor session-update prompts.

Added regression coverage for thinking preservation and advisor prompt redaction.

Fixes #3237
2026-06-22 07:05:23 +00:00
oldschoola bd3a35e13d fix(advisor): surface nested repo context 2026-06-20 15:55:47 -07:00
can1357 81a4af6684 perf(coding-agent): deduplicated primary context messages in advisor history
- Implemented message deduplication in `AdvisorRuntime` to collapse verbatim re-injected primary context (plan rules/approved plans) between turns.
- Reduced token consumption by replacing identical primary context segments with a status marker.
- Updated history formatter to allow expansion of load-bearing context types specifically, while retaining one-line summaries for other custom messages.
2026-06-20 22:59:30 +02:00
can1357 29d250fae2 feat(coding-agent): supported advisor transcript persistence
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
2026-06-19 03:50:00 +02:00
can1357 30a133cef7 feat(coding-agent): prevented stale advisor data leakage across conversations
- Add an epoch counter to `AdvisorRuntime` to discard in-flight advisor batches when a reset or disposal occurs.
- Introduce `resetAdvisorSessionState` to clear advisor-specific queues, latches, and pending cards, ensuring pre-reset state does not interfere with new conversations.
- Extend `YieldQueue.clear` to support conditional clearing by entry kind.
2026-06-18 18:58:01 +02:00
can1357 a050474af7 feat: migrated validation schemas and tool definitions from Zod to ArkType
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
2026-06-18 00:59:53 +02:00
can1357 d4317d3d20 feat(ai): consolidated OpenAI-family streaming and add OpenRouter API support
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.

Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
    - Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
    - Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
    - Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
2026-06-17 21:36:48 +02:00
can1357 54d212ec33 feat(coding-agent): added advisor immune-turn window for concern/blocker interruptions
- Tracked completed primary turns in `AgentSession` and started an immune-turn window after each interrupting advisor steer.
- Routed follow-on `concern`/`blocker` notes to the aside channel while the immune window is active, while preserving prior auto-resume-suppressed handling.
- Added the `advisor.immuneTurns` setting and tests for immune-turn detection and delivery-channel decisions.
2026-06-17 13:21:47 +02:00
can1357 3459371724 fix(coding-agent/advisor): fixed advisor concern/blocker notes being stranded after interrupts
- Added resolveAdvisorDeliveryChannel in advisor tooling to map each note to aside, steer, or preserve using severity, auto-resume suppression, core-streaming, and abort state.
- Updated AgentSession advice enqueuing to route concern/blocker notes through that resolver, preserving them only when the interrupted turn is idle or tearing down and steering them during active resumed turns.
- Added regression tests for resolveAdvisorDeliveryChannel covering nit versus interrupting severities across streaming, aborting, and suppression combinations.
2026-06-16 22:53:39 +02:00
can1357 712e859022 feat(coding-agent): restored the verbose /dump and /advisor dump raw output
- Rewrote `formatSessionDumpText` in `session-dump-format.ts` to emit the pre-16.x full dump: system-prompt prelude, model/thinking config, tool inventory with parameters, and the transcript as markdown role headings (`## User`, `## Assistant`, `### Tool Call`/`### Tool Result`), reusing `renderDelimitedThinking` for `<thinking>` blocks.
- Dropped the compact default and the `[raw]` flag from `/dump`: removed the `isRaw` parameter from `handleDumpCommand` in `command-controller.ts`, `interactive-mode.ts`, and `types.ts`, and removed the `inlineHint: "[raw]"`/`compact` plumbing in `builtin-registry.ts`.
- Updated the `formatSessionAsText` doc comment in `agent-session.ts` to describe the verbose dump shape.
- Removed the obsolete `formatSessionDumpText raw thinking` suite from `advisor.test.ts` and refreshed `session-dump-format.test.ts` to assert the verbose dump output.
- Recorded the revert in the coding-agent changelog and trimmed `/dump` from the compact transcript tool-intent-prefix entry.
2026-06-16 20:53:05 +02:00
can1357 5bce7ed6df feat: added advisory transcript formatting and one-shot benchmark metrics
- Introduced advisory note output as `<advisory>` tags with optional severity and guidance.
- Updated session transcript formatting to `### Session update` and inline watched role labels.
- Added shared `escapeXmlText` utility and escaped XML-sensitive text in advisor outputs.
- Added one-shot success run token metrics and one-shot statistics reporting.
2026-06-16 18:34:50 +02:00
roboomp e904585060 fix(ai): preserved sibling thinking envelopes
Parsed literal thinking envelopes independently before transcript rendering so interleaved thinking blocks do not collapse into one malformed wrapper. Added advisor raw dump regression coverage for sibling literal thinking blocks.\n\nFixes #2700
2026-06-15 21:05:08 +00:00
roboomp a0481ac070 fix(ai): unwrapped thinking envelopes in raw dumps
Normalized dialect thinking rendering so stored thinking content that already carries literal thinking tags is unwrapped before transcript serialization. Added advisor raw dump regression coverage.\n\nFixes #2700
2026-06-15 20:44:39 +00:00
can1357 9ba66ef7e7 feat(coding-agent/session): added tool intent comments to tool call history output
- Added an includeToolIntent option to session history formatting to control intent comments.
- Updated tool call rendering to prefix call lines with a leading comment when the tool argument contains a non-empty intent field.
- Enabled intent rendering in advisor history updates and added tests for intent-on and intent-off output paths.
2026-06-15 17:27:15 +02:00
can1357 dbf4c734fc feat(coding-agent): added advisor backlog sync and end-of-turn callback support
- Added awaited `onTurnEnd` and `setOnTurnEnd` wiring for turn-end callbacks.
- Added `advisor.syncBacklog` settings (off/1/3/5) and documented 30-second catch-up caps.
- Fixed advisor runtime backlog handling with failure counters, waiters, and retry requeue.
- Updated agent sessions to enqueue advisor updates on turn end and removed direct `turn_end` branch logic.
2026-06-15 17:25:04 +02:00
can1357 1524f7fb01 feat(coding-agent): added WATCHDOG.md discovery and advisor startup behavior updates
- Added discovery of local, user, and ancestor `WATCHDOG.md` files via `discoverWatchdogFiles`.
- Appended discovered watchdog prompts to advisor system prompts during session setup.
- Added protocol startup defaults that force `advisor.enabled` and `advisor.subagents` false.
- Handled `maintainContext` failures and drained pending updates before token estimation.
2026-06-15 17:13:35 +02:00
can1357 ecccff62f3 fix(coding-agent/advisor): fixed collapsed advisor notes to wrap instead of two-line truncation
- Removed the collapsed-mode truncation branch that forced advisor message cards to restrict bodies to two lines.
- Added a regression test covering long collapsed advisor notes to ensure they wrap at narrow widths instead of being cut short.
- Updated the package changelog to document the collapsed advisor note wrapping fix.
2026-06-15 16:54:20 +02:00
can1357 c6fd50b8e6 feat(coding-agent): added advisor context auto-maintenance with safer replay and compaction
- Added advisor context maintenance hook and token estimation before prompting for auto-upkeep.
- Added re-prime replay handling to reset advisor context and recover deferred prompts.
- Implemented session-level context compaction with model promotion and snapcompact-first fallback summarization.
- Surfaced advisor settings in the model tab and updated advisor system guidance text.
2026-06-15 16:49:21 +02:00
can1357 37ecd3e73a feat(advisor): added advisor agent for passive code review with severity-tagged advice
- Created AdvisorRuntime and AdviseTool to drive a read-only advisor agent that delivers severity-tagged advice (nit, concern, blocker) with interruption policy and transcript delta rendering.
- Added /advisor slash command with on/off/status/dump subcommands to control advisor lifecycle and inspect advisor metrics (model, messages, tokens, cost).
- Added advisor.enabled and advisor.subagents settings to enable passive advisor review on main agent and spawned task/eval subagents.
- Implemented advisor message rendering with severity-color badges (blocker=error, concern=warning, nit=muted) in chat log and status line indicator (++ badge).
- Extended yield-queue and session-history-format to support advisor batching and optional thinking block inclusion.
2026-06-15 16:32:13 +02:00