Commit Graph
69 Commits
Author SHA1 Message Date
can1357 5e74444c27 fix(test): repair auth-storage-rotation merge resolution and normalize changelogs 2026-07-14 18:50:47 +02:00
can1357 1dfec0d161 Merge PR #5090: fix(advisor): reduce stale advisories via delta coalescing, WIP markers, and delivery annotation (@apoc)
# Conflicts:
#	packages/coding-agent/src/advisor/__tests__/advisor.test.ts
#	packages/coding-agent/src/advisor/runtime.ts
#	packages/coding-agent/src/session/agent-session.ts
2026-07-14 18:44:57 +02:00
can1357 e96f5451db Merge PR #5216: fix(coding-agent): reject empty advisor stop completions (@roboomp) 2026-07-14 18:41:15 +02:00
can1357 b52a4cdede Merge PR #5186: fix(coding-agent): preserve late advisor notes after terminal answers (@roboomp)
# Conflicts:
#	packages/coding-agent/test/agent-session-advisor-suppression.test.ts
2026-07-14 18:40:35 +02:00
can1357 c0a0216867 fix(advisor): close quarantine containment gaps 2026-07-14 18:39:54 +02:00
roboomp 8608395eef fix(coding-agent): rejected empty advisor stops
Treat content-less advisor stop completions as failed turns so the advisor retry/drop path handles silent provider responses instead of accepting them as successful reviews.

Fixes #5212
2026-07-11 17:02:18 +00:00
Miroslav Drbal 74be4d5f67 style(advisor): apply biome formatting 2026-07-11 19:00:34 +02:00
Miroslav Drbal 59017f2616 fix(advisor): address final review: wip return type, import order, cap edge case, cap test
Blockers from final review:
- runtime.ts:457 TS2741: #collectAndMaintainBatch now returns wip in its
  result type; #drain destructures it and passes it to the retry-requeue
  unshift — a WIP batch that fails retry no longer silently loses its
  [in progress] heading.
- agent-session.ts import order: annotateForStaleness moved before
  formatAdvisorBatchContent (biome enforces case-insensitive alpha order).

Cap edge case (advisor nit + review CONCERN): final round of the
  coalescing for-loop now breaks BEFORE the late-item splice so any
  deltas that arrived during round MAX_COALESCE_ROUNDS-1's
  maintainContext call stay in #pending for the next drain iteration
  instead of being merged into an unbudgeted batch. Doc comment updated
  to match ('left for the next drain iteration' is now accurate).

Cap test: bounded version that only pushes new turns for the first 3
  maintainContext calls so the drain while-loop terminates cleanly after
  a second iteration. The previous unbounded version created an infinite
  drain loop (each maintainContext call unconditionally pushed another
  turn) and timed out.
2026-07-11 18:59:11 +02:00
Miroslav Drbal 441197b462 fix(advisor): address review: wip on PendingDelta, MAX_COALESCE_ROUNDS cap, annotateForStaleness extraction
Blocker: wip field added to PendingDelta and threaded into the reprime
  path. #collectAndMaintainBatch now captures the most-recent WIP state
  from each delta batch and forwards it to #renderDelta in both the
  normal and reprime branches, so a willContinue:true turn never loses
  its [in progress] heading through a reprime.

Safety cap: MAX_COALESCE_ROUNDS=3 constant defined and used in the
  coalescing for-loop, preventing indefinite dispatch stall under
  pathological fast-primary + slow-maintainContext conditions.

Testability: annotateForStaleness extracted as an exported pure function
  in advise-tool.ts and used in AgentSession#routeAdvice. Three unit
  tests added in advisor.test.ts covering the no-staleness, staleness,
  and note-preservation contracts. This addresses the ReviewSession
  regression-test concern without requiring a full AgentSession harness.

Reprime turn-tally coverage: new test 'backlog stays accurate when a
  delta arrives during the reprime-triggering maintainContext' asserts
  runtime.backlog === 0 after all three turns, catching a deleted
  turns += reduce(...) line.

Fragile double-await tests converted: sends-batch-when-maintenance-fails
  and expands-plan-mode-context now use Promise.withResolvers signals
  instead of counted await Promise.resolve() hops.

Docs: onTurnEnd JSDoc added; hasFreshBacklog comment broadened to cover
  all drain-busy phases (not just agent.prompt).
2026-07-11 18:59:11 +02:00
Miroslav Drbal 74715f8cca fix(advisor): reduce stale advisories via delta coalescing, WIP markers, and delivery annotation
Three related changes that address the pattern of the advisor flagging things
the primary already fixed:

Fix 1 — coalesce late-arriving deltas before agent.prompt (runtime.ts)
  Refactored #drain into a reusable #collectAndMaintainBatch helper that loops
  until the pending queue is stable (no new deltas arrive during a maintenance
  check) before calling agent.prompt. Previously, any turn queued during the
  maintainContext await was deferred a full extra model-call cycle; now it is
  merged into the current batch after re-checking the token budget for the
  expanded payload. Every await in the loop has an epoch guard so a
  reset/dispose mid-await cannot leak a stale batch. finalTurns always counts
  all merged turns so #backlog decrements correctly.

Fix 2 — hasFreshBacklog + delivery-time staleness annotation (runtime.ts, agent-session.ts)
  Added AdvisorRuntime.hasFreshBacklog getter (true when #pending.length > 0
  while agent.prompt is running — i.e., newer primary turns arrived after the
  reviewed window). #routeAdvice checks it at delivery time and appends a
  lightweight caveat to the note so the primary agent knows to verify before
  acting. Uses #pending.length not #backlog, which is always > 0 mid-call.

Fix 3 — willContinue WIP marker in rendered delta + system prompt (runtime.ts, agent-session.ts, system.md)
  onTurnEnd now accepts { willContinue } and passes it through to #renderDelta,
  which tags the heading '[in progress — more steps follow]' for intermediate
  turns. The agent-session.ts call site passes context.willContinue. The advisor
  system prompt instructs the model to withhold critique on WIP updates.

Also fixed pre-existing inline casts in #renderDelta and #dedupContextMessage
that suppressed the type checker instead of using the narrowing already provided
by the role discriminant.

All 75 advisor tests pass; pre-existing type errors in cursor.ts are unrelated.
2026-07-11 18:56:04 +02:00
roboomp ea5324fb10 fix(advisor): trusted advisor tool result provenance
Included advisor tool-result text in the quarantine source check so legitimate findings from granted read/grep tools are not treated as model-generated contamination.

Kept assistant text out of the source set to avoid laundering prior advisor hallucinations.

Fixes #5181
2026-07-11 13:17:46 +00:00
roboomp 77115fe16a fix(advisor): quarantined unsafe advise notes
Scanned allowed advise tool notes for output-only destructive directives before the tool can route them to the primary agent.

Kept provenance checks against the watched session update so legitimate warnings about user-provided dangerous text still pass.

Fixes #5181
2026-07-11 13:06:36 +00:00
roboomp 708eafaf8d fix(coding-agent): preserved late advisor terminal notes
Prevented late interrupting advisor findings from waking the primary after a terminal text answer when no queued work remains.

Added regression coverage for the advisor-confirmation path so duplicate primary turns are caught.

Fixes #4840
2026-07-11 12:59:22 +00:00
roboomp 9c961acd69 fix(advisor): cleared quarantined native payloads
Cleared provider-native replay payloads and stop details when Advisor output is quarantined so persisted transcripts only contain the sanitized error.

Added regression coverage for OpenAI Responses-style providerPayload leakage.

Fixes #5181
2026-07-11 12:56:39 +00:00
roboomp a58e8faa09 fix(advisor): quarantined unknown tool responses
Quarantined Advisor assistant turns that request tools outside the granted tool pool before they can enter the Advisor context.

Reset the Advisor runtime after quarantine so the next update re-primes from the primary transcript instead of replaying contaminated private context.

Fixes #5181
2026-07-11 12:47:26 +00:00
roboomp dabb2291a7 fix(advisor): kept defaults for invalid tools
Returned to default advisor tools when a non-empty configured tools list filters down to zero valid names.

Fixes #5155
2026-07-11 06:14:11 +00:00
roboomp 7d72ee9e0e fix(advisor): preserved empty tool lists
Kept explicit advisor tools: [] distinct from an omitted tools field so /advisor config can persist no tool access.

Fixes #5155
2026-07-11 06:01:16 +00:00
roboomp 325375f801 fix(advisor): used uuidv7 codex session ids
Separated advisor provider session identity from local advisor labels so Codex requests carry stable UUIDv7 values while transcripts keep their advisor-specific names.

Fixes #5040
2026-07-10 07:24:36 +00:00
can1357 cd7feac260 feat(coding-agent/advisor): implemented advisor credential blocking on usage limits
- Added `onTurnError` hook to `AdvisorRuntime` to handle failed turns before retries.
- Integrated credential blocking in `AgentSession` to prevent retrying usage-limited accounts when advisor turns fail.
- Included account key in `codex-auto-reset` debug logs to improve skip reason visibility.
2026-07-08 13:47:39 +02:00
can1357 59976e7d82 style: removed dead test helpers and formatted merged tests 2026-07-05 13:26:57 +02:00
can1357 17a94d2d5f fix: preserve hidden image descriptions in advisor history 2026-07-05 13:25:27 +02:00
Jeff Scott Ward 5731b861d0 fix: hide hidden custom session updates 2026-07-02 13:49:38 -04:00
can1357 95b91c7f73 feat(coding-agent/tools)!: replaced paths arrays with path strings
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
2026-07-02 08:30:33 +02:00
roboomp fbd9227a60 docs(advisor): described full-agent tool grants and WATCHDOG.yml roster
- Rewrote docs/advisor-watchdog.md 'Tools and isolation' to describe the
  read-only default plus the WATCHDOG.yml tools: grant surface (edit, write,
  bash, eval, browser, ...), and called out that grants do not bypass the
  session's approval mode (always-ask / write / yolo).
- Added a WATCHDOG.yml section documenting the advisor roster file (fields,
  legacy tool aliases, discovery locations) with an example that grants a
  fixer advisor edit + bash.
- Reworked the intro (title + first paragraphs) and the trailing peer
  sentence so they no longer promise a hard read-only observer.
- Updated the advisor system prompt to describe using whichever tools this
  session grants instead of asserting read-only access.
- Fixed the AdvisorConfig docstring in advisor/config.ts to match the
  runtime (any built-in name; default read/grep/glob).

Fixes #4044
2026-07-01 05:21:09 +00:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 0501addbd2 feat(coding-agent): allowed advisors to use mutating tools
- Removed the architectural restriction limiting advisors to read-only tools.
- Updated advisor configuration to permit any built-in tool, including `edit`, `write`, and `bash`.
- Defaulted advisor toolsets to `read`, `grep`, and `glob`, while maintaining strict session isolation for each advisor.
2026-06-28 16:07:46 +02:00
can1357 fbad280b57 feat: implemented multi-advisor concurrent runtime with tui management
- Introduced comprehensive support for multiple concurrent, independently-configured advisors via `WATCHDOG.yml` files.
- Implemented a full-screen TUI overlay for managing advisor rosters, models, tools, and instructions.
- Added session-wide advisor initialization, telemetry aggregation, and named transcript isolation.
- Enhanced advisor security and observability with secret redaction in tool results and secure XML attribute encoding.
2026-06-28 12:55:09 +02:00
roboompandcan1357 540eb73bcf fix(advisor): rolled back failed turns before retry or drop
Agent.#runLoop appends the user batch and a synthetic stopReason: "error" assistant turn to state.messages before resolving prompt() with state.error set. AdvisorRuntime now snapshots state.messages.length before each prompt, restores it on failure (via a new AdvisorAgent.rollbackTo hook that also resets the advisor's append-only sync cursor), and clears state.error so retries replay a clean baseline and the drop-after-3 path never leaks orphan failed turns into the next successful run's context.

Fixes #3635
2026-06-27 10:46:08 +02:00
roboompandcan1357 db377d417c fix(advisor): surfaced state.error from a cleanly resolved prompt
Agent.#runLoop catches provider/stream failures internally and resolves prompt() cleanly with the message recorded on state.error, so AdvisorRuntime treated the OpenRouter 404/no-endpoints turn as a success and never reached notifyFailure. Inspect state.error after each prompt and throw so the retry/notify path runs on real provider failures.

Fixes #3635
2026-06-27 08:20:46 +02:00
roboomp 22b840399c style: bun run fix 2026-06-27 06:07:06 +00:00
roboomp 38550db31d fix(advisor): notified when advisor model failed
Surfaced non-recovering advisor prompt failures through session notices so provider errors like OpenRouter ZDR endpoint rejection are visible in the main session.

Fixes #3635
2026-06-27 06:06:42 +00:00
can1357 a6ac86fe7e feat(coding-agent): enabled project context injection for advisor prompts
- Added formatAdvisorContextPrompt to render project context files into the advisor's system prompt.
- Updated AgentSession to accept and inject advisorContextPrompt into the session system prompt.
- Registered project context files for the advisor to ensure the reviewer evaluates the agent against standing project instructions like AGENTS.md.
2026-06-27 06:34:42 +02:00
can1357 26b3a22186 test: updated test suites and fix rendering test flakiness
- Update `umans-provider` test to remove references to deprecated GLM 5.1 model.
- Rename search tool reference to `grep` in `advisor` test.
- Improve test stability in TUI components by explicitly draining `setImmediate` queues before flushing terminal state.
2026-06-27 04:28:39 +02:00
can1357 ec03d3366e refactor(prompt): share prompt path normalization 2026-06-27 01:40:13 +02:00
can1357 5abe19eda1 Merge PR #3156: fix(advisor): surface nested repo context (@oldschoola)
# Conflicts:
#	packages/coding-agent/src/modes/components/status-line/component.ts
#	packages/coding-agent/src/sdk.ts
#	packages/coding-agent/src/system-prompt.ts
2026-06-27 01:40:12 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 1bd4f4b8c8 refactor(utils): moved shared JSON parser from pi-ai to pi-utils
- Relocated json-parse (repairJson/parseJsonWithRepair/parseStreamingJson/
  parseStreamingJsonThrottled) into @oh-my-pi/pi-utils; repointed all
  ai/agent/coding-agent import sites to @oh-my-pi/pi-utils.
- readSseJson now recovers a truncated or lightly malformed final SSE event
  through the shared streaming parser, removing the bespoke scanJsonState/
  getClosingSuffix/isJsonTruncated scanner from stream.ts.
- Dropped the @oh-my-pi/pi-ai/utils/json-parse legacy bundled-plugin registry
  entry; the symbols are exposed via the @oh-my-pi/pi-utils barrel.
- Moved json-parse tests into packages/utils/test.
2026-06-27 00:16:22 +02:00
can1357 4e3a0a0f46 Merge PR #3523: fix(advisor): filter content-free advisor notes at enqueue boundary (@roboomp)
# Conflicts:
#	packages/coding-agent/src/session/agent-session.ts
2026-06-26 23:42:40 +02:00
can1357 d81669a520 test(advisor): cover hidden argument transcript gap 2026-06-26 23:27:50 +02:00
can1357 d7a4f9e064 Merge PR #3488: fix(advisor): ground hidden-argument warnings (@roboomp) 2026-06-26 23:27:39 +02:00
roboomp ce4f5b8615 fix(advisor): enforced one-advise-per-update + dedupe + noise filter at the enqueueAdvice boundary
The advisor system prompt told the watcher model "at most one advise per
update" and "NEVER send the same advice twice", but nothing enforced
either rule. Issue #3520 captured a session where the advisor emitted
309 advise() calls covering 92 unique notes - 114x "Stop.", 52x "No
issue; continue.", 41x "Done." - landing 309 <advisory severity="blocker">
injections in the primary transcript and destabilizing the watched agent
after the task was already complete.

New AdvisorEmissionGuard sits on AgentSession#enqueueAdvice and:

- Normalizes notes (lowercase, NFKC, punctuation->space, trim) so every
  "Stop.", "*Stop*", "STOP!" variant keys to the same canonical form.
- Drops a small allowlist of content-free self-talk filler (stop, done,
  complete, no issue continue, lgtm, nothing to add, no further input,
  carry on, ...) - silence is the correct expression of "no concerns".
- Dedupes by exact normalized text across the session, FIFO-bounded at
  4096 entries.
- Rate-limits to one accepted advise per advisor model prompt cycle. The
  runtime calls host.beginAdvisorUpdate?.() before each agent.prompt(),
  so the new batch starts with a fresh budget. Suppressed calls don't
  consume the budget - a noise call never displaces a real concern.

Reset on advisor reset (compaction, session switch, /new) so a re-primed
reviewer can re-raise old concerns against the rewritten transcript.

Suppression is invisible to the advisor model: AdviseTool still returns
"Recorded." for a dropped call. Surfacing "suppressed" risks the model
rephrasing the same useless note ("Stop." -> "Halt." -> "Cease.") to
bypass the dedupe.

Fixes #3520
2026-06-26 02:48:27 +00:00
roboomp f62ee17080 fix(advisor): forwarded severity escalations past dedupe
Tracked the highest delivered severity per note so a nit can later land as a concern or blocker without being silently dropped.

De-escalation back to nit/concern stays treated as a duplicate so the model cannot flap severities to bypass dedupe.

Fixes #3511
2026-06-26 00:34:04 +00:00
roboomp cce5133bfa fix(advisor): reset advisory dedupe state
Cleared delivered-note memory when the advisor session state resets across conversation boundaries.

Added coverage that repeated advice is allowed again after the dedupe state resets.

Fixes #3511
2026-06-26 00:31:51 +00:00
roboomp 6d58901cc3 fix(advisor): suppressed repeated advisories
Deduplicated advisor notes inside the advise tool so a model cannot enqueue the same advisory repeatedly in one session.

Added focused regression coverage for duplicate advisory suppression.

Fixes #3511
2026-06-26 00:26:01 +00:00
roboomp 07f43bfb86 style: bun run fix 2026-06-25 15:40:32 +00:00
roboomp e6ec029fbc fix(advisor): grounded hidden argument warnings
Require advisor claims about tool arguments to cite transcript or inspected tool evidence instead of inventing hidden argument shapes. Add regression coverage for the prompt contract.\n\nFixes #3483
2026-06-25 15:40:18 +00:00
roboomp 426ca4bf03 fix(coding-agent): redacted nested advisor custom details
Walked custom/hook `details` recursively through the obfuscator so nested renderer fields (e.g. async-result `jobs[].label`) cannot leak configured secrets into the advisor prompt.

Fixes #3237
2026-06-22 07:24:51 +00:00
roboomp df71f9bc18 fix(coding-agent): redacted advisor file mentions
Rewrote file-mention path and content through the configured obfuscator before the advisor delta is formatted, matching the primary provider's hide-secrets behavior.

Fixes #3237
2026-06-22 07:20:11 +00:00
roboomp 2d65f29d90 fix(coding-agent): redacted advisor context before formatting
Redacted structured advisor delta messages before markdown/XML formatting so secrets containing XML-significant characters still match configured hide-secrets entries.

Added regression coverage for expanded primary-context redaction.

Fixes #3237
2026-06-22 07:14:05 +00:00
roboomp a7ee69ffff style: bun run fix 2026-06-22 07:05:46 +00:00