Commit Graph

543 Commits

Author SHA1 Message Date
can1357 bc5eeef398 feat(coding-agent): tagged read-only agents in task tool description
- Added `isReadOnlyAgent` and `READ_ONLY_TOOL_NAMES` to classify agents.
- Marked read-only agents and forbade edits, commands, and reasoning offload.
- Added tests for capability classification and description rendering.
2026-06-04 03:24:41 +02:00
can1357 dc4aeb7b88 refactor(coding-agent): renamed todo_write tool to todo
- Renamed `TodoWriteTool` to `TodoTool` and its source/prompt files.
- Updated tool registration, schema, renderers, and gating to `todo`.
- Adjusted cursor provider native tool names and tests to match.
- Renamed strike-animation constants and `todo-error-reminder` type.
2026-06-04 02:45:30 +02:00
can1357 384a206737 refactor(task): replaced numeric-prefix ids with name-first agent output ids
- Changed `AgentOutputManager` to use requested names verbatim, adding `-2`/`-3` suffixes only on repeats (e.g. `Anna`, `Anna-2`).
- Renamed main agent id from `0-Main` to `Main`; nested ids now use dot notation without numeric prefix (e.g. `Parent.Child`).
- Updated task widget to render dotted hierarchy as `Parent>Child` breadcrumb without leading index.
- Resume scan now tracks seen names instead of a counter to avoid clobbering prior outputs.
2026-06-02 06:50:03 +02:00
can1357 72cf371b9d Merge remote-tracking branch 'origin/farm/f30c73d1/demote-subagent-abort-log' 2026-06-01 14:37:10 +02:00
roboomp d14aa4dbe5 fix(task): demote subagent reminder-loop abort log
The catch around the subagent yield-reminder prompt previously logged
every exception at ERROR. User cancel (^C) and compaction-driven aborts
both surface as ToolAbortError through awaitAbortable, so benign control
flow generated 9 spurious 'Subagent prompt failed' errors in 2 days on
the reporter's instance.

Gate the ERROR branch on '!abortSignal.aborted && !(err instanceof
ToolAbortError)' and route the abort path to logger.debug. The outer
catch + finally still mark the run aborted, so observable behaviour is
unchanged.

Fixes #1623
2026-06-01 06:20:55 +00:00
can1357 7e11bea8f4 fix(coding-agent): repaired per-field double-encoded JSON in task tool
- Added `repairDoubleEncodedJsonString` to unescape fields double-encoded by the model (e.g. literal `\n`, `\"`, `\uXXXX` in `context`/`assignment`/`description`).
- Scoped repair to natural-language fields only, leaving code-bearing tools untouched.
- Applied repair on both render and execution paths in `TaskTool`.
2026-05-31 20:17:38 +02:00
can1357 a45747f96a refactor(coding-agent): replaced bracket-style section tags with underlined headers
- Converted [SECTION]...[/SECTION] markers to "SECTION\n===" format in system prompt templates.
- Updated system conventions doc to reference the new marker style.
- Updated tests to match against the new header pattern.
2026-05-31 20:10:56 +02:00
can1357 68430dee5c chore: renamed mnemosyne package to mnemopi
- Updated package name, directory, and binary from mnemosyne to mnemopi.
- Updated all lockfile references and workspace paths accordingly.
2026-05-31 08:45:12 +02:00
can1357 d1bd14f020 feat(coding-agent): propagated mcpManager and localProtocolOptions to subagents
- Stored mcpManager and localProtocolOptions on ToolSession so nested subagents inherit them without relying on process-global singletons.
- TaskTool now uses the session's localProtocolOptions and mcpManager when spawning sub-tasks, falling back to defaults if absent.
2026-05-31 06:49:20 +02:00
can1357 e22b31401b feat(packages/coding-agent): added orchestrate notices for session output
- Added orchestrate keyword detection and notice handling for non-synthetic prompts.
- Added orchestrate notice handling in session output paths, including streaming and append delivery.
- Added a system orchestrate notice specifying task-subagent delegation, phase workflow, and validation gates.
- Added shared gradient-highlighter utilities and switched ultrathink highlighting to use cached palettes.
- Removed embedded orchestrate prompt artifacts and updated usage tips for orchestration, ultrathink, and /login behavior.
2026-05-30 17:34:38 +02:00
can1357 99365385aa feat(mnemosyne): added mnemosyne parent-state sync in delegated sessions
- Added parentMnemosyneSessionState propagation from session state through SDK, executor, and task options into nested sessions.
- Added getMnemosyneSessionState() and rekeying logic to refresh Mnemosyne IDs during session sync, switch, and restore.
- Added Mnemosyne reset and teardown cleanup on unaliasing or restoration to avoid stale state.
2026-05-30 16:22:12 +02:00
can1357 baafa3c027 feat(render): added task renderContext propagation to hide preview rows
- Propagated task `renderContext` through `ToolExecutionComponent` so call rendering can detect result state.
- Suppressed task call-preview rows when a result snapshot exists to avoid duplicate task lines.
2026-05-30 16:22:12 +02:00
can1357 f559f6b10c ux(coding-agent): added ? fallback for %/window context in status/task
- Added formatContextUsage to render context as `%/window` with `?` fallback across status and task views.
- Updated task renderCall to show dispatched agents as tree entries with `Tasks (2)` header and `#3` fallback.
- Capped task preview collapse at 12 entries and added `... N more agents` overflow messaging.
- Replaced hard-coded dot separators with `theme.sep.dot` in subagent cost and context output.
- Added tests for streaming task preview rendering and updated nested-live expectations for percent/context output.
2026-05-30 15:30:05 +02:00
can1357 9d29fc9716 feat(mnemosyne): added configurable memory scoping with per-project-tagged mode
- Added `mnemosyne.scoping` setting: `global`, `per-project`, and `per-project-tagged`.
- `per-project-tagged` writes to a project-local bank while merging global memories on recall.
- Refactored `MnemosyneSessionState` to manage scoped recall/retain targets and deduplication.
- Updated hindsight tools to route recall/retain through scoped methods.
2026-05-30 14:47:53 +02:00
can1357 377ed34e08 refactor(coding-agent): replaced ctx/Σ labels with icon and cleaner cost separator
- Removed "ctx" suffix and cumulative Σ-token display from status lines.
- Replaced "N tools" text with tool count + extensionTool icon.
- Changed cost separator to ` . ` to visually distinguish it from dim stats.
- Added test asserting new format and absence of old labels.
2026-05-30 06:38:09 +02:00
can1357 535f7cfa89 fix(coding-agent/tools): reworked yolo approval resolution to honor user tool policies
- In `resolveApproval`, yolo mode now returns the user policy directly (`allow`/`prompt`/`deny`) and ignores tool `override` prompts.
- Updated approval-mode and approval unit tests to match the new behavior for critical bash patterns under yolo and auto-approve.
- Updated docs and settings metadata to describe yolo as user-policy-driven rather than override-driven.
2026-05-27 00:21:33 +02:00
can1357 e4a16451ec feat(coding-agent): added coding-agent approval types and mode options
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
2026-05-26 21:52:16 +02:00
oldschoola 384f429461 fix(coding-agent): tighten approval edge cases and rewrite mode docs
- approval: user 'tool: deny' now wins over critical-pattern override
  (the override only tightens allow->prompt; it must never re-arm a denied tool).
- approval: rename hindsight policy keys to match registered tool names
  (recall/retain/reflect, not hindsight_recall/hindsight_retain).
- approval: head+tail truncation for bash/ssh command prompts so a
  destructive suffix buried after a long benign preamble stays visible.
- task/executor: force tools.approvalMode='auto' in createSubagentSettings
  so subagents (which have no UI) cannot deadlock on per-tool prompts;
  the parent's approval of the task call is the authorization.
- docs/approval-mode: rewrite so every example surfaces tools.approvalMode
  and explains that tools.approval is ignored outside 'custom' mode.
2026-05-26 20:53:35 +02:00
Can Bölük e1b52714be Merge pull request #1398 from justadudewithtime/feat/resolved-model-badge
feat(coding-agent): surface resolved subagent model badge in task widget
2026-05-26 21:25:57 +03:00
can1357 8a5b3e9552 feat(eval): added shared executor inheritance for subagents with concurrent async cells
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
2026-05-26 14:37:56 +02:00
Magxm 03a29f3aec refactor(task): hoist settings.get into local before appendAgentStats calls 2026-05-26 18:00:14 +09:00
Magxm 075bfb94c1 fix(task): decouple appendAgentStats from global settings and sanitize resolvedModel
- Move settings.get("task.showResolvedModelBadge") out of appendAgentStats
  and into its callers, passing the value as an explicit opts field.
  This removes a hidden dependency on global state and makes the
  function a pure formatter.
- Sanitize resolvedModel with replaceTabs() + truncateToWidth(30)
  before rendering, consistent with the file's existing sanitization
  patterns and AGENTS.md requirements.
2026-05-26 17:41:21 +09:00
Magxm 8f5da835e7 feat(coding-agent): surface resolved subagent model badge in task widget 2026-05-26 17:30:11 +09:00
can1357 e0eae43fde feat(tools): added shared output schema validator for YieldTool
- Unified output schema construction and validation by adding buildOutputValidator and using it in YieldTool and task executor.
- Added MAX_SCHEMA_RETRIES so YieldTool now retries schema failures three times with hints before overriding.
- Updated failure handling to use shared summarizeValidationFailure and formatters for required-field reporting.
- Added tests for output-schema-validator and YieldTool covering malformed schemas and nested-array retry edge cases.
2026-05-26 06:34:30 +02:00
roboomp 0a3a48b92d fix(task): respected parent lsp disable for subagents
Combined task.enableLsp with the parent session enableLsp gate before spawning subagents, so --no-lsp remains authoritative even when subagent LSP is enabled in settings.

Fixes #1385
2026-05-26 03:13:28 +00:00
roboomp 28a0dc3ae4 feat(task): gated subagent LSP behind task.enableLsp setting
Added task.enableLsp (boolean, default false) and routed both regular and isolated subagent dispatch through it. Keeps subagents cheap by default while letting users opt in to LSP-aware delegation. Updated regression tests to cover the default-off, opt-in, plan-mode, and isolated paths.

Fixes #1385
2026-05-26 03:09:09 +00:00
roboomp 8948101e2b fix(task): forwarded parent enableLsp flag to subagents
Subagents now inherit the parent session's enableLsp value, so a top-level --no-lsp invocation propagates into spawned tasks instead of falling back to the executor's default of true.

Fixes #1385
2026-05-26 03:06:14 +00:00
roboomp 5ff1c747ce fix(task): inherited lsp for subagents
Removed the hardcoded subagent LSP disable flag so executor defaults and user settings control LSP availability. Passed the effective plan-mode agent definition into both regular and isolated subagent dispatch so plan-mode tool restrictions apply consistently.

Fixes #1385
2026-05-26 03:01:57 +00:00
Can Bölük d201442a16 Merge pull request #1372 from can1357/farm/e4c67c73/fix-subagent-session-start-busy
fix(agent): prevent subagent session_start busy race
2026-05-25 21:59:38 +03:00
roboomp 1217091557 fix(agent): prevented subagent session_start busy race
Queued extension-delivered user messages when deliverAs is set and waited for session_start extension message sends before prompting subagents.

Fixes #1343
2026-05-25 18:48:06 +00:00
Can Bölük 47dab57559 Merge branch 'main' into farm/5c2ff3c3/report-finding-tool-agent-output-schema- 2026-05-25 21:42:35 +03:00
can1357 3105870c86 feat(coding-agent/task): added live nested-subagent progress rendering
- Captured `tool_execution_update` snapshots for `task` calls into in-flight progress state for live nested rendering.
- Cleared in-flight task snapshots at task start and completion to prevent stale nested progress from persisting.
- Updated progress rendering to combine completed and in-flight task details through a dedicated nested task tree view.
2026-05-25 20:21:24 +02:00
roboomp 6e9cf81544 style: bun run fix 2026-05-25 18:01:45 +00:00
roboomp 6b14cf1f55 fix(coding-agent): coerced report_finding string priority to number for reviewer schema
The report_finding tool's priority is exposed as a string enum
("P0"-"P3") for ergonomics, but the reviewer agent and every
custom review agent declare priority as `type: number` in their
JTD output schema. The cast at executor.ts:1473 lied about the
runtime shape, so the auto-injected `findings[].priority` flowed
through as strings and every yield with at least one finding was
rejected with `findings.0.priority: expected number, received string`,
forcing the run into the schema_violation exit path.

Added `toReviewFinding(details)` in tools/review.ts that maps the
priority enum to its numeric ordinal via the existing PRIORITY_INFO
table and use it at the boundary in executor.ts. Render paths still
see the original `ReportFindingDetails` shape (string priority)
through normalizeReportFindings, so display formatting is unaffected.

Fixes #1350
2026-05-25 18:01:38 +00:00
can1357 5eec35367f chore: fix types 2026-05-25 12:42:26 +02:00
can1357 3801b4ee32 fix(coding-agent): added configurable retry delay cap and surfaced rate-limit failure state
- Added `retry.maxDelayMs` to the settings schema and interfaces, with a default cap for provider backoff delays.
- Updated session auto-retry logic to fail fast when a requested wait exceeds the cap without fallback, emitting terminal auto-retry failure state.
- Propagated retry state and failure data into task progress and rendering so children show retry/wait details and reminder prompts stop after terminal errors.
2026-05-25 12:37:32 +02:00
can1357 8e74996513 chore: adjust tests 2026-05-25 12:22:43 +02:00
Tommy Carlsson 5b35cd626b feat: add autoloadSkills frontmatter field for agent definitions
Adds optional autoloadSkills field to agent frontmatter that automatically loads listed skills when a sub-agent is spawned. Uses the same buildSkillPromptMessage + sendCustomMessage mechanism as interactive skill loading, queued via sendCustomMessage({ triggerTurn: false }) before the first session.prompt(task). No extra agent turns, no new injection path. Skills stay in listing for sub-resource access. Compaction behavior matches manual loading. Unknown skill names silently skipped.

Lore-id: f85fdbdc
Constraint: autoload must use buildSkillPromptMessage + sendCustomMessage, never modify systemPrompt or use contextFiles
Constraint: triggerTurn must be false to avoid extra agent turns
Rejected: append to systemPrompt | agent cannot distinguish skill content from own instructions
Rejected: contextFiles injection | agent sees opaque file blob, cannot discover sub-resources
Rejected: promptCustomMessage per skill | N extra agent turns with model inference
Directive: autoload skill names are resolved against parent session skill list at spawn time in task/index.ts
Tested: TypeScript compiles clean with tsc --noEmit
Tested: parseAgentFields parses array and CSV string frontmatter
Tested: parseAgentFields returns undefined for absent and empty fields
Not-tested: bun test cannot run locally due to missing pi_natives native addon (requires Rust toolchain)
Confidence: high
Scope-risk: moderate
Reversibility: clean
2026-05-24 19:35:40 +04:00
can1357 1228c96959 feat(coding-agent): added worktree list/clear CLI with orphan pruning
- Added the new `omp worktree` (`wt`) command with `list|clear`, `all/dry-run/json` options, and CLI registration.
- Added `listWorktrees`/`clearWorktrees` flows that scan worktrees, classify orphaned entries, emit JSON, and call `worktree.prune`.
- Replaced legacy path encoding with `hashPath` via `getWorktreeDir`, updating task isolation, storage keys, and PR checkout paths.
- Added bounded PR worktree path retries before `git worktree add` and updated checkout-path tests for hashed names.
2026-05-22 12:47:13 +09:00
can1357 6b671ff1c5 fix(coding-agent): enforce subagent output schema on yield
buildOutputValidator already exists in task/executor.ts but only ran on
the fallback JSON parse path. Subagents that called yield directly with
a data payload skipped validation entirely, so a schema-conforming-
empty-object (e.g. `{}` against a schema requiring `findings`) was
returned as a successful task.

finalizeSubprocessOutput now invokes the validator on every yield path
and on the fallback completion path. On failure the result carries
error="schema_violation", a typed message, the missing required field
list, and a truncated preview of the offending data. exitCode=1 and
isError=true so existing consumers in task/index.ts surface it as a
failed task without code changes.
2026-05-21 15:22:46 +09:00
can1357 6d1ade1200 fix(coding-agent): only claim Applied patches: yes when something applied
Both empty-patch and whitespace-only-patch branches in the isolated-task
merge unconditionally appended "Applied patches: yes". The marker now
fires only when the combined patch is non-empty AND applyText succeeded.
No-op merges report "No changes to apply.". The renderer marker matcher
accepts the new string.
2026-05-19 19:24:16 +09:00
can1357 be5451d5c3 fix(coding-agent): surface real jobId in background task announcement
The "Started N background task jobs" announcement listed the per-task
label (e.g. 7-MyScout) without the underlying jobId, but JobTool.execute
looks up by jobId. Append (job: ${jobId}) to each task entry and accept
the task label as a fallback alias in JobTool.execute.
2026-05-19 19:24:16 +09:00
can1357 0ed6edb8e2 fix(coding-agent): hide spawn-disabled agents from the task <agents> list
renderDescription filtered the listing only by disabledAgents. When a
subagent had parentSpawns empty the model still saw the full menu, tried
to spawn oracle, and got "Cannot spawn 'oracle'. Allowed: none". The
listing now intersects with the allowed-spawn set.
2026-05-19 19:24:16 +09:00
can1357 c39295eaac feat(coding-agent/prompts): added oracle advisory agent prompt and embedded registration
- Added a new `oracle.md` prompt defining a read-only diagnostic, architecture, and debugging advisor.
- Imported `oracle.md` into the agent prompt loader in `packages/coding-agent/src/task/agents.ts`.
- Included `oracle.md` in `EMBEDDED_AGENT_DEFS` so the new agent is available at runtime.
2026-05-19 05:33:43 +02:00
can1357 64fcdc308f refactor(coding-agent)!: removed StringEnum helper and shortened tool schema descriptions
- Replaced all StringEnum(...) usages with z.enum([...]) across tools, examples, and tests.
- Removed StringEnum re-export from @oh-my-pi/pi-coding-agent public API.
- Condensed verbose tool parameter descriptions to minimal lowercase phrases.
- Renamed AuthCredentialStore to SqliteAuthCredentialStore at usage sites.
2026-05-16 19:26:32 +02:00
can1357 88e5e8a9a2 fix(coding-agent/task): make subagent runtime timeout sticky and abortable through setup
- Runtime-limit timeout is now tracked with a sticky runtimeLimitExceeded flag, so later caller aborts during teardown cannot downgrade the timeout state and let a late yield report success.
- Awaited setup operations before the first prompt are now raced against the subagent abort signal, including auth discovery, model refresh, model resolution, session open, session creation, extension session_start, prompt, and waitForIdle.
2026-05-15 18:31:12 +02:00
can1357 c4672f11af fix(coding-agent/task): plugged wall-clock timer holes and propagated context stats
- Added defensive abort re-check after registering the abortSignal listener plus a checkAbort() immediately before await session.prompt(...), so a wall-clock timer that fires during pre-prompt setup is no longer lost between listener registration and the prompt call.
- Late yield events arriving after a wall-clock timeout no longer flip the result to success: a runtimeLimitExceeded flag derived from the internal abortReason forces wasAborted=true and exitCode=1 regardless of hasYield, while yield payloads remain captured in extractedToolData.
- Async task progress now copies contextTokens and contextWindow from the completed SingleResult onto AgentProgress, so backgrounded tasks still surface their context gauge to the UI.
2026-05-15 17:49:00 +02:00
can1357 45fe4df39e fix(ai): corrected AI tool handling via JSON-schema validation flow
- Replaced fromTypeBox conversion with a JSON-schema validator flow in ai tool handling and execution paths.
- Added recursive schema validation and expanded TypeBox checks for refs, enums, uniqueItems, and constraint keywords.
- Sanitized Azure/CCA tool schemas by dropping unsupported fields and rewriting oneOf tool branches as anyOf.
- Tightened argument and model-config validation, preserving unknown tool fields and adding apiKey plus compatibility flags.
2026-05-15 15:16:50 +02:00
can1357 2867e1f4e3 feat(deps): added pi.zod exports and removed TypeBox package exports
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
2026-05-15 14:46:54 +02:00
can1357 005c3cd81a feat(coding-agent/task): added subagent runtime timeout and context-window progress reporting
- Added a task.maxRuntimeMs setting with a default disabled state for per-subagent runtime limits.
- Updated subprocess execution to enforce the configured wall-clock timeout, mark timeouts as aborts, and include timeout-specific abort messaging.
- Tracked per-turn context size and context window through task progress/results and updated UI renderers to display current context against window with cumulative tokens rendered separately.
2026-05-15 14:46:54 +02:00