- Added `loadOverallPlanReference` to resolve a session plan reference from local storage and skip empty or missing files.
- Updated task execution to read the active plan reference (except in plan mode) and pass it into each spawned subagent.
- Extended the subagent system prompt and session SDK/tools plumbing so subagents receive and render the approved plan path and contents.
- Renamed `TodoWriteTool` to `TodoTool` and its source/prompt files.
- Updated tool registration, schema, renderers, and gating to `todo`.
- Adjusted cursor provider native tool names and tests to match.
- Renamed strike-animation constants and `todo-error-reminder` type.
The catch around the subagent yield-reminder prompt previously logged
every exception at ERROR. User cancel (^C) and compaction-driven aborts
both surface as ToolAbortError through awaitAbortable, so benign control
flow generated 9 spurious 'Subagent prompt failed' errors in 2 days on
the reporter's instance.
Gate the ERROR branch on '!abortSignal.aborted && !(err instanceof
ToolAbortError)' and route the abort path to logger.debug. The outer
catch + finally still mark the run aborted, so observable behaviour is
unchanged.
Fixes#1623
- Converted [SECTION]...[/SECTION] markers to "SECTION\n===" format in system prompt templates.
- Updated system conventions doc to reference the new marker style.
- Updated tests to match against the new header pattern.
- Added parentMnemosyneSessionState propagation from session state through SDK, executor, and task options into nested sessions.
- Added getMnemosyneSessionState() and rekeying logic to refresh Mnemosyne IDs during session sync, switch, and restore.
- Added Mnemosyne reset and teardown cleanup on unaliasing or restoration to avoid stale state.
- Added `mnemosyne.scoping` setting: `global`, `per-project`, and `per-project-tagged`.
- `per-project-tagged` writes to a project-local bank while merging global memories on recall.
- Refactored `MnemosyneSessionState` to manage scoped recall/retain targets and deduplication.
- Updated hindsight tools to route recall/retain through scoped methods.
- In `resolveApproval`, yolo mode now returns the user policy directly (`allow`/`prompt`/`deny`) and ignores tool `override` prompts.
- Updated approval-mode and approval unit tests to match the new behavior for critical bash patterns under yolo and auto-approve.
- Updated docs and settings metadata to describe yolo as user-policy-driven rather than override-driven.
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
- approval: user 'tool: deny' now wins over critical-pattern override
(the override only tightens allow->prompt; it must never re-arm a denied tool).
- approval: rename hindsight policy keys to match registered tool names
(recall/retain/reflect, not hindsight_recall/hindsight_retain).
- approval: head+tail truncation for bash/ssh command prompts so a
destructive suffix buried after a long benign preamble stays visible.
- task/executor: force tools.approvalMode='auto' in createSubagentSettings
so subagents (which have no UI) cannot deadlock on per-tool prompts;
the parent's approval of the task call is the authorization.
- docs/approval-mode: rewrite so every example surfaces tools.approvalMode
and explains that tools.approval is ignored outside 'custom' mode.
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
- Unified output schema construction and validation by adding buildOutputValidator and using it in YieldTool and task executor.
- Added MAX_SCHEMA_RETRIES so YieldTool now retries schema failures three times with hints before overriding.
- Updated failure handling to use shared summarizeValidationFailure and formatters for required-field reporting.
- Added tests for output-schema-validator and YieldTool covering malformed schemas and nested-array retry edge cases.
Queued extension-delivered user messages when deliverAs is set and waited for session_start extension message sends before prompting subagents.
Fixes#1343
- Captured `tool_execution_update` snapshots for `task` calls into in-flight progress state for live nested rendering.
- Cleared in-flight task snapshots at task start and completion to prevent stale nested progress from persisting.
- Updated progress rendering to combine completed and in-flight task details through a dedicated nested task tree view.
The report_finding tool's priority is exposed as a string enum
("P0"-"P3") for ergonomics, but the reviewer agent and every
custom review agent declare priority as `type: number` in their
JTD output schema. The cast at executor.ts:1473 lied about the
runtime shape, so the auto-injected `findings[].priority` flowed
through as strings and every yield with at least one finding was
rejected with `findings.0.priority: expected number, received string`,
forcing the run into the schema_violation exit path.
Added `toReviewFinding(details)` in tools/review.ts that maps the
priority enum to its numeric ordinal via the existing PRIORITY_INFO
table and use it at the boundary in executor.ts. Render paths still
see the original `ReportFindingDetails` shape (string priority)
through normalizeReportFindings, so display formatting is unaffected.
Fixes#1350
- Added `retry.maxDelayMs` to the settings schema and interfaces, with a default cap for provider backoff delays.
- Updated session auto-retry logic to fail fast when a requested wait exceeds the cap without fallback, emitting terminal auto-retry failure state.
- Propagated retry state and failure data into task progress and rendering so children show retry/wait details and reminder prompts stop after terminal errors.
Adds optional autoloadSkills field to agent frontmatter that automatically loads listed skills when a sub-agent is spawned. Uses the same buildSkillPromptMessage + sendCustomMessage mechanism as interactive skill loading, queued via sendCustomMessage({ triggerTurn: false }) before the first session.prompt(task). No extra agent turns, no new injection path. Skills stay in listing for sub-resource access. Compaction behavior matches manual loading. Unknown skill names silently skipped.
Lore-id: f85fdbdc
Constraint: autoload must use buildSkillPromptMessage + sendCustomMessage, never modify systemPrompt or use contextFiles
Constraint: triggerTurn must be false to avoid extra agent turns
Rejected: append to systemPrompt | agent cannot distinguish skill content from own instructions
Rejected: contextFiles injection | agent sees opaque file blob, cannot discover sub-resources
Rejected: promptCustomMessage per skill | N extra agent turns with model inference
Directive: autoload skill names are resolved against parent session skill list at spawn time in task/index.ts
Tested: TypeScript compiles clean with tsc --noEmit
Tested: parseAgentFields parses array and CSV string frontmatter
Tested: parseAgentFields returns undefined for absent and empty fields
Not-tested: bun test cannot run locally due to missing pi_natives native addon (requires Rust toolchain)
Confidence: high
Scope-risk: moderate
Reversibility: clean
buildOutputValidator already exists in task/executor.ts but only ran on
the fallback JSON parse path. Subagents that called yield directly with
a data payload skipped validation entirely, so a schema-conforming-
empty-object (e.g. `{}` against a schema requiring `findings`) was
returned as a successful task.
finalizeSubprocessOutput now invokes the validator on every yield path
and on the fallback completion path. On failure the result carries
error="schema_violation", a typed message, the missing required field
list, and a truncated preview of the offending data. exitCode=1 and
isError=true so existing consumers in task/index.ts surface it as a
failed task without code changes.
- Runtime-limit timeout is now tracked with a sticky runtimeLimitExceeded flag, so later caller aborts during teardown cannot downgrade the timeout state and let a late yield report success.
- Awaited setup operations before the first prompt are now raced against the subagent abort signal, including auth discovery, model refresh, model resolution, session open, session creation, extension session_start, prompt, and waitForIdle.
- Added defensive abort re-check after registering the abortSignal listener plus a checkAbort() immediately before await session.prompt(...), so a wall-clock timer that fires during pre-prompt setup is no longer lost between listener registration and the prompt call.
- Late yield events arriving after a wall-clock timeout no longer flip the result to success: a runtimeLimitExceeded flag derived from the internal abortReason forces wasAborted=true and exitCode=1 regardless of hasYield, while yield payloads remain captured in extractedToolData.
- Async task progress now copies contextTokens and contextWindow from the completed SingleResult onto AgentProgress, so backgrounded tasks still surface their context gauge to the UI.
- Replaced fromTypeBox conversion with a JSON-schema validator flow in ai tool handling and execution paths.
- Added recursive schema validation and expanded TypeBox checks for refs, enums, uniqueItems, and constraint keywords.
- Sanitized Azure/CCA tool schemas by dropping unsupported fields and rewriting oneOf tool branches as anyOf.
- Tightened argument and model-config validation, preserving unknown tool fields and adding apiKey plus compatibility flags.
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
- Added a task.maxRuntimeMs setting with a default disabled state for per-subagent runtime limits.
- Updated subprocess execution to enforce the configured wall-clock timeout, mark timeouts as aborts, and include timeout-specific abort messaging.
- Tracked per-turn context size and context window through task progress/results and updated UI renderers to display current context against window with cumulative tokens rendered separately.
- Task tool sessions now expose and forward parent OpenTelemetry config when creating subagent tasks.
- Subprocess execution now derives child telemetry from the parent config with the subagent identity and child session conversation handling.
- Subagent creation now records a handoff span using the resolved parent telemetry handle before running the child loop.
- Added a shared session command helper that aggregates extension, prompt, and skill slash commands.
- Updated ACP, extension UI, runtime-init, and task executor extension contexts to return session command data instead of empty arrays.
Token counter (token_total status-line segment, subagent progress tree,
session-observer stats line) previously included cacheRead in its cumulative
sum. With Anthropic prompt caching, cacheRead per turn equals the full cached
context, so summing across N turns gives N*context_size -- a session with a 1M
context and 5 turns showed ~5M tokens despite no compaction occurring.
Fix: display shows input + output + cacheWrite per turn. cacheWrite is kept
because each byte is written once; cacheRead re-reads the same context every
turn. Dedicated cache_read/cache_write status-line segments still show cache
activity; billing cost is unaffected.
Also adds per-subagent cost display (dollar amount, statusLineCost color)
accumulated incrementally from message_end events. Hidden when cost is zero
(subscription/OAuth providers). Brings token and cost display in line with
what Claude Code shows per-agent.
Adds `pi.on("credential_disabled", handler)` so extensions can react to
soft-disabled credentials (e.g. OAuth invalid_grant) without regex-matching
`agent_end` errorMessages.
`AuthStorage.onCredentialDisabled(listener)` returns an unsubscribe function;
multiple listeners fire for every event with per-listener exception isolation
and FIFO buffer-and-replay (cap 32) when none are attached. The constructor
option from #991 stays as sugar for an immediate permanent subscription.
`createAgentSession()` subscribes the per-session extension runner to
`modelRegistry.authStorage` immediately after resolution and unsubscribes on
dispose / startup failure. Events are forwarded via
`ExtensionRunner.emitCredentialDisabled(event)`, which buffers (cap 32,
drop-oldest) until `runner.initialize(...)` runs in the mode controller so
extension handlers see real UI/runtime context, not the constructor no-op
defaults.
Supersedes #997. Builds on #991.
Co-Authored-By: omp <noreply@oh-my-pi.dev>
- Added parent-to-subagent artifact manager adoption so subagents reuse the parent `ArtifactManager` and write artifacts into a shared directory with shared IDs.
- Passed the shared artifact manager through tool/session context into subagent executor startup and exposed it via `SessionManager` and `ToolSession` for lookup.
- Updated kernel environment and artifact-resolution paths to prefer `PI_ARTIFACTS_DIR`, falling back to existing session-file-based behavior when absent.
- Skipped writing compact conversation context files when IRC is enabled so subagents avoid using stale markdown snapshots.
- In runSubprocess, passed an undefined context file for IRC-enabled paths and kept context file prompting for non-IRC executions.
Two-edge fix for a runtime circular-import TDZ that manifested as
`ReferenceError: Cannot access 'TaskTool' / 'SUBAGENT_WARNING_*' /
'MAX_OUTPUT_BYTES' before initialization.` whenever the executor module
graph and the task tool module graph were evaluated together (e.g.
`bun test executor-warnings.test.ts task-simple-mode.test.ts` in either
order, or the wider `task/ discovery/ task-simple-mode` combination).
The runtime cycle is
task/index.ts → ./executor → ../sdk → ./tools → ../task
closed by `tools/index.ts:285` eagerly dereferencing `TaskTool.create`
while the `task` module's body had not yet reached its `export class
TaskTool` declaration. The throw aborted `tools/index.ts`, which
propagated back through `sdk → executor`, leaving executor's body
suspended before its post-import `const`s (`SUBAGENT_WARNING_*`,
`MAX_OUTPUT_BYTES`) were initialized.
Two structural changes:
1. `tools/index.ts:285`: replace `task: TaskTool.create` with
`task: s => TaskTool.create(s)`. Defers the `TaskTool` binding
dereference to factory-call time, by which point the cycle has
fully unwound. Matches the lazy-factory shape every other
`BUILTIN_TOOLS` entry already uses.
2. `task/executor.ts:31`: split the `"../tools"` import. Source
`truncateTail` directly from its leaf module
`../session/streaming-output` (the barrel was just re-exporting
it). Keep `ContextFileEntry` as a type-only import — erased at
runtime, so no participation in the cycle.
Neither change touches tests, test runners, or removes the cycle in
source. They eliminate the eager dereferences that turned a benign
linker-level cycle into a TDZ at evaluation time.
Verified:
- `bun test executor-warnings.test.ts task-simple-mode.test.ts`
passes in both orderings (11/11).
- Wider `bun test packages/coding-agent/test/task/
packages/coding-agent/test/discovery/
packages/coding-agent/test/tools/task-simple-mode.test.ts` now
100/100 (was 14 fail / 1 error).
- `bun --cwd=packages/coding-agent run check` clean apart from the
pre-existing unrelated `anthropic.ts:1206 'stop_details'` error
from 4e0ca3c0e.
Co-Authored-By: omp <noreply@oh-my-pi.dev>
- Added explicit prompt markers and wrappers across system templates, including `[env]`, `[role]`, `[coop]`, `[closure]`, and `[now]`.
- Removed `renderTemplate` and `sectionSeparator` flows, deleted `task/template.ts`, and switched to per-task `renderSubagentUserPrompt` rendering.
- Updated system prompt assembly to `shortenPath`-normalize `cwd`, append rendered now metadata, and preserve trailing `[now]` blocks.
- Removed legacy template tests and added prompt-composition tests for ordered `[contract]`->`[project]`->`[now]` blocks and context-only system placement.
- Reworked shared prompt utilities by collapsing consecutive blank lines and removing obsolete `OPENING_HBS`/`LIST_ITEM` helper behavior.
- Added `listWorkspace` native binding and API types, exporting bounded workspace trees with AGENTS.md candidates.
- Reworked `buildWorkspaceTree` and `buildDirectoryTree` to call `listWorkspace` with 5s timeout defaults.
- Replaced startup AGENTS.md discovery with workspace-tree-only scanning and removed legacy AgentsMdSearch session plumbing.
- Updated `WorkspaceTree` and system prompt context to expose `agentsMdFiles` and aligned tests/changelog expectations.
- Introduced a flag to detect when an explicit modelRegistry was supplied to runSubprocess.
- Skipped modelRegistry.refresh() when reusing the parent registry and retained refresh when creating a new one.
- Added a debug log to indicate when a parent modelRegistry is reused and refresh is bypassed.
Reporter screenshot showed a parent session on DeepSeek V4 Pro dispatching
a task subagent that resolved to `qwen3.6-plus-free` — an opencode-zen
model the user had no working credentials for. The dispatch hit a
provider that could not serve the model and surfaced a confusing API
rejection instead of using the parent's already-authenticated model.
Adds `resolveModelOverrideWithAuthFallback`, an auth-aware wrapper
around `resolveModelOverride` that checks the resolved subagent model's
credentials via `modelRegistry.getApiKey` + `isAuthenticated` and
falls back to the parent session's active model pattern when the
primary has no working auth. The parent's active model is plumbed
through `ExecutorOptions.parentActiveModelPattern` from `TaskTool`
into `runSubprocess`. If neither has working auth (or they resolve to
the same model), the primary resolution is preserved so the existing
error path still surfaces a meaningful failure downstream.
Fixes#985
- Updated system prompts to require every active turn to end with a tool call and clarified that reminders should default to resuming work unless the task was complete or genuinely blocked.
- Reworked the yield-reminder text to disallow fake blocker reasons and to permit continued tool-calling instead of forced immediate yields.
- In `runSubprocess`, applied `toolChoice` only on the final yield retry by adding an `isFinalRetry` check before sending the reminder.
Subagents previously re-ran buildAgentsMdSearch and buildWorkspaceTree on
every spawn, repeating the slowest part of system-prompt construction for
each task tool invocation. On large/pathological repos those scans
exceeded the 5s preparation deadline and tripped the per-subagent
'system prompt preparation timed out' warning.
Forward the parent's already-resolved AgentsMdSearch and WorkspaceTree
through createAgentSession (alongside the existing contextFiles, skills,
and promptTemplates inheritance):
- Add agentsMdSearch and workspaceTree to CreateAgentSessionOptions;
createAgentSession short-circuits the parallel scan promises when
these are provided.
- Resolve them with contextFiles before constructing ToolSession; expose
on ToolSession so the task tool can read the parent's values.
- Thread them through ExecutorOptions (task/executor.ts) into the
subagent's createAgentSession call, and pass them from the task tool
(task/index.ts) on both the worktree-isolated and non-isolated paths.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Added `HindsightSessionState` to `AgentSession` and bound hindsight lifecycle hooks to session state.
- Removed global hindsight state/queue handling and replaced it with per-session `HindsightRetainQueue` batching and scoped flushing.
- Reworked recall, reflect, and retain tools to use `session.getHindsightSessionState()` instead of sessionId-based lookup.
- Updated SDK/task/backend/controller flows to pass `session`/`parentHindsightSessionState`, scope `/memory` behavior, and document it in changelog.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
- Added a localProtocolOptions field to session and executor option types for configurable local:// behavior.
- Passed localProtocolOptions through to LocalProtocolHandler creation and subagent session bootstrap.
- Propagated parent session local protocol settings from TaskTool so subprocess subagents share the same local:// artifacts and session context.
- Updated `formatMatchLine` to emit `*` for matched lines, a leading space for context, and a `|` anchor/content separator.
- Revised grep/hashline mismatch messages and prompts to describe the new marker and separator format.
- Aligned affected atom and hashline tests with the updated match-line prefixes and separators.