- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Added DirectoryTree, DirectoryTreeOptions, and buildDirectoryTree exports for configurable tree rendering.
- Changed read tool directory output to use buildDirectoryTree with depth and exclusion limits.
- Added read.summarize settings and summary-mode parseable read output behavior when no selector is used.
- Added tests for truncated root/child listings and hidden or excluded entry filtering.
- Added `WorkspaceTree` and `buildWorkspaceTree` APIs for working-directory tree rendering with limits.
- Extended `buildSystemPrompt` and `createAgentSession` to resolve and pass workspace tree context for system prompts.
- Updated the system prompt template to include a `<workspace-tree>` section with truncation notices before append output.
- Added workspace-tree and system-prompt tests covering sorting, truncation, exclusions, and prompt ordering.
- Added `HindsightSessionState` to `AgentSession` and bound hindsight lifecycle hooks to session state.
- Removed global hindsight state/queue handling and replaced it with per-session `HindsightRetainQueue` batching and scoped flushing.
- Reworked recall, reflect, and retain tools to use `session.getHindsightSessionState()` instead of sessionId-based lookup.
- Updated SDK/task/backend/controller flows to pass `session`/`parentHindsightSessionState`, scope `/memory` behavior, and document it in changelog.
- Added `memory.backend` and `hindsight.*` settings schema with migration from `memories.enabled` legacy mode.
- Added Hindsight memory backend runtime modules for resolved config, client creation, bank ID derivation, and state lifecycle.
- Added off/local/hindsight backends and resolver wiring across SDK, commands, and compaction context.
- Added `hindsight_recall`, `hindsight_reflect`, and `hindsight_retain` tools with schema validation and backend gating.
- Added Memory tab metadata and symbols to expose backend selection in the settings UI.
- Added package export barrels and tests for bank ID, content formatting, and hindsight config env precedence.
- Consolidated AI provider imports through register-builtins and moved Gemini/Antigravity header helpers to a shared module.
- Added lazy loading for heavy providers and SDK-backed modules with cached initialization to trim startup cost.
- Converted markdown conversion helpers to async and awaited htmlToBasicMarkdown in affected scraper and kernel output paths.
- Parsed bundled agent definitions on-demand and moved BrowserTool prompt rendering behind a memoized getter.
- Added cached validation/error handling paths by replacing AJV runtime checks with Value.Check and trimming validation error output.
The warmup path no longer produces prelude docs, so the cached-session
warmup it implemented added no value over the create-on-first-execute
path that withKernelSession already covers. Remove warmPythonEnvironment,
the backend warm() hook, the eval-tool warmup loop, the createTools
warmup preflight, and the forcePythonWarmup option. Simplify
ExecutorBackendCallOptions into ExecutorBackendExecOptions since execute
is now the only consumer.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
The signature in #computeAppliedToolSignature hashed the full raw
instructions string, but rebuildSystemPrompt truncates each server
instruction at 4000 chars before embedding it. A change past character
4000 produced an identical prompt but a different signature, causing a
spurious rebuild and a cache miss on every such reconnect.
Fix: hoist MAX_MCP_INSTRUCTIONS_LENGTH to module scope in sdk.ts and
apply the same truncation in the getMcpServerInstructions callback
before returning. The session now hashes exactly the strings that end
up in the prompt.
Regression test: changes only past char 4000 do not trigger a rebuild;
changes within the first 4000 chars do.
Two cache-stability fixes for Anthropic prompt caching during MCP server
reconnects, which happen routinely (~5 min per server) in long sessions
due to SSE transport keepalive timeouts.
1) MCPManager: deterministic tool ordering
`#tools` is now sorted by name after every mutation. The previous
filter-out + push-to-end pattern in `#replaceServerTools` moved the
reconnecting server's tools to the end of the array, producing a new
byte order whenever the reconnect sequence differed from the initial
discovery sequence. With multiple healthy servers, each reconnect of
the non-last server flipped the order and invalidated the tools
cache breakpoint sent to Anthropic.
Sort applies in `discoverAndConnect` (initial population) and
`#replaceServerTools` (used by `reconnectServer` and
`refreshServerTools`). The comparator is character-code based,
locale-independent and deterministic. `sortMCPToolsByName` is
exported as a small generic helper and unit-tested.
2) AgentSession: skip system-prompt rebuild when inputs are unchanged
`#applyActiveToolsByName` (called from `refreshMCPTools` after every
reconnect) used to unconditionally call `rebuildSystemPrompt` and
`setSystemPrompt` even when the resulting prompt was byte-identical.
This wasted CPU on every flap and risked silent cache invalidation
if the rebuild path ever became non-deterministic.
Now `#applyActiveToolsByName` computes a stable signature of the
inputs `rebuildSystemPrompt` reads and skips the rebuild when the
signature matches the last successful one. The signature covers:
- active tool names in render order
- active tool labels and descriptions (rendered as `{{label}}:
\`{{name}}\`` in the prompt body)
- when MCP discovery is on, every registry tool's name + label +
description (the prompt summarizes discoverable-but-inactive
MCP tools)
- per-server MCP `instructions` text (embedded under "## MCP
Server Instructions" in the appended prompt; can change on
server upgrade while tool list stays identical)
Server instructions are read via a new optional
`getMcpServerInstructions` callback on `AgentSessionConfig`, wired
from the SDK as `() => mcpManager.getServerInstructions()`.
`refreshBaseSystemPrompt()` continues to rebuild unconditionally and
refreshes the cached signature, so explicit refreshes still pick up
ambient changes (edit-mode toggles, memory writes, etc.) that the
signature does not cover.
Signature inputs deliberately NOT covered: tool input schemas, memory
instructions read from disk, and other ambient state. Callers that
mutate those must call `refreshBaseSystemPrompt()` explicitly; existing
hooks (`#syncEditToolModeAfterModelChange`, memory hooks, `/clear`)
already do.
- Parallelized startup by deferring plugin preload and running AGENTS.md scan plus context/template/command discovery in parallel.
- Added AgentsMdSearch exports and options so prebuilt search results were passed into system-prompt construction.
- Reworked logger timing to use AsyncLocalStorage-backed nested spans, initialize a root span, and emit hierarchical summaries.
- Added PI_TIMING-gated TS/TSX module-load timing via side-effect module-timer registration and wrapped key init/request paths with logger.time.
Add optional defaultLevel to ThinkingConfig schema/type so models.yml can
declare a preferred starting thinking level per model. On model switch
the agent session adopts model.thinking.defaultLevel when present (with
explicit caller-supplied level still winning); otherwise current behavior
is preserved. SDK initial selection prefers the model's defaultLevel
before falling back to the global defaultThinkingLevel setting.
Fixes#775
- Added a localProtocolOptions field to session and executor option types for configurable local:// behavior.
- Passed localProtocolOptions through to LocalProtocolHandler creation and subagent session bootstrap.
- Propagated parent session local protocol settings from TaskTool so subprocess subagents share the same local:// artifacts and session context.
- Renamed MCP tool IDs from `mcp_<server>_<tool>` to `mcp__<server>_<tool>`, and changed built-in `grep` to `search`.
- Updated `parseMCPToolName()` and bridge helpers to require and trim the `mcp__` prefix.
- Updated cursor, manager, and session discovery flows to require `mcp__`-prefixed tool names.
- Updated MCP tests and assertion fixtures to use `mcp__`-prefixed tool IDs and expected system prompts.
- Renamed the built-in `grep` content-search tool to `search` across settings, schemas, and SDK exports.
- Switched execution wiring so `Task`, `Plan`, cursor, and shell mapping now invoke `search` instead of `grep`.
- Updated prompts, plan-mode docs, and example tool lists to replace `grep`/`ls` references with `search` guidance.
- Aligned `Grep*`/`grep` event, renderer, and hook types to `Search*`/`search` across runtime and tests.
- Documented and fixed `search` result rendering budget behavior and added internal-URL/path-list transcript notes.
- Added agent identity and registry fields to session configuration and session creation, enabling relay routing metadata for agent sessions.
- Implemented non-persistent IRC relay emission to forward incoming and reply observations from non-main agents into the main session UI.
- Updated IRC UI rendering to support `irc:relay` messages with participant-aware arrow formatting and body display.
- Updated `formatMatchLine` to emit `*` for matched lines, a leading space for context, and a `|` anchor/content separator.
- Revised grep/hashline mismatch messages and prompts to describe the new marker and separator format.
- Aligned affected atom and hashline tests with the updated match-line prefixes and separators.
- Added `AgentRegistry` singleton with session registration/unregistration and IRC routing metadata for peer lookups.
- Added IRC messaging prompts and tooling with `irc.enabled` setting, `list/send` tool paths, and peer roster rendering.
- Changed `/btw` to session-side `runEphemeralTurn`, added background IRC exchange flushing, and fixed empty-input checks.
- Added unit tests for IRC tool and BtwController ephemeral behavior, including disabled, busy, not-found, and abort cases.
- Added `openai` and `openai-codex` as image providers and let `providers.image=auto` prefer GPT images.
- Updated settings, selector, and SDK wiring so OpenAI image providers pass through `setPreferredImageProvider`.
- Replaced Gemini-only image tooling with `image-gen` and added OpenAI/Codex hosted-image execution with SSE parsing.
- Added image-gen and handoff tests, including final-yield no-compaction regression and OpenAI payload/header assertions.
- Renamed subagent completion flow from `submit_result` to `yield` across SDK tools, prompts, and docs.
- Updated executor/task handling to require and parse `yield` calls, replacing legacy submit-result extraction and state flags.
- Added `subagent-yield-reminder` and updated system prompts to require `yield` with `result.data` or `result.error`.
- Renamed hidden-tool and registration plumbing to `yield`, including discovery helpers and renderer/test surface.
- Canonicalized file and CLI defaults from `read` to `open` across tool registration and prompts.
- Added `resolveToolAlias()` and applied alias-normalized tool selection so legacy `read` maps to `open`.
- Updated runtime, UI, and export layers to treat `open` as first-class while preserving `read` compatibility.
- Renamed read prompt docs to `open.md`/`open-chunk.md` and refreshed system guidance to recommend `open`.
- Updated tool-related tests and expectations from `read` to `open` (including test fixtures and aliases).
Subagents created with enableMCP=false had toolSession.mcpManager
unset, causing depth-2+ sub-subagents to re-discover and spawn
duplicate MCP server processes.
- Add mcpManager option to CreateAgentSessionOptions
- Set toolSession.mcpManager unconditionally after MCP block
- Pass options.mcpManager from executor to createAgentSession
- Guard callback registration: only wire onToolsChanged/onPromptsChanged/
onResourcesChanged when the session owns the manager (created via
discovery), not when reusing a parent's — prevents child sessions
from clobbering the parent's live MCP refresh handlers
Latent since 91da560cc (in-process subagent migration), observable
since f82a5d121 added task.maxRecursionDepth allowing depth-2 agents.
- Removed the standalone vim tool and normalized built-in/requested tooling to edit.
- Updated session and SDK tool activation to dedupe lowercase names and track edit state via the edit key.
- Added vim-mode argument detection and delegated edit rendering/execution into Vim handlers under edit.
- Updated Vim step handling to auto-reorder numeric-positioned commands, including cc/C/S/s/i/I/A cases.
- Renamed prompt/changelog text and test expectations to reflect edit-only tool naming and usage.
- Removed `SearchDb` APIs and `searchDb` fields, dropping db-backed state from native and agent sessions.
- Replaced crate export `fff` with `fd`, moving fuzzy-find bindings into `fd.rs`.
- Removed `SearchDb`/picker fast-path logic from `glob` and `grep`, simplifying scan flow and dropping db args.
- Removed `SearchDb`/`getSearchDb` wiring from extension, tool, and task context constructors across coding-agent.
- Added over-indentation validation warnings in chunk-edit normalization for suspicious `~` body line formatting.
- Removed `bytes`, `fff-grep`, `fff-search`, and `blake3` deps, adding `grep-searcher = "0.1"`.
- Updated `AgentSession` and `createAgentSession` to resolve active edit tool names via `resolveEditToolName`.
- Filtered inactive edit variants with `filterInactiveEditToolName` so tool listings expose the active edit mode.
- Normalized requested edit-capable tool names with `normalizeToolNamesForEditMode` during activation and startup.
- Synced edit-tool mode after model changes so the active tool swaps between `edit` and `vim` automatically.
- Enabled per-file partial updates by threading `onUpdate` through `EditTool` and `executePerFile`.
- Updated `ToolExecutionComponent` to render multi-file edit `perFileResult` entries as separate boxes with pending state.
fixed retained-kernel restart and owner cleanup edge cases during recovery and disposal
tracked async user_python hooks during disposal-sensitive execution paths and hardened startup warmup tracking
strengthened cleanup and kernel lifecycle regressions to remove deadlocks, false positives, and timing flakes
scoped retained-kernel ownership to agent sessions and cleaned it up on session disposal
cleaned up warmed python owners on session startup failure and rejected new direct and tool-based python starts during disposal, including async hook, preflight, and warmup races
- Added bash.autoBackground.enabled and bash.autoBackground.thresholdMs settings with defaults for background job behavior.
- Added auto-backgrounding support for long foreground bash commands with managed job state and timeout threshold.
- Updated bash prompts, job-protocol messages, and tool activation checks to use async and auto-background support.
- Added background bash completion integration with AsyncJobManager and end-to-end tests for short/long auto-background scenarios.
- Added canonical model equivalence types, cache helpers, and registry APIs for provider variant lookup.
- Changed model resolution to apply canonical ID overrides/excludes with provider order before fallback matching.
- Added canonical and provider model views in list-models and selector UI with canonical sorting/persistence.
- Updated role/model persistence to store selectors while runtime now resolves concrete canonical-backed provider models.
- Migrated language classifiers from imperative methods to declarative semantic rule tables with ClassifierTables and StructuralOverrides.
- Extracted node shape analysis into dedicated shape module with field priority constants and helper functions for AST traversal.
- Added schema module with language-aware node metadata and thread-local language context management.
- Centralized environment variable parsing across codebase using $flag(), $envpos(), and isBunTestRuntime() utilities.
- Added PI_CHUNK_AUTOINDENT configuration to control indentation normalization in chunk read/edit operations.
- Enhanced system prompt with instruction priority, output contract, tool persistence, and completeness guidelines.
- Added `/force` slash command and `ToolChoiceQueue` system for managing tool-choice directives with lifecycle callbacks and requeue semantics.
- Added `setForcedToolChoice()`, `peekQueueInvoker()`, `buildToolChoice()`, `steer()`, and `getToolChoiceQueue()` methods to AgentSession and ToolSession APIs.
- Refactored tool-choice override mechanism from simple override to queue-based system with generator directives, lifecycle callbacks, and requeue preservation.
- Removed `PendingActionStore` class and replaced with `ToolChoiceQueue`; updated `ResolveTool` and custom tool loader to use queue invokers.
- Fixed tool-choice queue cleanup on agent loop abort and requeue semantics to preserve callbacks across abort cycles.
- Remove pi-ref.ts deferred barrel import (no longer needed after circular dep fix)
- Update extension/hook/custom-tool/custom-command loaders to use direct imports
- Fix warmPythonEnvironment mock return type in python tool tests (add docs field)
- Made Codex websocket prewarm asynchronous and non-blocking during session creation for faster startup.
- Added Codex websocket status updates in interactive mode when prewarm completes or fails.
- Introduced `getOpenAICodexTransportDetails()` to determine transport preferences before prewarm initialization.
- Refactored Codex prewarm to fire in background without awaiting, matching LSP server warmup pattern.
- Updated CHANGELOG entries for ai and coding-agent packages to reflect Codex startup event monitoring and async prewarm behavior.
- Refactored logger.time() API to accept function references and arguments separately instead of wrapped callbacks.
- Replaced RingBuffer-based timing with wall-clock markers and async span tracking for improved timing accuracy.
- Updated 20+ call sites across executor, kernel, main, and tools modules to use new logger.time() signature.
- Removed parsePath() helper function and inlined path.split() calls in settings module.
- Added test case verifying model selection without auth validation during startup.
- Added asynchronous LSP server discovery and warmup at startup via `discoverStartupLspServers()` function.
- Added LSP startup event channel (`lsp:startup`) for real-time server status notifications during warmup.
- Added `LspStartupServerInfo` type to track LSP server status (connecting, ready, error) throughout lifecycle.
- Changed LSP server warmup from blocking synchronous operation to non-blocking background task.
- Refactored interactive mode to subscribe to LSP startup events and display real-time server status updates.
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).
The four extension/hook/custom-tool/custom-command loaders statically
imported the package barrel `@oh-my-pi/pi-coding-agent` to expose it as
`pi` to user code. This created a self-referential cycle:
tools/index -> task -> sdk -> custom-commands/loader
-> @oh-my-pi/pi-coding-agent (package barrel)
-> modes/components -> tool-execution -> renderers
-> tools/read (TDZ on readToolRenderer)
Triggered at startup by 'omp --help' / 'omp stats --help' depending on
ESM evaluation order, manifesting as:
ReferenceError: Cannot access 'readToolRenderer' before initialization
Introduce `extensibility/pi-ref.ts` that resolves the barrel lazily via
`require` on first call. All four loaders use `getPiRef()` instead of a
static `import * as piCodingAgent`. The barrel is fully initialized by
the time any loader function actually runs, so the lookup is safe.
- Fixed memory leak by cancelling idle compaction timer on event controller disposal.
- Fixed session resumption to preserve last non-empty session when starting fresh.
- Fixed stash detection to use git ref resolution instead of output parsing for reliability.
- Fixed secret obfuscation to deobfuscate restored session messages locally while keeping LLM messages obfuscated.
- Fixed stash pop operation to preserve staged changes with --index flag after task branch merges.
- Changed idle compaction settings from enum to numeric type for flexible configuration.
Wire EventBus through task runtime so subagent lifecycle and progress
events propagate to the TUI. Add SessionObserverRegistry to track
active sessions, and a session observer overlay accessible via Ctrl+S
that shows a picker of running subagents and a read-only transcript
viewer that reads the subagent's session JSONL file to display
thinking, text, tool calls, and results.
- Add TASK_SUBAGENT_LIFECYCLE_CHANNEL for start/end events
- Add SubagentProgressPayload.sessionFile for session file tracking
- Pass eventBus through sdk.ts -> main.ts -> InteractiveMode
- SessionObserverRegistry with multi-listener onChange pattern
- Overlay picker preserves selection position across live refreshes
- Viewer renders full transcript from session JSONL (thinking, text,
tool calls with smart arg summaries, inline tool results)