- Loaded recall, reflect, and retain tool descriptions from new markdown prompt files.
- Updated the retain tool to accept batched `items`, enqueue each memory entry, and return the queued-memory count.
- Reworked retain tests to pass the array payload, assert pluralized queue messages, and confirm single-batch flush behavior.
- Added per-session retain queues with size/time auto-flush, recursive drain, and lifecycle flushes on end/clear/enqueue.
- Changed `hindsight-retain` to validate session state, enqueue writes, and return `Memory queued.` immediately.
- Added `notice` event support and handlers that route error/warning/status messages with source-aware formatting.
- Updated client and tests with shared request mapping, `RequestOptions`, `buildMemoryItem`, and expanded batch/list/doc APIs.
- Added an optional `aliasOf` link and a `pickPrimaryState` helper to track reusable parent Hindsight state.
- Updated task-depth subagent startup to reuse the latest primary state's bank, client, config, tags, and missions while returning silently when no parent exists.
- Skipped alias states in recall-snippet selection and enqueue so auto-recall/retain work only on the parent session.
- Added inline hashline parse and apply support for `<` prepend and `+` append operations with prefix+suffix edits.
- Added fail-fast behavior to reject inline modify ops combined with delete or replace on same line.
- Renamed HASHLINE_* and mode symbols to HL_* in prompt tooling, read/search checks, and prompt templates.
- Standardized separators to `PI_HL_SEP`/`HL_EDIT_SEP` and fixed `HL_BODY_SEP='|'`, updating parser formatting behavior.
- Updated benchmark subtype constants and python cleanup test setup to use HL_* values and AgentRegistry mock failure injection.
- Replaced duplicated plugin-registry cache invalidation blocks with clearPluginRootsAndCaches in command and selector setup paths.
- Updated OMP plugin registry path resolution to use getPluginsDir for reads and cache invalidation, matching marketplace write locations.
- Removed a redundant project-scope marketplace test after centralizing cache-root invalidation logic.
- Removed the `@vectorize-io/hindsight-client` dependency from both package manifests.
- Added a local `HindsightApi` client and changed `createHindsightClient` to instantiate it.
- Replaced `HindsightClient` imports/types with `HindsightApi` across backend, bank, and tests.
- Implemented `HindsightApi` `#request` handling with `fetch`, safe JSON parsing, and `HindsightError` on failures.
- Added configurable hashline separator support via `PI_HASHLINE_SEP` with `|` fallback.
- Replaced hardcoded `|` payload markers with `{{hsep}}`/`$HSEP$` in prompts, grammar, and parsing.
- Updated payload parsing to strip shared separator prefixes while preserving whitespace-only prefixes.
- Replaced `resolveLarkLidPlaceholders` with `resolveHashlineGrammarPlaceholders` and added a compatibility alias.
- Updated `collectPayload` to accept a warnings accumulator, detect doubled payload prefixes, and strip one extra leading `|` when all payload lines used the doubled form.
- Changed `parseHashlineWithWarnings` to return collected parser warnings instead of an empty list.
- Added tests confirming auto-stripping, single-line `|` preservation, and mixed-prefix payload behavior.
- Added `hindsight.scoping` and `HINDSIGHT_SCOPING` with global, per-project, and per-project-tagged modes.
- Migrated legacy `raw.hindsight` handling to scoping-based behavior and removed deprecated `dynamicBankId`/`agentName` keys.
- Replaced `deriveBankId` with `computeBankScope`, returning `bankId` plus optional retain/recall tags by mode.
- Forwarded scoped tags through memory ops so recall and retain/reflect calls use `tags`, `tagsMatch`, and `recallTags`.
- Validated `scoping` values from env/settings, warning and defaulting to `per-project-tagged` when invalid.
- Standardized Hindsight context output to emit `<memories>` blocks and updated retention stripping to remove both `<memories>` and legacy `<hindsight_memories>`/`<relevant_memories>` tags.
- Renamed built-in Hindsight tool entry points and labels from `hindsight_retain`/`hindsight_recall`/`hindsight_reflect` to `retain`, `recall`, and `reflect`, including tool availability checks.
- Updated runtime memory config loading and backend resolution so `memory.backend === "local"` now enables the local pipeline directly, while `memories.enabled` is treated only as a legacy migration input.
- Switched the `memory.backend` default to `off` and kept `memories.enabled` hidden from the Memory tab UI for migration compatibility.
- Adjusted memory runtime, resolver, and documentation/tests to match the new backend-selection semantics.
- Added a new optional `beforeAgentStartPrompt` hook to `MemoryBackend` and implemented it in the Hindsight backend to recall long-term context for the first turn.
- Updated `AgentSession` startup flow to inject the recalled context into the turn-specific system prompt before the first response is generated.
- Preserved `<hindsight_memories>` tags in Hindsight developer instructions and added tests for first-turn injection and state caching.
- Added `memory.backend` and `hindsight.*` settings schema with migration from `memories.enabled` legacy mode.
- Added Hindsight memory backend runtime modules for resolved config, client creation, bank ID derivation, and state lifecycle.
- Added off/local/hindsight backends and resolver wiring across SDK, commands, and compaction context.
- Added `hindsight_recall`, `hindsight_reflect`, and `hindsight_retain` tools with schema validation and backend gating.
- Added Memory tab metadata and symbols to expose backend selection in the settings UI.
- Added package export barrels and tests for bank ID, content formatting, and hindsight config env precedence.
- Updated search, find, ast-edit, and ast-grep to skip missing paths and run on existing ones.
- Added multi-path error handling to throw ToolError only when all provided paths are missing.
- Updated search and find outputs to surface skipped `missingPaths` in text and renderer warnings.
- Updated CHANGELOG with multi-path tool behavior updates and `search_repos` global-search notes.
- Added `PartitionedPaths` and `partitionExistingPaths` to classify existing versus missing path inputs.
- Added test helpers and fixtures for multi-path missing-path cases in search/find behavior tests.
- Added `search_code`, `search_commits`, and `search_repos` to GitHub tool schema, docs, and titles.
- Extended `GithubTool.execute` and search arg building to dispatch new search ops and handle args correctly.
- Added richer search formatters for code, commit, and repo results with short SHA and repository metadata.
- Added tests for code, commit, and repo search outputs and `--repo` argument behavior.
- Status-line rendering now treats `statusLine.sessionAccent` as disabled only when explicitly set to false.
- Updated status-line overflow tests to assert gap colors use session accent when enabled and theme border when disabled.
- Test teardown now restores `WSL_INTEROP` and `WSL_DISTRO_NAME` environment variables after mutation.
- Initialized in-memory settings during test setup and reset them after tests.
- Updated the status-line border assertion to verify border glyph text and styling directly.
- Removed reliance on session accent helper functions in the session-accent test case.
- Added a new boolean statusLine.sessionAccent setting to settings schema and propagated it through status-line preview, controller, and component setting updates.
- Updated status line rendering and interactive border coloring to disable session-based accent colors when the setting is false.
- Added a regression test verifying the status-line gap uses theme border color instead of session accent when session accents are disabled.
Fixes#918
POSIX permits short syscalls (incl. open/read/stat) to be interrupted
by a signal and surface as EINTR. Optional sync git metadata helpers
(used by the status line) propagated the raw error and crashed the
agent. Add a small bounded retry around the sync stat/read calls and
classify persistent EINTR as 'metadata unavailable' so optional reads
fall back to undefined instead of throwing.
Fixes#899
The --list-models handler in runRootCommand short-circuited to
listModels() right after Settings.init and modelRegistry.refresh,
exiting before extension loading ran in createAgentSession. As a
result, providers contributed via pi.registerProvider() (from -e
paths or settings.extensions) never appeared in the listing.
Extract a runListModelsCommand entry point in cli/list-models.ts
that loads extensions (CLI -e paths and settings.extensions) into
the supplied ModelRegistry, mirroring sdk.ts's handoff of pending
provider registrations, and then delegates to listModels. The load
is intentionally narrow: no agent loop, no MCP servers, no custom
tools.
Fixes#905
- Updated `githubToolRenderer.renderCall` to use operation-specific titles and contextual metadata.
- Fixed `githubToolRenderer.renderResult` fallback path to show the real operation and clearer status messages.
- Changed non-`run_watch` fallback output to truncate lines by terminal width and show `+N more lines` hints.
- Switched search/ast_grep/ast_edit/find inputs from scalar path fields to required `paths` arrays.
- Reworked path resolution to normalize and expand each `paths` entry via explicit helpers in `src/tools/path-utils.ts`.
- Updated search and find tool-call rendering to display explicit `paths` values in mode overlays and export views.
- Updated tool prompt docs and examples to document `paths` array inputs for search, grep, find, and AST tools.
- Raised the default search match limit from 20 to 500 and updated limit-reached messaging.
- Centralized `line-hash.ts` hash regex sources and resolved `atom.lark` via `resolveLarkLidPlaceholders`.
- Expanded `computeLineHash` to emit `>[a-z]` and `[a-z]<` hashes for brace-context anchors.
- Replaced atom/hashline parsers' hard-coded lid regex with shared `HASHLINE_HASH_RE_SRC` and lax counterparts.
- Removed `\\TEXT` continuation handling in atom rewrites and switched multi-line replacements to `+TEXT`.
- Added brace-body insertion warning when `@Lid` on `{`-ending lines inserts at non-body-safe indent.
- Suppressed duplicate auto-rebase warnings and kept unmatched `-`/`+` ranges separate in compact previews.
- Removed deprecated file helper APIs (find, glob, grep, rgrep, sed, and stat) from JS and Python eval preludes.
- Removed status icons and formatting branches for find/grep/rgrep/glob/stat/sed from tools/eval.ts.
- Updated eval helper docs and Python prelude tests to match the reduced exposed helper surface.
- Removed PI_STRICT_EDIT_MODE gating from edit-mode resolution so model fallbacks now always apply.
- Stopped injecting PI_STRICT_EDIT_MODE in edit-benchmark.py and rate-edit-tool.py execution environments.
- Removed PI_STRICT_EDIT_MODE from environment-variable documentation and strict-mode test coverage.
- Added branch-aware session loading so autoresearch state only rehydrates for current branch.
- Replaced user-specified experiment commands with fixed `bash autoresearch.sh` execution flow.
- Enforced safer setup checks, including missing `autoresearch.sh` and uncommitted-worktree errors.
- Added branch-specific storage helpers, baseline-commit persistence, and expanded tests for dirty-path cases.
- Adjusted atom mutation conflict validation to skip throwing on repeated delete edits for the same anchor line.
- Added tests confirming duplicate delete edits on one anchor are idempotent and do not trigger conflicts.
- Added coverage ensuring explicit deletes within replace ranges are treated as redundant and ignored.
- Added `disableReasoning` to `SimpleStreamOptions` and OpenAI completions, sending `reasoning: { enabled: false }` for OpenRouter requests to prevent reasoning models from consuming the full output budget on small calls like title generation.
- Fixed `canAppend` to accept `response.completed` as a terminal event, restoring websocket append reuse after codex sessions end.
- Replaced async blob-decoding and `addEventListener` with synchronous `onmessage`/`onopen`/`onerror`/`onclose` handlers and `binaryType = "nodebuffer"` for simpler, reliable message decoding.
- Simplified title generator to discard per-role thinking level and always pass `disableReasoning: true`.
- Removed title-source aware branching from session terminal-title and accent helpers, and updated callers to use session name plus cwd only.
- Dropped UUID-based recent-session naming by preferring explicit header titles or first user prompts and generating an "Untitled · <time>" fallback.
- Adjusted welcome session-row rendering for width-aware name truncation and disabled reasoning in title generation requests to keep terminal titles concise.
- Replaced file-backed autoresearch contracts with sqlite-backed session/run storage in `~/.omp/autoresearch`.
- Added `AutoresearchStorage` and rewired `init_experiment`, `run_experiment`, and `log_experiment` to persist sessions and runs.
- Added `update_notes` tool with `body`/`append_idea` inputs and updated prompts to use active-session context.
- Removed `autoresearch.md` contract parsing and checks flow, including `runChecks`, `force`, and timeout schema options.
- Updated autoresearch state/types to persist `goal`, `notes`, `branch`, and `baselineCommit` plus run justification/flag metadata.
- Fixed path resolution to emit per-target fanout when common base collapses to filesystem root, preventing full-filesystem scans.
- Fixed match/file counts and pagination to aggregate correctly across all targets.
- Fixed returned paths to be normalized relative to the original search scope.
- Defined ASI as an object with explicit `hypothesis`, `rollback_reason`, and `next_action_hint` fields while allowing additional keys.
- Updated `validateAsiRequirements` to clarify guidance when ASI data is missing or missing a valid hypothesis.
- Adjusted autoresearch state tests to assert the revised ASI validation error messages.
- Consolidated AI provider imports through register-builtins and moved Gemini/Antigravity header helpers to a shared module.
- Added lazy loading for heavy providers and SDK-backed modules with cached initialization to trim startup cost.
- Converted markdown conversion helpers to async and awaited htmlToBasicMarkdown in affected scraper and kernel output paths.
- Parsed bundled agent definitions on-demand and moved BrowserTool prompt rendering behind a memoized getter.
- Added cached validation/error handling paths by replacing AJV runtime checks with Value.Check and trimming validation error output.
- Updated AGENTS.md discovery to use glob search honoring .gitignore, depth limits, and deduped results.
- Updated eval tool flow so Python preflight runs only when needed and exec now maps to eval when available.
- Deferred canonical model-index rebuilds during refresh/rebuildProvider and replayed pending rebuilds after resume.
- Added memoized model-equivalence resolution with trailing-marker and canonical reference caches.
- Optimized frontmatter key normalization to keep unchanged keys/arrays/objects without extra cloning.
- Updated JS executor tests to use base-path concatenation for nested fixture filesystem calls.
The warmup path no longer produces prelude docs, so the cached-session
warmup it implemented added no value over the create-on-first-execute
path that withKernelSession already covers. Remove warmPythonEnvironment,
the backend warm() hook, the eval-tool warmup loop, the createTools
warmup preflight, and the forcePythonWarmup option. Simplify
ExecutorBackendCallOptions into ExecutorBackendExecOptions since execute
is now the only consumer.
Manual handoff starts a fresh session and seeds it with a displayed custom handoff message, not an assistant message. Session persistence normally waits for an assistant message before creating the session file, which made the new handoff session exist only in memory until later activity.
Persist the seeded handoff session after injecting the handoff context, and record the previous session file as its parent so lineage remains discoverable.
Drives AgentSession + Agent + BashTool end-to-end through the patched
pi-natives binding. Asserts a spawned child reports its own session id
(setsid was called) and that pipelines still produce both stages'
output. Validated by reverting commands.rs and confirming the test
fails with a named diagnostic before restoring.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
- Added `rename_file` LSP action with path checks, capped pair enumeration, and apply preview flow.
- Added `request` LSP action with auto-built params, optional JSON `payload`, and method-not-found handling.
- Added `capabilities` LSP action to inspect server capabilities for a file or all servers.
- Added schema updates and client file-operation capabilities to advertise rename support.
- Added regression tests for rename behavior, preview mode, and request/capabilities flows.
The signature in #computeAppliedToolSignature hashed the full raw
instructions string, but rebuildSystemPrompt truncates each server
instruction at 4000 chars before embedding it. A change past character
4000 produced an identical prompt but a different signature, causing a
spurious rebuild and a cache miss on every such reconnect.
Fix: hoist MAX_MCP_INSTRUCTIONS_LENGTH to module scope in sdk.ts and
apply the same truncation in the getMcpServerInstructions callback
before returning. The session now hashes exactly the strings that end
up in the prompt.
Regression test: changes only past char 4000 do not trigger a rebuild;
changes within the first 4000 chars do.
buildSystemPrompt injects today's date into the prompt body. The
applied-tool signature skips rebuilds when tools are byte-identical,
but did not cover the date — so a session spanning midnight with only
tool-stable MCP reconnects would keep yesterday's date indefinitely.
Append the current YYYY-MM-DD date as a suffix to the signature so any
reconnect after midnight triggers exactly one rebuild, then resumes
skipping normally for the rest of the new day.
- Documented and removed `utils/oauth` from the `ai` package entrypoint, noting it as a breaking change.
- Refactored `cli`, `auth-storage`, and `utils/oauth` to load provider modules via scoped dynamic `import()` calls.
- Removed top-level provider imports and barrel exports from `utils/oauth/index.ts`, streamlining oauth module loading.
- Consolidated OAuth symbol, type, and provider imports in coding-agent and tests to `@oh-my-pi/pi-ai/utils/oauth` modules.
- Defined `DEFAULT_LOCAL_TOKEN` locally in model-registry and removed its cross-package OAuth import usage.
Built-in tools whose prompt-rendered metadata depends on settings
(`TaskTool`, `SearchToolBm25Tool`, `EditTool`) expose `description`/
`label` via getters that re-evaluate on every access. The skip
optimization in `#applyActiveToolsByName` is correctness-safe for these
because `#computeAppliedToolSignature` reads `tool.description` live
each call, so a settings flip mutates the rendered string and differs
the signature on the next refresh.
This contract was implicit; a future refactor that caches per-tool
description strings would silently break it. Defending it explicitly:
- Added a regression test that wires a getter onto a CustomTool's
`description`, verifies `refreshMCPTools` skips while the underlying
state is unchanged, then mutates the state (without changing tool
object identity) and verifies the rebuild fires.
- Expanded the `#computeAppliedToolSignature` docstring to document
the getter-based coverage path and the SDK-init-time closure
constants in `sdk.ts` that genuinely cannot change at runtime
(`repeatToolDescriptions`, `eagerTasks`, `intentField`,
`mcpDiscoveryEnabled`, `secretsEnabled`).
Triggered by a review question on whether the skip breaks settings-
based prompt changes. It does not, but the property is non-obvious.