- Dropped the terminal split artifact so repeated saves no longer grow a
blank line per write.
- Replaced the partition-and-concat merge with an ordered line list:
headings, prose, and footers keep their positions; new lessons insert
at the head of the first bullet run and the cap trims oldest bullets.
- Strengthened tests to byte-exact idempotence and mixed-Markdown order.
- Review follow-up for PR #6774.
- Serialized backend transitions across runtime state, tools, and prompts.
- Rehydrated Mnemopi listeners after clear and enqueue maintenance.
- Made memory.backend the sole post-migration local runtime gate.
Fixes#5638
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Use the pre-session memory prompt snapshot as the learned.md baseline when startup consolidation refreshes before a session-scoped cache exists. This keeps active-session learn writes out of the refreshed prompt while still surfacing the new consolidated summary.
Refs #3743
Refresh the startup consolidation summary without rereading learned.md for the active session, so lessons captured while startup is still running remain deferred to the next session.
Refs #3743
Background memory startup writes memory_summary.md after the initial system-prompt build has already cached an empty value for the session. Drop the cached snapshot before refreshBaseSystemPrompt so the active session actually picks up the freshly consolidated summary instead of returning the stale cached one.
Refs #3743
Stopped passive autolearn from adding hidden conversation messages and froze local memory developer instructions per session so learn writes land in future sessions instead of mutating the active Anthropic prompt prefix.
Fixes#3743
- Added idle recap functionality to EventController to display goals and next actions when the agent is inactive.
- Updated session parsing in gc-cli and memories to correctly support optional title entries in session files.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
`truncateByApproxTokens` appends a truncation marker, so a truncated memory
summary can exceed `limit * 4` chars and drive the remaining lessons budget
negative. Clamp to 0 and skip lessons entirely when the summary already fills the
budget, instead of feeding a negative limit to the truncator.
Addresses review thread on PR #2542 (thread 15).
The `learn` tool previously required a `hindsight`/`mnemopi` backend. It now
also works when `memory.backend` is `local` (the file-based rollout backend):
lessons append to a `learned.md` under the project's memory root, kept separate
from the consolidation artifacts so a consolidation pass never clobbers them,
and are injected into future sessions alongside the memory summary.
- memories: `saveLearnedLesson` (newest-first, deduped, count- and per-field
size-capped, secret-redacted, injection-neutralized) with per-path write
serialization; `buildMemoryToolDeveloperInstructions` reads `learned.md` and
shares one injection budget with the summary; `redactSecrets` extended with
GitHub/npm/Slack/Google token prefixes.
- local backend: implements `save()`; status reports `writable: true`.
- learn tool: `local` execute branch; `createIf`/`isToolAllowed`/auto-include
and the standing guidance extended to `local`; local saves tier as a `write`
approval.
- read-path prompt: renders the learned-lessons block when present.
- Lessons are injection-neutralized and secret-redacted on BOTH write and read
(they render unescaped into the system prompt).
Also moves the auto-learn CHANGELOG entry out of the released [15.12.6] section
(a cherry-pick artifact) back under [Unreleased] and notes the local backend.
Tests: local storage (format, dedup, cap, redaction incl. provider/delimiter-
split tokens, concurrency), read-back (with/without summary, off-gating, raw
hand-edited file), tool gating + write-approval tiering.
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
Concurrent omp --session restores after an unclean shutdown crashed
in SqliteAuthCredentialStore.#initializeSchema() with
SQLITE_BUSY_RECOVERY because the multi-statement schema run installed
PRAGMA busy_timeout=5000 AFTER PRAGMA journal_mode=WAL, the first
lock-taking statement during WAL recovery. Bun's default busy_timeout
is 0, so the lock conflict surfaces immediately.
- packages/ai/src/auth-storage.ts: hoisted PRAGMA busy_timeout to a
standalone first statement, dropped it from the multi-statement
schema run, wrapped SqliteAuthCredentialStore.open() in a 4-attempt
exponential-backoff retry loop on the SQLITE_BUSY family, and the
exhausted-retry error now includes the DB path. Exported
isSqliteBusyError(err) (matches code prefix 'SQLITE_BUSY').
- packages/coding-agent/src/session/agent-storage.ts: same hoist and
the existing retry loop now uses isSqliteBusyError so
SQLITE_BUSY_RECOVERY / _SNAPSHOT / _TIMEOUT also trigger backoff.
- Hoisted busy_timeout before journal_mode=WAL in every other shared
SQLite open path: history-storage, autoresearch/storage,
memories/storage, github-cache, report-tool-issue (auto-QA),
catalog/model-cache; stats/db.ts now sets busy_timeout at all.
- packages/ai/test/auth-storage-sqlite-busy.test.ts pins the contract:
isSqliteBusyError matches every BUSY extended code (rejects
SQLITE_LOCKED, non-errors, strings); open() leaves the connection in
WAL mode (proves busy_timeout ran before journal_mode); open() retries
through synthetic SQLITE_BUSY_RECOVERY; non-BUSY errors (SQLITE_CORRUPT)
short-circuit; exhausted retries throw an error mentioning the DB path
with exactly 3 sleeps for a 4-attempt budget.
Fixes#2421
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.
Fixes#2198
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Added first-party-first provider priority defaults for model ranking.
- Consolidated model resolution to use getModelMatchPreferences from session settings.
- Prioritized providerPriorityRank ahead of usage rank when picking preferred models.
- Added second-pass fallback to default-model or API-key-valid matching order.
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
The memory pipeline hardcoded `Effort.Low` (stage1) and `Effort.Medium` (phase2
consolidation) when calling `completeSimple`. On models whose supported efforts
exclude those levels (e.g. `deepseek/deepseek-v4-pro` → [high, xhigh]),
`completeSimple → mapOptionsForApi → resolveOpenAiReasoningEffort →
requireSupportedEffort` threw "Thinking effort low is not supported by
<provider>/<model>" and every stage1 job was recorded as failed, blocking phase2
and producing no memory artifacts.
Route both call sites through `clampThinkingLevelForModel(model, requested)` —
the same helper already used by compaction (#1182). For `[high, xhigh]` both
`low` and `medium` lift to `high`; non-reasoning models continue to receive
`undefined`, preserving prior behaviour.
Fixes#1480
Anthropic counts sessions by metadata.user_id. Without this fix, OMP
generated fresh random entropy on every API request, inflating the
session count and preventing backend attribution to the authenticated
account.
Changes:
packages/ai:
- resolveAnthropicMetadataUserId() now accepts JSON-format user_id
matching real Claude Code's getAPIMetadata shape
({ session_id, account_uuid, ... }). Previously only the legacy
cloaking format was accepted on OAuth, causing stable caller-supplied
values to be silently discarded.
- AnthropicOAuthFlow.exchangeToken() and refreshAnthropicToken() now
populate OAuthCredentials.{accountId, email} from the token response
account block, removing the need for a separate /api/oauth/profile
round-trip.
- AuthStorage.getOAuthAccountId(provider, sessionId) returns the OAuth
accountId for the session-sticky credential, used to build
account_uuid in metadata.user_id. Guards against misattribution for
API-key, runtime-override, env-key, and fallback-resolver paths that
do not record a session credential.
packages/agent:
- Agent.metadataForProvider(provider) resolves request metadata for
the given provider via the installed resolver, or returns the static
metadata value. The plain metadata getter now returns only the static
value; provider-aware resolution is explicit.
- Agent.setMetadataResolver(fn) installs a (provider: string) resolver
evaluated per LLM request in agent-loop, after getApiKey records the
session-sticky credential, so account_uuid reflects the credential
actually used.
- AgentLoopConfig.metadataResolver is called with config.model.provider
after getApiKey, overriding the static metadata field.
packages/coding-agent:
- AgentSession.#syncAgentSessionId installs a metadata resolver that
builds { user_id: JSON.stringify({ session_id, account_uuid? }) },
matching the Anthropic session attribution format. account_uuid is
only included for provider="anthropic" to avoid leaking the OAuth
identity to third-party Anthropic-format-compatible providers.
- sessionId getter prefers providerSessionId when supplied via
AgentSessionConfig so all API paths (getApiKey, direct calls,
metadata resolver) share the same provider-facing session ID.
- prepareSimpleStreamOptions stamps session metadata on direct calls
(runEphemeralTurn, compaction, branch summary, title generation) so
they share the same session bucket as Agent.prompt requests.
- generateBranchSummary and generateSessionTitle accept a
(provider: string) metadata resolver evaluated after their own
getApiKey call for correct credential attribution.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Updated runtime memory config loading and backend resolution so `memory.backend === "local"` now enables the local pipeline directly, while `memories.enabled` is treated only as a legacy migration input.
- Switched the `memory.backend` default to `off` and kept `memories.enabled` hidden from the Memory tab UI for migration compatibility.
- Adjusted memory runtime, resolver, and documentation/tests to match the new backend-selection semantics.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
Phase1 caught every per-claim failure, recorded the reason in
jobs.last_error, and surfaced only an aggregate failed count via the
phase1 completion debug line. Users hitting setup-time failures (e.g.
WSL2 stale rollout paths producing ENOENT before any LLM call) had no
diagnostic. Emit logger.error per failed claim with threadId,
rolloutPath, and reason so the actual error is visible in omp.log.
Fixes#846
- Renamed the built-in `grep` content-search tool to `search` across settings, schemas, and SDK exports.
- Switched execution wiring so `Task`, `Plan`, cursor, and shell mapping now invoke `search` instead of `grep`.
- Updated prompts, plan-mode docs, and example tool lists to replace `grep`/`ls` references with `search` guidance.
- Aligned `Grep*`/`grep` event, renderer, and hook types to `Search*`/`search` across runtime and tests.
- Documented and fixed `search` result rendering budget behavior and added internal-URL/path-list transcript notes.
- Canonicalized file and CLI defaults from `read` to `open` across tool registration and prompts.
- Added `resolveToolAlias()` and applied alias-normalized tool selection so legacy `read` maps to `open`.
- Updated runtime, UI, and export layers to treat `open` as first-class while preserving `read` compatibility.
- Renamed read prompt docs to `open.md`/`open-chunk.md` and refreshed system guidance to recommend `open`.
- Updated tool-related tests and expectations from `read` to `open` (including test fixtures and aliases).
- Added canonical model equivalence types, cache helpers, and registry APIs for provider variant lookup.
- Changed model resolution to apply canonical ID overrides/excludes with provider order before fallback matching.
- Added canonical and provider model views in list-models and selector UI with canonical sorting/persistence.
- Updated role/model persistence to store selectors while runtime now resolves concrete canonical-backed provider models.
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).
* feat(utils): full XDG Base Directory support for all path helpers
Implement XDG-first resolution across all omp path helpers, extend the
migration command to cover every data/state/cache location, and fix
five data-safety issues found in review.
dirs.ts:
- Add getXdgCachePath() helper ($XDG_CACHE_HOME/omp/<subpath>)
- Add isDefaultAgentDir() helper: XDG lookup is only valid when the
resolved agentDir equals the process default (~/.omp/agent); custom
profiles set via PI_CODING_AGENT_DIR or setAgentDir() are never
silently redirected to the global XDG database
- Update 15 functions to XDG-first resolution:
data: getPluginsDir, getRemoteDir, getRemoteHostDir, getPythonEnvDir,
getWorktreeBaseDir
state: getReportsDir, getSshControlDir, getCrashLogPath, getDebugLogPath
cache: getPuppeteerDir, getGpuCachePath, getNativesDir
- Guard XDG lookup with isDefaultAgentDir(agentDir ?? getAgentDir()) in
all 9 agent-subdir helpers so that callers passing the global default
agentDir still resolve to the migrated XDG location, while callers
passing a non-default agentDir or running under a custom profile via
setAgentDir() bypass XDG entirely
- Plugin-derived helpers delegate to getPluginsDir() and follow XDG
resolution automatically
migrate-xdg.ts:
- Add getXdgCacheHome() helper
- Extend MigrationItem.category to include 'cache'
- Add 12 new migration entries: reports, plugins, remote, ssh-control,
remote-host, python-env, puppeteer, wt, gpu_cache.json, natives,
omp-crash.log, omp-debug.log
- Refuse to run when PI_CODING_AGENT_DIR points to a non-default
profile: migration only makes sense for the default ~/.omp/agent tree
- copyDirectory returns skipped source paths (target existed, non-force)
- verifyIntegrity: remove size-mismatch early-return that masked stale
targets as successful copies
- executeMigration: delete only entries that were actually copied;
use rmdir on source dir so it is removed only when empty, preserving
any skipped files for a subsequent --force run
- executeMigration: rename partial target to <target>.bak on integrity
failure instead of deleting; preserves pre-existing user data while
preventing getXdgDataPath from treating the partial tree as
authoritative; source remains intact for re-copy on next run
test isolation:
- Set XDG_DATA_HOME/XDG_STATE_HOME to non-existent paths in
memories-runtime.test.ts beforeEach/afterEach to prevent
getXdgDataPath/getXdgStatePath from resolving to real user data
* fix(utils,coding-agent): fix XDG support issues
dirs.ts:
- Refactor path resolution into DirResolver class. XDG base dirs are
resolved once at construction from env vars (Linux only, no
existsSync). setAgentDir creates a fresh instance, naturally
invalidating all cached paths and recomputing isDefaultProfile.
- getRootSubdir/agentSubdir accept optional XdgCategory parameter;
when set, the XDG base replaces the config root. Every accessor
is a one-liner delegate.
- Non-Linux platforms: XDG fields are null, zero overhead. No
filesystem probing, no string comparisons on the hot path.
- Config-only subdirs (themes, tools, commands, prompts, modules)
have no XDG category — they stay under the config root.
- Remove `import { env } from 'bun'`, use process.env consistently.
- Restore JSDoc comments to document actual defaults (~/.omp/...).
migrate-xdg.ts:
- Gate migrateToXdg on Linux — exits with clear error on other
platforms.
- Fix data loss bug: verifyIntegrity now accepts a Set of skipped
paths and skips verification for files that were intentionally not
copied (pre-existing at target in non-force mode).
- Fix nested directory source deletion: recursive removeSourceEntries
walks the tree and only deletes files not in the skipped set.
- Remove dead _sourcePath variable and unused force parameter from
verifyIntegrity.
- Remove `import { env } from 'bun'`, use process.env consistently.
logger.ts:
- Revert JSDoc to document ~/.omp/logs/ as default.
oauth.ts:
- Replace direct getAgentSubdir call with getTestAuthPath().
CHANGELOG.md:
- Merge duplicate section headers under [Unreleased].
- Add missing blank line before [13.11.1].
---------
Co-authored-by: can1357 <me@can.ac>
* fix(memories): isolate Phase 2 consolidation per project working directory
The global memory consolidation job used a single 'global' job key,
causing all projects to share one Phase 2 slot. Whichever project
claimed it first got every project's stage1 outputs written into its
memory directory — cross-project contamination.
Root cause:
- GLOBAL_KEY = 'global' was a single key shared by all projects
- listStage1OutputsForGlobal() had no cwd filter — returned ALL
stage1 outputs across all projects
- markGlobalPhase2Succeeded/Failed/Unowned all operated on the same
single global job row
Fix:
- Replace GLOBAL_KEY constant with globalJobKey(cwd) function that
namespaces the job key per project: 'global:/path/to/project'
- Add cwd parameter to all Phase 2 storage functions so each project
maintains its own job slot in the jobs table
- Filter listStage1OutputsForGlobal() by t.cwd = ? so each project
only consolidates its own thread outputs
- Thread cwd from session.sessionManager.getCwd() through runPhase2()
to all storage calls
- Add cwd to markStage1SucceededWithOutput/NoOutput so enqueueGlobal
Watermark targets the correct per-project job key
- Add cwd parameter to enqueueMemoryConsolidation() public API and
update its only external caller (command-controller)
Closes#369
* test(memories): add isolation tests for per-project Phase 2 consolidation
Three tests covering the regression fixed in the previous commit:
- listStage1OutputsForGlobal filters outputs by cwd (no cross-project leak)
- enqueueGlobalWatermark creates separate job rows keyed per-project
- tryClaimGlobalPhase2Job claims only the requested project's slot and
leaves the other project's job independently claimable
* docs: add inline comments for memory isolation fix and tests
---------
Co-authored-by: Rens Tillmann <rens@super-forms.com>
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Consolidated @oh-my-pi/pi-utils subpath imports into single package root import across 100+ files.
- Moved tryParseJson utility from local web scrapers module to @oh-my-pi/pi-utils package for centralized JSON parsing.
- Renamed loadSkillsFromDir to scanSkillsFromDir and refactored skill discovery to use fs.promises.readdir instead of glob-based approach.
- Replaced custom parseJSON with tryParseJson across discovery modules for consistent error handling.
- Removed emitCustomToolSessionEvent method and cleanupSshResources function, consolidating shutdown logic into dispose method.
- Updated glob pattern construction to use GlobBuilder with literal_separator(true) for improved path handling.
- Standardized XML tag naming from snake_case to kebab-case across 50+ prompt files for consistency.
- Replaced imperative language with RFC 2119 keywords (MUST/SHOULD/MAY/MUST NOT) throughout system and tool prompts for clarity.
- Removed artifactsDir parameter from Python executor and simplified environment variable handling to use PI_SESSION_FILE only.
- Renamed read_path.md to read-path.md and updated memory guidance with hierarchy rules and conflict resolution workflow.
- Added noEscape option to bash URL expansion and extracted cwd parameter from leading cd commands for improved path handling.
- Exported NO_PAGER_ENV constant from bash-interactive module for centralized environment variable management.
- Modified memory storage to isolate memories by project working directory, preventing cross-project memory contamination.
- Added encodeProjectPath function to safely encode working directory paths for use in memory storage directory names.
- Updated getMemoryRoot function to accept cwd parameter and include encoded project path in memory directory structure.
- Updated clearMemoryData function signature to accept cwd parameter for project-specific memory cleanup.
- Updated test fixtures to use project-aware memory root paths.
- Added autonomous memory extraction and consolidation system with two-phase pipeline for extracting durable knowledge and consolidating into reusable skills.
- Added /memory slash command with subcommands (view, clear, reset, enqueue, rebuild) for inspecting and managing memory maintenance.
- Added configurable memory settings including concurrency limits, lease timeouts, token budgets, and rollout age constraints.
- Added memory injection payload that includes learned context in system prompts with operational rules for reading and using memory artifacts.
- Improved system prompt building to support memory guidance injection and enhanced error handling in resolvePromptInput function for multiline input.