18 KiB
18 KiB
Changelog
[Unreleased]
[16.3.9] - 2026-07-06
Fixed
- Fixed extractor JSON parsing to correctly unwrap object-shaped facts, instructions, preferences, and timeline items from known text fields instead of persisting literal
[object Object]rows.
[16.3.7] - 2026-07-05
Added
- Added
RecallOptions.contentPreviewCharsto allow customizing or disabling the content preview cap (default is 500, set to 0 for full content). - Added
RecallResult.truncatedandRecallResult.full_lengthproperties to easily identify clipped previews without parsing trailing markers.
Fixed
- Fixed background LLM fact extraction to preserve specific extractor categories (
instructions,preferences,timelines, andkgtriples) in MEMORIA tables and graph triples instead of flattening them into genericfact/entityrows. - Improved recall previews and
factLinecontext to append a trailing ellipsis (…) when content is clipped, preventing mid-word truncation without a marker.
[16.3.5] - 2026-07-04
Fixed
- Fixed
remember(..., { embedText })so hosts can store full transcripts while embedding, FTS-indexing, and rebuild-reembedding a marker-free projection. (#4395)
[16.2.2] - 2026-06-27
Fixed
- Improved resilience during API extraction calls by enhancing the handling of rate limits and transient errors.
[16.1.17] - 2026-06-24
Fixed
- Fixed
remember(..., { extract: true })fact/entity extraction accepting anextractTextoverride so hosts can store full transcripts while mining facts from a safer projection; also tightened deterministicInstruction:extraction to require an explicitI/yousubject instead of treating everyalways/neverclause as a user instruction. (#3372)
[16.1.8] - 2026-06-20
Fixed
- Capped per-input length in
embed()atMNEMOPI_EMBEDDING_MAX_INPUT_CHARS(default 8192 chars, override via the env var orembeddings.maxInputCharsruntime option;0disables) so a long retention transcript can no longer overflow the embedding model's context window. Oversized inputs are clipped with a head/tail split so chronological transcripts keep both the opening setup and the most recent turns instead of losing the latest content under a naive prefix slice. llama.cpp's/embeddingsserver used to reject the request withrequest (N tokens) exceeds the available context size, silently dropping vector recall for that memory (#3126). - Fixed the proactive-linking write path ignoring host configuration:
proactiveLinkIfEnabledreadMNEMOPI_PROACTIVE_LINKINGdirectly, so a host that enabled proactive linking throughconfigureRecallFeatures()had no effect unless the environment variable was also set.proactiveLinkingis now aRecallFeatureFlagsoption resolved through aproactiveLinkingEnabled()fallback, matching the existing polyphonic and enhanced recall flags, with theMNEMOPI_PROACTIVE_LINKINGenvironment variable still taking precedence whenever it is set. (#2440)
[16.1.3] - 2026-06-19
Added
- Exposed
setLocalModelInitializer(and theLocalEmbeddingModel,LocalModelInitializer,LocalModelInitOptions,StandardEmbeddingModeltypes) so hosts can route fastembed loads through a dedicated subprocess and keeponnxruntime-node's NAPI constructor + finalizer out of their own address space. Same wipe semantics as the existingsetLocalModelInitializerForTestsseam; the agent CLI uses it to crash-proof Windows whenmemory.backend: mnemopiis enabled (#3031).
Fixed
- Fixed background fact extraction skipping runtime-configured remote LLM endpoints when
MNEMOPI_LLM_BASE_URLwas unset, soremember(..., { extract: true })now stores remote-distilled facts frommnemopi.llmconfig instead of falling back to regex heuristics. (#3041) - Fixed local fastembed startup on macOS ARM64 by letting
fastembed@2.1.0install its matchingonnxruntime-node@1.21.0native runtime instead of forcing1.26.0, and by repairing missing tokenizer sidecars from the upstream Hugging Face model cache when a stale fastembed archive lacks them. (#3054)
[16.0.6] - 2026-06-18
Fixed
- Forced the on-demand fastembed runtime install to override fastembed's archived
onnxruntime-node@1.21.0transitive pin with Mnemopi'sonnxruntime-node@1.26.0pin, fixing local embedding startup on macOS ARM64. (#2920)
Changed
- Updated OpenRouter request headers to use standard shared headers from the pi-ai package
[16.0.5] - 2026-06-17
Fixed
- Capped
sleep_consolidationepisodic rows atmaxEpisodeChars(default 100KB,MNEMOPI_MAX_EPISODE_CHARS) so raw session transcripts cannot be stored and extracted as multi-megabyte episodes. (#2869) - Skipped regex-only entity and pattern fact extraction for oversized raw transcripts so progress/log noise cannot flood MEMORIA with junk facts. (#2868)
[15.13.1] - 2026-06-15
Added
- Added a wipe-and-rebuild reconcile (
reconcileEmbeddingModel) that runs when the configured embedding model changes. At store open, if the model stamped on storedmemory_embeddingsrows differs from the activecurrentEmbeddingModel(), the stale embeddings and their binary vectors are dropped and every existing memory is enqueued for background re-embedding (in bounded batches) at the new model/dimension. The destructive wipe is skipped whenever it could not be rebuilt — embeddings disabled via the runtime option or theMNEMOPI_NO_EMBEDDINGSenv, an unresolved (empty) active model, or a read-only open (reconcile: false, used by ephemeral stats readers that would exit before the async rebuild finished) — so a stale-but-valid corpus is never destroyed without a replacement. Recall degrades gracefully (FTS-only) for memories whose vectors are not yet rebuilt (#2476)
Fixed
- Normalized enhanced recall fact scoring against lexical coverage so high-confidence facts that only match generic query tokens no longer outrank exact working-memory hits. (#2441)
[15.12.4] - 2026-06-13
Fixed
- Fixed
consolidateToEpisodic(the function backingsleep/sleepAllSessions) never populating the episodic graph: thegistsandgraph_edgestables stayed at 0 rows across every bank even after multiple consolidation cycles, so Polyphonic Recall'sgraphvoice (BFS overfindGistsByParticipant/findRelatedMemories) always returned nothing. Consolidation now best-effort ingests the new episodic memory intoEpisodicGraphso the gist row, gist→memoryctxedge, fact edges, and cross-memory similarity/entity/temporal edges land alongside the episodic row. Independent of the existingMNEMOPI_PROACTIVE_LINKINGflag, which still gates the same enrichment on theremember()write path. (#2435)
[15.12.0] - 2026-06-12
Changed
- Moved
fastembedandonnxruntime-nodefromdependenciesto optionalpeerDependenciespinned to exact versions. When the peers are absent (bundled CLI, compiled binary, or installs that skip optional peers), the local embedding pathbun installs the pinned pair into~/.omp/cache/fastembed-runtime/<version-key>on first use and loads fastembed from there — restoring local embeddings in bundled distributions and removing ~270MB of eager native downloads from default installs (#2389)
[15.11.4] - 2026-06-12
Added
- Added
configureRecallFeatures()(exported from the package root,core, andconfig) so hosts can enable the polyphonic recall engine and the enhanced recall query cache programmatically.polyphonicRecallEnabled(),enhancedRecallEnabled(), andisEnhancedRecallEnabled()now fall back to these configured defaults, with theMNEMOPI_POLYPHONIC_RECALL/MNEMOPI_ENHANCED_RECALLenvironment variables still taking precedence whenever they are set. (#2323)
Fixed
- Fixed the embedding pipeline's silent
catch {}blocks (runEmbedding(),getLocalModel(), and the local-model path ofembed()) swallowing failures with zero diagnostics. These best-effort paths still degrade gracefully (returnnull/ skip the write), but now emit structuredlogger.debugentries with the error and per-site context (item count, model name). Themnemopi.debugconfig flag now propagates into the core library via runtime options (MnemopiOptions.debug→ResolvedMnemopiRuntimeOptions.debug) and escalates these logs towarnso they surface at the default log level. (#2322)
Changed
- Extraction, embedding, and remote-LLM clients now accept an
ApiKey(static string or resolver) and resolve it per request throughwithAuth, so 401s force-refresh and rotate credentials via the central auth-retry policy instead of failing with a stale key. Empty-key setups (local/proxy endpoints withoutAuthorization) and pinned literal keys behave exactly as before. - Embedding and remote-LLM 401 errors now throw pi-ai's typed
ProviderHttpErrorinstead ofObject.assign-patchedErrors, keeping the same structural.statuscontract for the auth-retry classifier. - SHMR consolidation clustering (
core/shmr) now uses the real embedding provider when one is configured instead of always hashing:embed(), the newembedBatch(),clusterBySimilarity(),computeHarmonyScore(),harmonize(), andrecallBeliefs()are now async, batch-embed candidate texts in a single provider call, and reuse precomputed vectors frommemory_embeddingsfor episodic candidates. The SHA1 bag-of-words hash remains as the deterministic fallback when no provider is available or embedding fails. (#2324)
[15.10.12] - 2026-06-10
Changed
- Reworked the in-memory fallback vector search to build a normalized exact vector index per query, matching the shape needed for future quantized or TurboVec-style backends without adding a new dependency yet.
[15.10.11] - 2026-06-10
Fixed
- Fixed embedding provider detection to match
openrouterby URL host, so custom embedding endpoints are now recognized correctly instead of being misclassified by substring matching - Fixed the check for OpenRouter base URLs so only true
openrouterhosts are treated as non-custom
[15.10.8] - 2026-06-09
Added
- Added a
fetchoption toExtractionClientto inject a custom fetch implementation for remote LLM requests - Added an optional
fetchoption toextractFactsto control the transport used for remote extraction calls - Added support for passing a custom
fetchimplementation throughcompleteandsummarizeMemoriesvia remote LLM options
[15.9.1] - 2026-06-04
Breaking Changes
- Changed
Mnemopi.recall(),Mnemopi.recallEnhanced(),Mnemopi.search(),Mnemopi.query(), the module-levelrecall/recallEnhanced/search/queryexports, theBeamMemory.recall/recallEnhancedmethods, the freerecall/recallEnhancedfunctions incore/beam/recall, andorchestrateRecallto returnPromise<RecallResult[]>so the recall pipeline can auto-derivequeryEmbeddingfrom the query text viaembedQuery. Callers mustawaitrecall calls; passqueryEmbedding: nullto opt out of auto-embedding and stay on FTS-only. - Changed the MCP entrypoints
handleToolCall,callToolJson, andhandleJsonRpcinmcp-server/mcp-toolsto async so the recall/shared-recall handlers can await the newPromise<ToolResult[]>shape; external MCP transports mustawaitthese.
Fixed
- Fixed
memory_embeddingsnever being populated by the productionremember/rememberBatch/updateWorking/consolidateToEpisodicpaths; embedding generation is now scheduled as a background task onbeam.pendingExtractions(mirroringscheduleFactExtraction), so configured providers (fastembed, OpenAI-compatible API, custom) actually run and rows land inmemory_embeddings(memory_id, embedding_json, model). (#1832) - Fixed
recall()/recallEnhanced()never deriving a query embedding from the query text, which silently degraded every deployment to FTS-only regardless of provider configuration. The recall pipeline now auto-callsembedQuery(query)whenoptions.queryEmbeddingis undefined; passnullto keep the old FTS-only behaviour. (#1832) - Fixed
toRecallOptionsdroppingqueryEmbeddingbetween theMnemopifacade and the beam layer, so callers can now explicitly pin or disable the query vector through the public API. - Fixed
withMemory(CLI) andwithBeam/withSharedBeam(MCP) closing the SQLite handle before background fact-extraction and embedding tasks finished, so short-livedmnemopi store/mnemopi sleepand MCPremember/updatepaths now drainflushExtractionsbefore close instead of silently droppingmemory_embeddingsrows. CLI handlers and MCPhandleRemember/handleUpdate/handleSleep/etc. are async as a result. (#1832, follow-up to #1833 review) - Fixed the process-wide
embedQuery()cache incore/embeddings.tskeying by query text alone, which let twoMnemopiinstances in the same process with different providers/models cross-contaminate theirdense_scorerankings. The cache key now includes a WeakMap-assigned provider identity, the resolved model name, and the configuredapiUrl, so disjoint runtimes never read each other's cached vectors. (#1832, follow-up to #1833 review)
[15.7.4] - 2026-05-31
Fixed
- Fixed the
darwin-x64release build failing inbun build --compilebecause the Windows ORT 1.24 preload pulledonnxruntime-nodeinto the static graph and there is nodarwin/x64prebuilt for that line. The preload is now guarded behind aprocess.platform === "win32"literal that Bun dead-code-eliminates on non-Windows targets; macOS/Linux load fastembed's bundled ORT 1.21 binding as before.
[15.7.3] - 2026-05-31
Changed
- Changed embedding result normalization to return
Float32Arrayvectors soembedandembedQuerynow cache and emit float32 rows - Changed the embedding provider contract to a single typed
EmbeddingOutput(AsyncIterable<number[][]>) instead ofunknown, matching fastembed'sembed(), soEmbeddingProvider.embedand theproviderruntime option stream the embedding matrix as async batches (async *embed(texts) { yield texts.map(embedOne); }) - Changed local model cache directory resolution for
fastembedto usegetFastembedCacheDirinstead of the hard-coded~/.hermes/cache/fastembedpath
Fixed
- Fixed cosine similarity behavior across retrieval, clustering, and caching to consistently handle mismatched vector lengths as zero-padded and ignore non-finite values
- Fixed embedding API requests to retry transient failures with backoff via shared retry logic before returning null
- Fixed compiled
ompbinaries losing local Mnemopi embeddings by keepingfastembedandonnxruntime-nodereachable to Bun's static compiler while preserving lazy runtime loading.
[15.7.2] - 2026-05-31
Fixed
- Fixed Windows startup crashes by keeping fastembed's older ONNX Runtime binding lazy until local embeddings are used.
- Fixed a segfault at startup from eagerly loading fastembed: importing the embeddings module pulled in
fastembed, which eagerly loads theonnxruntime-nodenative addon. The import is now deferred until a local fastembed model is actually initialized, so API-model, disabled-embeddings, and test runtimes never load the native addon.
[15.6.0] - 2026-05-30
Added
- Added
llm.extractionPromptruntime option to override the fact-extraction prompt template using{text}and{lang}placeholders - Added
llm.consolidationPromptruntime option to override the consolidation sleep prompt template using{memories},{source}, and{memory_count}placeholders - Published
@oh-my-pi/pi-mnemopito npm: the local SQLite memory engine is now built, checked, tested, and released through the monorepo CI pipeline alongside the other workspace packages. - Exported the diagnostic inspector as the
@oh-my-pi/pi-mnemopi/diagnosesubpath for coding-agent memory maintenance commands. - Added
flushExtractions()(onMnemopi,BeamMemory, and as a module-level export) to drain in-flight background fact extraction; used by tests and graceful shutdown so facts are persisted before the database closes.
Changed
- Changed fact extraction to prefer a configured runtime LLM completion path before host extraction, with automatic fallback when the configured completion returns no output or fails
Fixed
- Fixed
rememberBatch(..., { extract: true })to run background fact extraction for batch uploads (including per-itemextractflags) so extracted facts are generated and recallable after extraction - Fixed
extract: truefact extraction to continue safely when no LLM is configured by turning extraction failures into no-op background tasks - Fixed configured LLM fact extraction by using temperature 0 so re-ingesting the same text is deterministic and avoids near-duplicate extractions
- Fixed
remember(..., { extract: true })silently dropping the flag: it now schedules the LLM fact extractor (extractFactsSafe) over the stored content and persists the extracted facts so they become recallable. Previously the LLM extractor had no production callers andextractwas dead.