The memory-extraction prompt concatenated its instructions, few-shot
examples, and the user message into a single user turn, so a small local
model could not distinguish instructions from input and frequently echoed
the Globex/weather examples instead of extracting facts.
Send the instructions as a real system turn and the raw text as the user
turn. The tiny worker protocol gains a systemPrompt field, and Mnemopi
completion input carries task metadata so the backend selects the right
prompt per call.
Drop the code-built MEMORY_EXTRACTION_TEMPLATE rather than porting it:
prompt text belongs in .md files, and resolveMemoryCompletionInput already
overrides that template for every extraction call, so Mnemopi rendered it
only for the result to be discarded.
Measured on ONNX q4 CPU, LFM2.5-1.2B memory extraction improved from 1/8
to 5/8 once the roles were separated.
buildWhere() appended a redundant hard `channel_id = ?` clause on top of
the `(session_id = ? OR scope = 'global' OR channel_id = ?)` visibility
clause. The hard AND nullified the `scope='global'` branch, so any global
row whose channel_id differed from the recall channelId — including rows
imported via importFromDict() with channel_id NULL — was silently dropped.
Channel isolation is fully preserved by the visibility OR-clause alone
(other-channel, non-global, cross-session rows still fail all three
disjuncts), so removing the hard clause restores global visibility without
leaking other channels.
Fixes#8525
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Updated user agent and openrouter titles in packages/ai from Oh-My-Pi to omp.
- Integrated getOpenRouterHeaders into image generation tool requests in packages/coding-agent.
- Updated provider documentation and test assertions to reflect the rebrand.
An interrupted fastembed model download leaves <cacheDir>/<model>/ with
sidecars and a truncated model.onnx_data but no model.onnx. Upstream
retrieveModel short-circuits on the existing dir (and reuses a leftover
partial <model>.tar.gz), so FlagEmbedding.init throws "Model file not
found at .../model.onnx" every session: semantic recall silently dies
machine-wide and reconcileEmbeddingModel re-enqueues the same
never-embeddable rows on every store open.
quarantineCorruptModelFile only matched "Protobuf parsing failed", never
this partial-extraction variant. Add clearIncompleteModelCache: a
"Model file not found" init failure now removes the incomplete model dir
and the leftover partial archive (containment-guarded to a direct child
of the fastembed cache root) and retries init exactly once, so the next
attempt re-downloads cleanly and recall self-heals.
Fixes#7916
- mnemopi provider parity 'diagnose, validate, graph' does ~6.7s of real
work under bun --parallel=8 on loaded runners; raised its per-test
timeout to 30s (default 5s flaked twice in three CI runs).
- utils LRUCache updateAgeOnGet drove a 30ms TTL with real 20ms sleeps
(10ms margin); now drives performance.now() via a mocked clock, so the
contract is asserted deterministically with no wall-clock wait.
- Restricted the reasoning-wrapper removal to leading <think> blocks so a
literal think tag after real content is preserved instead of deleted.
- Added coverage asserting mid-content think tags survive cleanOutput.
Fixes#7231
- Removed MiniMax-style think wrappers in cleanOutput so reasoning-model
responses no longer leak into consolidation summaries or corrupt fact
extraction (the wrapper previously survived parsing and every stored fact
became reasoning prose).
- Added remote consolidation and remote extraction regression coverage.
Fixes#7231
Six suites build a real DB under a temp dir and then `rmSync` it in
`afterEach`. Each leaked its own one-shot statement the same way the source
did, so the teardown hit EBUSY on Windows and failed tests whose assertions
had already passed.
Only the suites that touch a file-backed DB are changed. The remaining
`prepare()` calls in the package tests run against `:memory:`, where there
is nothing to unlink and no cleanup to fail.
Every `db.prepare(...)` in the package was a one-shot: prepared, stepped
once via run/get/all, then dropped without `finalize()`. There were 65 such
sites and no `finalize()` call anywhere. An unfinalized statement keeps the
SQLite connection alive, so `BeamMemory.close()` -> `closeQuietly(db)` ->
`db.close()` never released the file. `closeQuietly` swallows the error;
calling `db.close(true)` instead surfaces "database is locked".
One-shot writes now go through `db.run(sql, params)`, which prepares and
finalizes internally. Statements that are read from, or reused across a
loop, are bound with `using` so they are finalized when they leave scope.
On Windows the effect is user-visible beyond tests: the memory DB and its
`-wal`/`-shm` sidecars stay locked after mnemopi closes them, so the file
cannot be deleted, moved, or rotated. On POSIX the leak is silent because
open files can be unlinked.
The new suite pins both halves: that `db.close(true)` does not throw after
the store paths run, and that a closed bank leaves no `-wal`/`-shm` behind
and can be deleted. All three cases fail on the unfixed sources.
Same class as 14252e71c and #6762, neither of which covered this package.
Swap the batch-shaped scoring loops to the pi-natives kernels behind the
existing TS function signatures, so every caller keeps its guards and
observable behavior: mmrRerank (default Jaccard path), searchExactVectorIndex,
clusterBySimilarity, BinaryVectorStore.search, and FastBinarySearch.search.
cosine_similarity_pairs now takes Float64Array so f64 vectors round-trip
without f32 narrowing. Scalar one-off call sites stay in TS (query-cache
cosine probe, shmr centroid/confidence loops, recall per-row cosine map,
custom-similarity mmrRerank). Adds a seeded parity suite: 1e-9 relative
tolerance for floats, exact equality for Hamming distances and pair lists,
identical MMR index sequences across lambda/topK grids.
defaultLocalModelInitializer (exported @internal for tests; production
seam stays setLocalModelInitializer) now provably retries
FlagEmbedding.init EXACTLY once after a Protobuf-corruption failure,
quarantining the cached file in between, and surfaces the error without
looping when the retry fails too. The heal callsite now passes
options.cacheDir so containment checks the caller's cache root, not
only the global default.
A truncated model_optimized.onnx (observed live: 5.7MB file from June,
'Protobuf parsing failed' on every load) blocks local embeddings
forever: the downloader treats the existing file as complete, so every
init re-parses the same broken bytes and local recall/retain loses its
embedder. defaultLocalModelInitializer now quarantines the exact file
named by the loader (atomic rename to *.corrupt-<ts>) and retries init
ONCE so the model re-downloads.
The extracted path is error-message content: it is honored only when
it resolves inside the fastembed cache directory, so a dependency
emitting an unexpected message can never rename an arbitrary file.
Contract tests: protobuf failure quarantines exactly the named cache
file and reports retry-safe; unrelated errors touch nothing; a missing
file (concurrent heal) stays retry-safe; a path outside the cache root
is refused untouched.
Working-memory TTL trim treated every consolidated_at IS NULL row as
scratch, so restored or imported durable rows disappeared on the next
write, and the trim delete left annotations, embeddings, facts, and
memoria projections orphaned.
- Exclude IMPORTED-tier rows from the trim eligibility query so restored
banks survive a later remember()/rememberBatch().
- Stamp imported working-memory rows as consolidated in importFromDict so
restored backups are durable regardless of trust tier.
- Add purgeWorkingMemoryArtifacts() and route trim, forgetWorking, and
force-import overwrite through it to cascade annotations, embeddings,
facts (source_msg_id), memoria_* (source_memory_id), gists, and the
graph edges tied to those memory/gist/fact node ids.
Fixes#4819
- Resolved ONNX Runtime from fastembed's own dependency graph.
- Prepended the cached DLL directory before loading the native binding.
- Added regression coverage for inherited stale runtime paths.
Fixes#4849
recall (includeFacts) surfaces facts.fact_id as a result id, but
store.get only searched working_memory + episodic_memory, so every
surfaced fact id was a dead end for 'read memory://<id>' and
memory_edit ('not found in any scoped bank').
- store.get now falls back to the facts table (visibility mirrors
factRecall: same-session or scope='global'), returning a read-only
row with memory_store 'fact' and the full triple as content.
- coding-agent labels the store honestly ('fact') in memory:// reads
and reports not_editable (instead of not_found) for memory_edit ops
on fact ids; the facts table stays immutable.
Fixes#4725
- Accepted natural wrapper fields for instruction and preference objects.
- Preserved timeline descriptions with their dates when object-shaped timeline items are parsed.
- Extended the parseFacts regression to cover category-specific fields.
Fixes#4649
Recall silently clipped every result to 500 chars mid-word with no marker,
and memory_edit update replaces content wholesale by id. There was no way
for an agent to inspect the full row before overwriting it: Mnemopi.get
existed but nothing surfaced it, and the advertised URI only served the
file-backed memory summary. The natural recall/inspect/update loop had no
inspect step.
Three fixes across mnemopi and coding-agent:
- recall now appends an ellipsis marker when it clips content and reports
truncated=true plus full_length. The cap is exposed as
RecallOptions.contentPreviewChars (default 500, 0 disables). The factLine
200-char clip used by the enhanced-context sandwich gets the same marker.
- Under the mnemopi backend the read-tool URI scheme now routes an id host
to Mnemopi.get() across every session's scoped banks, returning the row
as text/markdown with a YAML-frontmatter header (bank, store, source,
timestamp, importance, veracity). The root namespace remains for the
file-backed summary. Miss errors now name the backend explicitly.
- Updated the recall and memory_edit tool prompts to document the
truncation marker and require reading the full memory before any
wholesale content update.
Fixes#4443
Scaled enhanced fact recall with the requested topK, bucketed formatted context without duplicate rows, and ignored flat fact storage fields during scoring.
Added regression coverage for extracted flat fact recall and storage-only fact/entity noise.
Fixes#4402
Carried working_memory.embed_text into sleep consolidation summaries so episodic content, FTS, and embeddings do not reintroduce retention protocol markers after working rows age out.
Fixes#4395
Scored working-memory recall candidates against embed_text when present so FTS matches from the projection survive the lexical gate even without dense embeddings.
Fixes#4395
Added an embedText projection for remember() so stored transcripts can remain readable while embeddings, FTS indexing, and embedding-model rebuilds use marker-free text. Updated coding-agent retention to pass the marker-free projection and strip retained protocol markers from recall display.
Fixes#4395
Short-circuit successfully parsed structured extractor output even when all extraction arrays are empty, so the JSON response body is not stored as a fallback memory.
Added coverage for plain and fenced empty extraction JSON.
Fixes#4390
- Added an extractText override to pi-mnemopi remember paths so stored content and mined facts can use different text.
- Routed coding-agent mnemopi retention to store the full transcript while extracting only user-authored turns.
- Tightened deterministic Instruction extraction to require an explicit I/you subject.
Fixes#3372
`MnemopiSessionState.retainMessages` hands `embed()` the chronological
multi-turn transcript (oldest -> newest). A naive `slice(0, max)` cap landed
on the oldest turns and dropped the most recent content, so every retained
episode past the cap collapsed onto essentially the same prefix vector and
dense recall could not match topics introduced after the first 8192 chars.
`capInputs` now routes oversized inputs through `clipToWindow`, which keeps
roughly half the cap from the head and half from the tail with a small
`[...]` elision marker between them. Short inputs and array reference pass-
through are unchanged. Falls back to a tail-only clip when `max` is too small
to fit a useful split. Locked in by a new `embedding-input-cap.test.ts` case
that pins markers at both ends of a 50k transcript and asserts both survive
the clip.
Fixes#3126
Defaulted MNEMOPI_EMBEDDING_MAX_INPUT_CHARS to 8192 so the automatic guard matches bge-m3 and OpenAI text-embedding context limits by default. Larger local embedding servers such as Qwen3-Embedding with 32k ctx can still raise the cap, and 0 still disables truncation.
Fixes#3126
`MnemopiSessionState.retainMessages` (`packages/coding-agent/src/mnemopi/state.ts:352`)
always calls `prepareRetentionTranscript(messages, true)` and hands the whole
multi-turn transcript to `embed([transcript])`. Long sessions (especially CJK
content) routinely outgrow the embedding model's context window, and llama.cpp's
`/embeddings` server rejects oversized requests with
`request (N tokens) exceeds the available context size` — every retain after
that point silently lost its vector row, leaving recall on FTS-only.
Capped per-input length inside `embed()` (the single chokepoint every retain /
query / consolidate flow funnels through) at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`
(default 32000 chars ≈ 8k English tokens / 16–32k CJK tokens, override via env
or `embeddings.maxInputChars` runtime option; `0` disables). The new array
is allocated only when at least one input is oversized, so the typical short-
query path through `embedQuery` still passes the original array through; the
truncation also emits a debug-or-warn log so the resize is no longer silent.
Fixes#3126
- Replaced usage of `ReturnType<typeof setTimeout>` and `ReturnType<typeof setInterval>` with the explicit `Timer` type across the codebase.
- Updated several type definitions and function signatures to use concrete types instead of inferred return types for improved clarity and maintainability.
- Downloaded config.json with tokenizer sidecars when repairing stale fastembed model caches.\n- Treated fastembed Config file missing errors as repairable initializer failures.\n- Extended cache repair coverage for the missing-config state.\n\nFixes #3054
- Preserved fastembed's transitive ONNX Runtime instead of forcing an ABI-mismatched runtime cache install.\n- Repaired stale fastembed model caches by downloading missing tokenizer sidecars from the matching Hugging Face model repo.\n- Added regression coverage for the runtime install plan and tokenizer sidecar repair.\n\nFixes #3054