feat(coding-agent): added local model registry for title completion

- Added Mnemosyne runtime `extractionPrompt` and `consolidationPrompt` options and wired them into resolved LLM config.
- Added fact-extraction branch to call configured completion first (temp 0), then parse facts and safely fall back.
- Added tiny local model registry features for memory/title, including keys, specs, and validation helpers.
- Added `complete` protocol messages and abort-aware client/worker paths for local title completion generation.
- Added local-models documentation for tiny/memory transformer paths, defaults, and known parser caveats.
This commit is contained in:
can1357
2026-05-30 18:41:48 +02:00
parent 91513cdbf3
commit 70202360ff
11 changed files with 415 additions and 36 deletions
+132
View File
@@ -0,0 +1,132 @@
# Embedded Local Tiny-Model Experiments
This document summarizes the experiments behind the optional **local** tiny-model paths for two
coding-agent tasks: session-title generation (`providers.tinyModel`) and Mnemosyne memory
extraction/consolidation (`providers.memoryModel`). It is a factual engineering record for
maintainers: what we measured, which recipes won, and which models we shipped. Both settings
default to `online`, so existing users incur no downloads or CPU cost unless they opt in.
## Runtime / environment findings
- **Stack**: `@huggingface/transformers` (transformers.js) v4 running under Bun. In Bun the library
loads the **native `onnxruntime-node` backend** (not the WASM build). Available device options are
`cpu` / `coreml` / `webgpu` — there is **no `wasm` device** in the node build.
- **Device verdict: use `device:"cpu"`.** CPU is the only reliable path.
- `coreml` EP is **broken for decoder LLMs**: it rejects the dynamic KV-cache `past_key_values`
zero-element first-token shape.
- `webgpu` runs but is **slower and numerically divergent** (worse output quality).
- **Quantization: q4 is the sweet spot** — smaller on disk, faster to load, and fast at inference.
q8/int8 loads slower *and* infers slower on CPU.
- **Load-time correction (important).** An earlier belief that "q4 >=1B models take minutes to load"
was a **measurement artifact** caused by running ~5 multi-GB HuggingFace downloads in parallel
(I/O saturation). Clean, isolated **warm** loads are all sub-3s:
- TinyLlama-1.1B q4: ~0.5s
- Llama-3.2-1B q4: ~2.8s (`graphOpt=all`) / ~0.5s (`disabled`)
- LFM2-1.2B q4: ~0.36s
- Qwen2.5-1.5B q4: ~1.5s
- Qwen3-1.7B q4: ~1.6s
- gemma-3-1b q4: ~1.1s
- Conclusion: **1B–1.7B models are viable on CPU.**
- **`session_options.graphOptimizationLevel`** trades load vs inference speed: `disabled` = fastest
load, slightly slower inference; `all` = default.
- **First run** downloads weights from the HF Hub to a cache dir (q4 weights ~200MB–1.1GB depending
on model); subsequent **warm** loads are sub-second to ~3s. Inference is async and
background-friendly for memory tasks; titles are semi-interactive.
## Task 1: Session title generation (`providers.tinyModel`)
**Task**: turn the first user message into a 3–6 word title. Tiny models (sub-1B) suffice.
**Winning recipe**:
- Plain system prompt (no few-shot).
- **Prefill** the assistant turn with `<title>` and **stop at `</title>`**, then take the first line.
- Greedy decoding (`do_sample:false`), `enable_thinking:false` in the chat template.
**What we learned**:
- **Few-shot examples HURT sub-0.6B models** for titles; the tag-prefill rescues even 270M models.
- **Token biasing (`bad_words_ids`) is a confirmed no-op** here — the prefill already controls the
opener.
**Leaderboard** (tag trick, CPU, warm):
| Model | Verdict |
| --- | --- |
| LFM2-350M | Best speed/quality balance (~212MB) |
| Qwen3-0.6B | Most robust |
| gemma-3-270m | Smallest viable |
| Qwen2.5-0.5B | Acceptable |
| SmolLM2-135M | Too small |
| flan-t5-small | Rejected — just echoes the input |
**Shipped local options**: `lfm2-350m`, `qwen3-0.6b`, `gemma-270m`, `qwen2.5-0.5b`, `lfm2-700m`.
**Default**: `online` (pi/smol).
## Task 2: Mnemosyne memory (`providers.memoryModel`)
Mnemosyne runs two small-LLM tasks:
1. **Extraction** — pull durable, structured items from a single message.
2. **Consolidation** — summarize a list of memories into 1–3 faithful sentences.
These need **bigger models than titles: 1B–1.7B**. We tested LFM2-1.2B, Qwen2.5-1.5B, Qwen3-1.7B,
and gemma-3-1b (q4, CPU) via four parallel agents each running 27–31 experiments.
### Extraction findings
The stock 5-category JSON prompt fails on small models in two ways:
1. The all-empty example `{"facts":[],...}` gets **copied verbatim** → 0 facts extracted.
2. Capable models emit **JSON objects inside arrays**, which Mnemosyne's `String(item)` coerces into
the literal string `[object Object]`.
The robust fix is a **one-item-per-line output format** (consumed by Mnemosyne's parser line-fallback)
or a **flat JSON array of strings**. Every model also over-extracts pure small talk; an explicit
chit-chat → NONE example is the best mitigation.
### Technique polarity flips vs titles
- At 1B+, **few-shot is the dominant quality lever**: e.g. Qwen2.5-1.5B extraction F1 0.52 → 0.83
going 1 → 3 shots; gemma recall 0.65 → 0.92 with 2 shots.
- **Prefill HURTS extraction** — it forces output on small talk, producing false positives.
- **System-split** (instructions in the system role) helps models that have a system role.
- **Greedy >= temperature** for both tasks.
- **Token biasing** is again a no-op.
### Per-model verdicts (head-to-head, 16-fixture set)
- **Qwen3-1.7B** — most disciplined extraction: returns empty on small talk, no buried-fact leak,
preserves language, clean flat JSON. Weaknesses: coarse granularity, missed a multi-turn value
update.
- **Qwen2.5-1.5B** — best extraction granularity (atomic facts), caught the value update, zero
small-talk leakage. Weaknesses: weakest consolidation (run-on, no dedup) and one degenerate
buried-fact output.
- **gemma-3-1b** — best consolidation (dedup works, faithful, clean single-memory). Weaknesses: leaks
small talk and translated German.
- **LFM2-1.2B** — solid and fastest to load. Weaknesses: `Label: value` noise, small-talk + buried
leaks, a fluffy single-memory summary.
### Recommendation
Extraction favors **precision** (do not pollute long-term memory) → **Qwen3-1.7B is the best single
pick** (its consolidation is good enough). If running a second model for consolidation, **gemma-3-1b**
wins that task.
**Shipped local options**: `qwen3-1.7b` (recommended), `gemma-3-1b`, `qwen2.5-1.5b`, `lfm2-1.2b`.
**Default**: `online` (the configured smol model).
### Known Mnemosyne parser bugs (surfaced by these experiments)
- `String(item)` produces `[object Object]` on object array items.
- The line-fallback drops items `<=10` chars, so a correct short fact like `Name: Can` is discarded.
## Integration notes
- Both settings default to `online`, so existing users get **no downloads or CPU cost** unless they
opt in.
- Local inference runs **in a worker** (off the main thread); models are cached on disk and
downloaded on first use.
- The memory local path applies the refined recipes (line-format + small-talk-guarded extraction
prompt, hardened consolidation prompt) via Mnemosyne prompt overrides; the **online path is
unchanged**.
+4 -9
View File
@@ -1,9 +1,10 @@
# Changelog
## [Unreleased]
### Added
- Added Mnemosyne memory inference model selection with an online mode or local transformers.js options (`qwen3-1.7b`, `gemma-3-1b`, `qwen2.5-1.5b`, `lfm2-1.2b`) so memory extraction and consolidation can run via the shared tiny-model worker
- Changed memory tiny-model handling to route local memory prompts through the same queueed tiny-model worker pipeline with bounded completion output
- Added a Providers → Tiny Model setting for session titles, defaulting to the online `pi/smol` path with five optional local CPU transformers.js models. A local model — and the one-time `@huggingface/transformers` runtime install in compiled binaries — is downloaded and loaded only when explicitly selected (or via `omp tiny-models download`); the default online path never spawns the title worker for inference. Selecting a local model adds a delayed `pi/smol` fallback so titles never block, plus in-chat download progress.
- Added a persistent live agent roster pinned below the editor (focus it with `Ctrl+S` or `Alt+Down`), including view-as switching into delegated agent sessions with human-readable delegate names and UI pinning to suppress idle reaping while viewed. The roster stays hidden until at least one delegated agent exists and releases focus back to the editor once the last one is gone.
- Recorded the originating session ID alongside each prompt in `history.db` (new `session_id` column, surfaced as `HistoryEntry.sessionId`), so recalled prompts can be traced back to the session they came from. Existing history databases gain the column automatically on next launch.
@@ -21,11 +22,6 @@
- Changed Mnemosyne `recall` tool output to include memory ids for explicit recall results so agents can target `memory_edit`; auto-injected memory context and `reflect` remain id-free.
- Changed the system prompt to advertise `memory://root` only when the local memory backend is active.
### Fixed
- Fixed a native crash (`malloc: pointer being freed was not allocated` / `NAPI FATAL ERROR`) when quitting after the local transformers.js title model had run. The tiny-title worker no longer calls `pipeline.dispose()` on shutdown — disposing the onnxruntime session freed native memory that Bun's worker/NAPI teardown then freed again. The worker is torn down immediately after, so the OS reclaims the model memory regardless.
- Fixed the tiny-title download progress bar flashing on every first message even when the local model was already downloaded. A cached model emits the same `download`/`progress` events as a real download, so the bar is now revealed only when in-flight progress events keep arriving past a short grace window — cache hits finish (or fall silent during onnxruntime init) before then and never show the bar.
### Removed
- Removed the standalone `ask`, `task`, and `yield` tools along with their obsolete prompts, docs, and tests; delegation now routes through persistent `delegate` agents plus IRC coordination.
@@ -33,6 +29,8 @@
### Fixed
- Fixed a native crash (`malloc: pointer being freed was not allocated` / `NAPI FATAL ERROR`) when quitting after the local transformers.js title model had run. The tiny-title worker no longer calls `pipeline.dispose()` on shutdown — disposing the onnxruntime session freed native memory that Bun's worker/NAPI teardown then freed again. The worker is torn down immediately after, so the OS reclaims the model memory regardless.
- Fixed the tiny-title download progress bar flashing on every first message even when the local model was already downloaded. A cached model emits the same `download`/`progress` events as a real download, so the bar is now revealed only when in-flight progress events keep arriving past a short grace window — cache hits finish (or fall silent during onnxruntime init) before then and never show the bar.
- Fixed the Mnemosyne memory backend lifecycle so auto-retain counts the full session transcript, delegated agents inherit the parent Mnemosyne state, `/memory clear` removes scoped project-bank databases, session disposal closes Mnemosyne SQLite handles, session switches rekey/reset Mnemosyne tracking, and project bank names include an absolute-root hash with safe bank-name sanitization.
- Fixed the streaming edit preview showing no diff for single-line hashline edits. The preview-diff coalescing keyed only on the arg text, so the final (args-complete) pass — which computes an untrimmed diff — was skipped because the payload was byte-identical to the last streamed chunk whose trailing line had been trimmed. The dedup key now pairs the streaming state with a content hash.
- Fixed `Esc` in a delegated agent view returning to the main session instead of aborting the delegated agent's active turn.
@@ -40,9 +38,6 @@
- Fixed the agent roster staying pinned under the editor when all delegated agents are idle or dormant; it now reappears when explicitly focused with `Alt+Down` / session observe.
- Fixed selector-style UI components to honor `tui.select.up` and `tui.select.down` keybindings instead of hard-coding raw Up/Down arrow bytes ([#1535](https://github.com/can1357/oh-my-pi/issues/1535)).
- Fixed the bash (and `recipe`) tool result footer not rendering for failed commands. A non-zero exit threw a `ToolError`, which dropped the result details, so the styled `⟨Wall … | Timeout …⟩` footer was replaced by the raw `Wall time: … seconds` / `Command exited with code N` lines. Non-zero exits now resolve as a non-throwing error result that keeps `wallTimeMs`/`timeoutSeconds`/`exitCode`, and the footer shows `⟨Wall … | Timeout … | Exit: N⟩` with the textual notices folded out of the output pane. Aborts, timeouts, and missing-exit-status still throw as before.
### Fixed
- Fixed selector-style UI components to honor `tui.select.up` and `tui.select.down` keybindings instead of hard-coding raw Up/Down arrow bytes ([#1535](https://github.com/can1357/oh-my-pi/issues/1535)).
## [15.5.15] - 2026-05-30
+108
View File
@@ -101,3 +101,111 @@ export function getTinyTitleModelSpec(key: TinyTitleLocalModelKey): (typeof TINY
if (!spec) throw new Error(`Unknown tiny title model: ${key}`);
return spec;
}
/** Default memory model: the online path (the configured smol / remote LLM; no local download). */
export const ONLINE_MEMORY_MODEL_KEY = "online";
/** Recommended local model for memory tasks when none is named. */
export const DEFAULT_MEMORY_LOCAL_MODEL_KEY = "qwen3-1.7b";
/**
* Local models for Mnemosyne memory tasks (fact extraction + consolidation).
* These are larger (1B-1.7B) than the title models: structured extraction and
* faithful summarization need more capacity than 3-6 word titles. All q4, CPU.
* Ranking/recipe rationale lives in docs/local-models.md.
*/
export const TINY_MEMORY_LOCAL_MODELS = [
{
key: "qwen3-1.7b",
repo: "onnx-community/Qwen3-1.7B-ONNX",
dtype: "q4",
label: "Qwen3 1.7B",
description:
"Recommended; most disciplined extraction (ignores chit-chat), good consolidation, about 1.1 GB cached.",
contextNote: "Best single-model pick for memory from the CPU experiment.",
},
{
key: "gemma-3-1b",
repo: "onnx-community/gemma-3-1b-it-ONNX",
dtype: "q4",
label: "Gemma 3 1B",
description: "Best consolidation/dedup; lighter footprint, but leaks small talk during extraction.",
contextNote: "Use when consolidation quality and size matter most.",
},
{
key: "qwen2.5-1.5b",
repo: "onnx-community/Qwen2.5-1.5B-Instruct",
dtype: "q4",
label: "Qwen2.5 1.5B",
description: "Best extraction granularity (atomic facts); weaker consolidation.",
contextNote: "Use when fine-grained, deduplicatable facts matter more than summaries.",
},
{
key: "lfm2-1.2b",
repo: "onnx-community/LFM2-1.2B-ONNX",
dtype: "q4",
label: "LFM2 1.2B",
description: "Fastest load; solid all-rounder, slightly noisier extraction labels.",
contextNote: "Use when local startup cost is the priority.",
},
] as const satisfies readonly TinyTitleLocalModelSpec[];
export const TINY_MEMORY_MODEL_VALUES = [
ONLINE_MEMORY_MODEL_KEY,
"qwen3-1.7b",
"gemma-3-1b",
"qwen2.5-1.5b",
"lfm2-1.2b",
] as const;
export type TinyMemoryModelKey = (typeof TINY_MEMORY_MODEL_VALUES)[number];
export type TinyMemoryLocalModelKey = (typeof TINY_MEMORY_LOCAL_MODELS)[number]["key"];
type MissingTinyMemoryModelValue = Exclude<
typeof ONLINE_MEMORY_MODEL_KEY | TinyMemoryLocalModelKey,
TinyMemoryModelKey
>;
type ExtraTinyMemoryModelValue = Exclude<TinyMemoryModelKey, typeof ONLINE_MEMORY_MODEL_KEY | TinyMemoryLocalModelKey>;
const TINY_MEMORY_MODEL_VALUES_MATCH_REGISTRY: MissingTinyMemoryModelValue extends never
? ExtraTinyMemoryModelValue extends never
? true
: never
: never = true;
void TINY_MEMORY_MODEL_VALUES_MATCH_REGISTRY;
export const TINY_MEMORY_MODEL_OPTIONS = [
{
value: ONLINE_MEMORY_MODEL_KEY,
label: "Online (smol/remote)",
description: "Use the configured Mnemosyne LLM mode (smol or remote); no local model download or CPU inference.",
},
...TINY_MEMORY_LOCAL_MODELS.map(model => ({
value: model.key,
label: model.label,
description: model.description,
})),
] satisfies ReadonlyArray<{ value: TinyMemoryModelKey; label: string; description: string }>;
export function isTinyMemoryLocalModelKey(value: string): value is TinyMemoryLocalModelKey {
return TINY_MEMORY_LOCAL_MODELS.some(model => model.key === value);
}
export function getTinyMemoryModelSpec(key: TinyMemoryLocalModelKey): (typeof TINY_MEMORY_LOCAL_MODELS)[number] {
const spec = TINY_MEMORY_LOCAL_MODELS.find(model => model.key === key);
if (!spec) throw new Error(`Unknown tiny memory model: ${key}`);
return spec;
}
/** Any local model key (title or memory), used by the shared inference worker. */
export type TinyLocalModelKey = TinyTitleLocalModelKey | TinyMemoryLocalModelKey;
/** Resolve a local model spec by key across both the title and memory registries. */
export function getTinyLocalModelSpec(key: string): TinyTitleLocalModelSpec | undefined {
return (
TINY_TITLE_LOCAL_MODELS.find(model => model.key === key) ??
TINY_MEMORY_LOCAL_MODELS.find(model => model.key === key)
);
}
export function isTinyLocalModelKey(value: string): value is TinyLocalModelKey {
return getTinyLocalModelSpec(value) !== undefined;
}
+57 -6
View File
@@ -1,5 +1,12 @@
import { isCompiledBinary, logger } from "@oh-my-pi/pi-utils";
import { isTinyTitleLocalModelKey, type TinyTitleLocalModelKey } from "./models";
import {
isTinyLocalModelKey,
isTinyMemoryLocalModelKey,
isTinyTitleLocalModelKey,
type TinyLocalModelKey,
type TinyMemoryLocalModelKey,
type TinyTitleLocalModelKey,
} from "./models";
import type { TinyTitleProgressEvent, TinyTitleWorkerInbound, TinyTitleWorkerOutbound } from "./title-protocol";
interface WorkerHandle {
@@ -11,7 +18,8 @@ interface WorkerHandle {
type PendingRequest =
| { kind: "generate"; modelKey: TinyTitleLocalModelKey; resolve: (title: string | null) => void }
| { kind: "download"; modelKey: TinyTitleLocalModelKey; resolve: (ok: boolean) => void };
| { kind: "complete"; modelKey: TinyMemoryLocalModelKey; resolve: (text: string | null) => void }
| { kind: "download"; modelKey: TinyLocalModelKey; resolve: (ok: boolean) => void };
export interface TinyTitleDownloadOptions {
signal?: AbortSignal;
@@ -145,8 +153,44 @@ export class TinyTitleClient {
}
}
async complete(
modelKey: string,
prompt: string,
options: { maxTokens?: number; signal?: AbortSignal } = {},
): Promise<string | null> {
if (!isTinyMemoryLocalModelKey(modelKey)) return null;
if (options.signal?.aborted) return null;
try {
const worker = this.#ensureWorker();
const id = String(++this.#nextRequestId);
const { promise, resolve } = Promise.withResolvers<string | null>();
this.#pending.set(id, { kind: "complete", modelKey, resolve });
const abort = (): void => {
const pending = this.#pending.get(id);
if (pending?.kind !== "complete") return;
this.#pending.delete(id);
pending.resolve(null);
};
options.signal?.addEventListener("abort", abort, { once: true });
try {
worker.send({ type: "complete", id, modelKey, prompt, maxTokens: options.maxTokens });
return await promise;
} finally {
options.signal?.removeEventListener("abort", abort);
this.#pending.delete(id);
}
} catch (error) {
logger.debug("tiny-model: local completion failed", {
modelKey,
error: error instanceof Error ? error.message : String(error),
});
return null;
}
}
async downloadModel(modelKey: string, options: TinyTitleDownloadOptions = {}): Promise<boolean> {
if (!isTinyTitleLocalModelKey(modelKey)) return false;
if (!isTinyLocalModelKey(modelKey)) return false;
if (options.signal?.aborted) return false;
const unsubscribe = options.onProgress ? this.onProgress(options.onProgress) : undefined;
@@ -189,7 +233,7 @@ export class TinyTitleClient {
this.#unsubscribeError = null;
for (const pending of this.#pending.values()) {
this.#emitProgress({ modelKey: pending.modelKey, status: "error" });
if (pending.kind === "generate") pending.resolve(null);
if (pending.kind === "generate" || pending.kind === "complete") pending.resolve(null);
else pending.resolve(false);
}
this.#pending.clear();
@@ -232,9 +276,13 @@ export class TinyTitleClient {
if (pending.kind === "download") pending.resolve(true);
return;
}
if (message.type === "completion") {
if (pending.kind === "complete") pending.resolve(message.text);
return;
}
logger.debug("tiny-title: worker returned error", { error: message.error });
this.#emitProgress({ modelKey: pending.modelKey, status: "error" });
if (pending.kind === "generate") pending.resolve(null);
if (pending.kind === "generate" || pending.kind === "complete") pending.resolve(null);
else pending.resolve(false);
}
@@ -246,7 +294,7 @@ export class TinyTitleClient {
logger.warn("tiny-title: worker error", { error: error.message });
for (const pending of this.#pending.values()) {
this.#emitProgress({ modelKey: pending.modelKey, status: "error" });
if (pending.kind === "generate") pending.resolve(null);
if (pending.kind === "generate" || pending.kind === "complete") pending.resolve(null);
else pending.resolve(false);
}
this.#pending.clear();
@@ -256,6 +304,9 @@ export class TinyTitleClient {
export const tinyTitleClient = new TinyTitleClient();
/** Alias for the shared tiny-model worker client (titles + memory completions). */
export const tinyModelClient = tinyTitleClient;
export async function shutdownTinyTitleClient(): Promise<void> {
await tinyTitleClient.terminate();
}
@@ -1,4 +1,4 @@
import type { TinyTitleLocalModelKey } from "./models";
import type { TinyLocalModelKey, TinyTitleLocalModelKey } from "./models";
export type TinyTitleProgressStatus =
| "initiate"
@@ -15,7 +15,7 @@ export interface TinyTitleProgressFileState {
}
export interface TinyTitleProgressEvent {
modelKey: TinyTitleLocalModelKey;
modelKey: TinyLocalModelKey;
status: TinyTitleProgressStatus;
name?: string;
file?: string;
@@ -30,12 +30,14 @@ export interface TinyTitleProgressEvent {
export type TinyTitleWorkerInbound =
| { type: "ping"; id: string }
| { type: "generate"; id: string; modelKey: TinyTitleLocalModelKey; message: string }
| { type: "download"; id: string; modelKey: TinyTitleLocalModelKey }
| { type: "complete"; id: string; modelKey: TinyLocalModelKey; prompt: string; maxTokens?: number }
| { type: "download"; id: string; modelKey: TinyLocalModelKey }
| { type: "close" };
export type TinyTitleWorkerOutbound =
| { type: "pong"; id: string }
| { type: "title"; id: string; title: string | null }
| { type: "completion"; id: string; text: string | null }
| { type: "downloaded"; id: string }
| { type: "error"; id: string; error: string }
| { type: "progress"; id: string; event: TinyTitleProgressEvent }
+52 -11
View File
@@ -11,7 +11,7 @@ import type {
import { getTinyModelsCacheDir, isCompiledBinary, prompt } from "@oh-my-pi/pi-utils";
import packageJson from "../../package.json" with { type: "json" };
import tinyTitleSystemPrompt from "../prompts/system/tiny-title-system.md" with { type: "text" };
import { getTinyTitleModelSpec, type TinyTitleLocalModelKey } from "./models";
import { getTinyLocalModelSpec, type TinyLocalModelKey, type TinyTitleLocalModelKey } from "./models";
import { formatTitleUserMessage, normalizeGeneratedTitle } from "./text";
import type {
TinyTitleProgressEvent,
@@ -24,6 +24,7 @@ const TITLE_PREFILL = "<title>";
const TITLE_CLOSE = "</title>";
const TITLE_MAX_NEW_TOKENS = 20;
const STOP_DECODE_WINDOW_TOKENS = 32;
const MEMORY_COMPLETION_MAX_NEW_TOKENS = 256;
const TINY_TITLE_SYSTEM_PROMPT = prompt.render(tinyTitleSystemPrompt);
const TRANSFORMERS_PACKAGE = "@huggingface/transformers";
const sourceRequire = createRequire(import.meta.url);
@@ -51,7 +52,7 @@ interface TransformersRuntime {
) => Promise<TextGenerationPipeline>;
}
const pipelines = new Map<TinyTitleLocalModelKey, Promise<TextGenerationPipeline>>();
const pipelines = new Map<TinyLocalModelKey, Promise<TextGenerationPipeline>>();
function resolveTransformersVersionSpec(): string {
const manifest = packageJson as {
@@ -175,7 +176,7 @@ async function runRuntimeInstall(runtimeDir: string): Promise<void> {
function sendRuntimeInstallProgress(
transport: TinyTitleTransport,
requestId: string,
modelKey: TinyTitleLocalModelKey,
modelKey: TinyLocalModelKey,
status: "initiate" | "download" | "done",
): void {
transport.send({
@@ -192,7 +193,7 @@ function sendRuntimeInstallProgress(
async function ensureCompiledTransformersRuntime(
transport: TinyTitleTransport,
requestId: string,
modelKey: TinyTitleLocalModelKey,
modelKey: TinyLocalModelKey,
): Promise<string> {
const runtimeDir = getTinyTitleRuntimeDir();
if (await isCompiledRuntimeInstalled(runtimeDir)) return runtimeDir;
@@ -221,7 +222,7 @@ function configureTransformers(transformers: TransformersRuntime): TransformersR
async function loadTransformers(
transport: TinyTitleTransport,
requestId: string,
modelKey: TinyTitleLocalModelKey,
modelKey: TinyLocalModelKey,
): Promise<TransformersRuntime> {
if (transformersRuntime) return transformersRuntime;
transformersRuntime = (async () => {
@@ -304,13 +305,14 @@ function sendProgress(
}
async function loadPipeline(
modelKey: TinyTitleLocalModelKey,
modelKey: TinyLocalModelKey,
transport: TinyTitleTransport,
requestId: string,
): Promise<TextGenerationPipeline> {
const spec = getTinyLocalModelSpec(modelKey);
if (!spec) throw new Error(`Unknown tiny local model: ${modelKey}`);
const cached = pipelines.get(modelKey);
if (cached) {
const spec = getTinyTitleModelSpec(modelKey);
void cached
.then(() => {
transport.send({
@@ -323,7 +325,6 @@ async function loadPipeline(
return cached;
}
const spec = getTinyTitleModelSpec(modelKey);
const transformers = await loadTransformers(transport, requestId, modelKey);
const startedAt = performance.now();
const loaded = transformers
@@ -334,7 +335,7 @@ async function loadPipeline(
})
.then(
generator => {
sendLog(transport, "debug", "tiny-title: local model loaded", {
sendLog(transport, "debug", "tiny-model: local model loaded", {
modelKey,
repo: spec.repo,
elapsedMs: Math.round(performance.now() - startedAt),
@@ -396,6 +397,41 @@ async function generateTitle(
return extractTinyTitle(output[0]?.generated_text ?? "");
}
function buildCompletionPrompt(generator: TextGenerationPipeline, promptText: string): string {
const chat = [{ role: "user", content: promptText }];
return `${generator.tokenizer.apply_chat_template(chat, {
add_generation_prompt: true,
tokenize: false,
enable_thinking: false,
})}`;
}
/**
* Generic single-turn completion used by Mnemosyne memory tasks (fact extraction
* and consolidation). The caller (Mnemosyne) supplies the full task prompt; we
* wrap it as the user turn, decode greedily, and return the raw text for the
* caller's own parser. Output is capped to keep CPU latency bounded.
*/
async function generateCompletion(
transport: TinyTitleTransport,
requestId: string,
modelKey: TinyLocalModelKey,
promptText: string,
maxTokens: number | undefined,
): Promise<string | null> {
const generator = await loadPipeline(modelKey, transport, requestId);
const text = buildCompletionPrompt(generator, promptText);
const requested = maxTokens ?? MEMORY_COMPLETION_MAX_NEW_TOKENS;
const maxNewTokens = Math.min(Math.max(1, requested), MEMORY_COMPLETION_MAX_NEW_TOKENS);
const output = (await generator(text, {
max_new_tokens: maxNewTokens,
do_sample: false,
return_full_text: false,
})) as TextGenerationStringOutput;
const generated = (output[0]?.generated_text ?? "").trim();
return generated === "" ? null : generated;
}
function releasePipelines(): void {
// Intentionally NOT calling `pipeline.dispose()`. transformers.js disposes the
// underlying onnxruntime InferenceSession, freeing native memory that Bun's
@@ -408,7 +444,7 @@ function releasePipelines(): void {
function enqueueRequest(
transport: TinyTitleTransport,
request: Extract<TinyTitleWorkerInbound, { type: "generate" | "download" }>,
request: Extract<TinyTitleWorkerInbound, { type: "generate" | "complete" | "download" }>,
): void {
generateQueue = generateQueue.then(
async () => {
@@ -422,7 +458,7 @@ function enqueueRequest(
async function handleQueuedRequest(
transport: TinyTitleTransport,
request: Extract<TinyTitleWorkerInbound, { type: "generate" | "download" }>,
request: Extract<TinyTitleWorkerInbound, { type: "generate" | "complete" | "download" }>,
): Promise<void> {
try {
if (request.type === "download") {
@@ -430,6 +466,11 @@ async function handleQueuedRequest(
transport.send({ type: "downloaded", id: request.id });
return;
}
if (request.type === "complete") {
const text = await generateCompletion(transport, request.id, request.modelKey, request.prompt, request.maxTokens);
transport.send({ type: "completion", id: request.id, text });
return;
}
const title = await generateTitle(transport, request.id, request.modelKey, request.message);
transport.send({ type: "title", id: request.id, title });
} catch (error) {
+10 -1
View File
@@ -1,8 +1,17 @@
# Changelog
## [Unreleased]
### Added
- Added `llm.extractionPrompt` runtime option to override the fact-extraction prompt template using `{text}` and `{lang}` placeholders
- Added `llm.consolidationPrompt` runtime option to override the consolidation sleep prompt template using `{memories}`, `{source}`, and `{memory_count}` placeholders
- Published `@oh-my-pi/pi-mnemosyne` to npm: the local SQLite memory engine is now built, checked, tested, and released through the monorepo CI pipeline alongside the other workspace packages.
- Exported the diagnostic inspector as the `@oh-my-pi/pi-mnemosyne/diagnose` subpath for coding-agent memory maintenance commands.
### Changed
- Changed fact extraction to prefer a configured runtime LLM completion path before host extraction, with automatic fallback when the configured completion returns no output or fails
### Fixed
- Fixed configured LLM fact extraction by using temperature 0 so re-ingesting the same text is deterministic and avoids near-duplicate extractions
+30 -2
View File
@@ -1,6 +1,7 @@
import { getDiagnostics, safeForLog } from "./extraction/diagnostics";
import { callHostLlm, getHostLlmBackend } from "./llm-backends";
import { callLocalLlm, callRemoteLlm, cleanOutput, llmAvailable } from "./local-llm";
import { callLocalLlm, callConfiguredCompletion, callRemoteLlm, cleanOutput, configuredLlmWillHandleCall, llmAvailable } from "./local-llm";
import { getMnemosyneRuntimeOptions } from "./runtime-options";
const TRUE_VALUES: Record<string, true> = { "1": true, true: true, yes: true, on: true };
@@ -64,7 +65,8 @@ User message: {text}
Extraction:`;
export function buildExtractionPrompt(text: string, detectedLang = "en"): string {
return EXTRACTION_PROMPT_TEMPLATE.split("{text}").join(text).split("{lang}").join(detectedLang);
const template = getMnemosyneRuntimeOptions()?.llm?.extractionPrompt ?? EXTRACTION_PROMPT_TEMPLATE;
return template.split("{text}").join(text).split("{lang}").join(detectedLang);
}
function stripFence(raw: string): string {
let s = raw.trim();
@@ -229,6 +231,32 @@ export async function extractFacts(text: string | null | undefined): Promise<str
}
const prompt = buildExtractionPrompt(text);
// Configured completion (host-injected runtime LLM, e.g. the coding-agent's smol
// or a local on-device model). Mirrors consolidation's precedence: when a
// complete() fn is wired, it is the chosen path. Extraction is deterministic
// (temperature 0) so re-ingesting the same content does not create near-dupes.
if (configuredLlmWillHandleCall()) {
diag.recordAttempt("host");
try {
const raw = await callConfiguredCompletion(prompt, 0, { maxTokens: llmMaxTokens() });
if (typeof raw === "string" && raw.trim() !== "") {
const facts = parseFacts(raw);
if (facts.length > 0) {
diag.recordSuccess("host", facts.length);
diag.recordCall({ succeeded: true });
return facts;
}
}
diag.recordNoOutput("host");
} catch (exc) {
diag.recordFailure("host", exc, "configured_completion_raised");
diag.recordCall({ succeeded: false });
console.warn(`extractFacts: configured completion raised: ${safeForLog(exc)}`);
return [];
}
return localFallback(prompt, text, diag);
}
try {
const [attempted, hostText] = await tryHostExtraction(prompt);
if (attempted) {
+4 -3
View File
@@ -125,7 +125,8 @@ function memoryLines(memories: readonly string[]): string {
}
function formatSleepPrompt(memories: readonly string[], source = ""): string | null {
const template = sleepPrompt();
const override = getMnemosyneRuntimeOptions()?.llm?.consolidationPrompt;
const template = override !== undefined && override !== "" ? override : sleepPrompt();
if (template === "") {
return null;
}
@@ -151,7 +152,7 @@ export function buildPrompt(memories: readonly string[], source = ""): string {
return `/no_think\n${header}\n\n${memoryLines(memories)}\n\nSummary:`;
}
async function callConfiguredCompletion(
export async function callConfiguredCompletion(
prompt: string,
temperature: number,
opts: MnemosyneLlmCompleteOptions = {},
@@ -214,7 +215,7 @@ function hostBackendWillHandleCall(): boolean {
return llmEnabled() && hostLlmEnabled() && getHostLlmBackend() !== null;
}
function configuredLlmWillHandleCall(): boolean {
export function configuredLlmWillHandleCall(): boolean {
return llmEnabled() && (activeCustomCompletion() !== undefined || activePiAiModel() !== undefined);
}
+7 -1
View File
@@ -194,13 +194,17 @@ function resolveRuntimeOptions(options: MnemosyneOptions): ResolvedMnemosyneRunt
const llmApiKey = options.llmApiKey ?? nestedLlm?.apiKey;
const llmMaxTokens = nestedLlm?.maxTokens;
const llmComplete = nestedLlm?.complete;
const llmExtractionPrompt = nestedLlm?.extractionPrompt;
const llmConsolidationPrompt = nestedLlm?.consolidationPrompt;
if (
llmEnabled !== undefined ||
llmBaseUrl !== undefined ||
llmApiKey !== undefined ||
llmModel !== undefined ||
llmMaxTokens !== undefined ||
llmComplete !== undefined
llmComplete !== undefined ||
llmExtractionPrompt !== undefined ||
llmConsolidationPrompt !== undefined
) {
llm = {
enabled: llmEnabled,
@@ -209,6 +213,8 @@ function resolveRuntimeOptions(options: MnemosyneOptions): ResolvedMnemosyneRunt
model: llmModel,
maxTokens: llmMaxTokens,
complete: llmComplete,
extractionPrompt: llmExtractionPrompt,
consolidationPrompt: llmConsolidationPrompt,
};
}
}
@@ -34,6 +34,10 @@ export interface MnemosyneLlmRuntimeOptions {
model?: string | Model<Api>;
maxTokens?: number;
complete?: MnemosyneLlmCompletion;
/** Override the fact-extraction prompt template ({text}/{lang}). Used to feed small local models a friendlier format. */
extractionPrompt?: string;
/** Override the consolidation/sleep prompt template ({memories}/{source}/{memory_count}). */
consolidationPrompt?: string;
}
export interface MnemosyneRuntimeOptions {
@@ -56,6 +60,8 @@ export interface ResolvedMnemosyneLlmRuntimeOptions {
model?: string | Model<Api>;
maxTokens?: number;
complete?: MnemosyneLlmCompletion;
extractionPrompt?: string;
consolidationPrompt?: string;
}
export interface ResolvedMnemosyneRuntimeOptions {