49415712f2
The memory-extraction prompt concatenated its instructions, few-shot examples, and the user message into a single user turn, so a small local model could not distinguish instructions from input and frequently echoed the Globex/weather examples instead of extracting facts. Send the instructions as a real system turn and the raw text as the user turn. The tiny worker protocol gains a systemPrompt field, and Mnemopi completion input carries task metadata so the backend selects the right prompt per call. Drop the code-built MEMORY_EXTRACTION_TEMPLATE rather than porting it: prompt text belongs in .md files, and resolveMemoryCompletionInput already overrides that template for every extraction call, so Mnemopi rendered it only for the result to be discarded. Measured on ONNX q4 CPU, LFM2.5-1.2B memory extraction improved from 1/8 to 5/8 once the roles were separated.