- Changed tiny-device preference resolution to always default to CPU instead of platform-specific DirectML/CUDA heuristics.
- Updated settings schema, documentation, and changelog text to describe the CPU default while keeping accelerated providers behind explicit `providers.tinyModelDevice`/`PI_TINY_DEVICE` choices.
- Added persistent settings in the Providers tab for ONNX execution provider and quantization/precision, replacing env-var-only configuration.
- PI_TINY_DEVICE and PI_TINY_DTYPE env vars still override the matching setting at spawn time.
- Added tinyWorkerEnvOverlay to map settings onto worker env without clobbering explicit env vars.
- Updated docs and tests to reflect the new setting-first resolution order.
- Added `PI_TINY_DEVICE` env var to control ONNX execution provider (`gpu` default, `cpu`, `metal`, `cuda`, `dml`, `coreml`).
- Local tiny-model inference now tries accelerated GPU provider first and retries on CPU if initialization fails.
- Added `device.ts` module with normalization, preference resolution, and load-order helpers.
- Updated docs and model descriptions to drop CPU-specific language.
- Added Mnemosyne runtime `extractionPrompt` and `consolidationPrompt` options and wired them into resolved LLM config.
- Added fact-extraction branch to call configured completion first (temp 0), then parse facts and safely fall back.
- Added tiny local model registry features for memory/title, including keys, specs, and validation helpers.
- Added `complete` protocol messages and abort-aware client/worker paths for local title completion generation.
- Added local-models documentation for tiny/memory transformer paths, defaults, and known parser caveats.