Downloaded missing ONNX Runtime CUDA provider sidecars when the compiled tiny-model side runtime is used with PI_TINY_DEVICE=cuda.
Preserved actionable CUDA worker diagnostics in tiny-models text output and added focused regression coverage.
Fixes#4475
resolveModels("all") expanded the full TINY_LOCAL_MODELS registry, which now
includes the qwen3-1.7b entry marked unsupportedReason. loadPipeline() throws
for such specs, so the download worker reported it as failed and the bulk
command exited with "One or more tiny title models failed to download" even
when every usable model downloaded. Filter unsupported specs out of the `all`
prefetch path; explicit single-model requests are unchanged. Addresses the
unaddressed Codex P2 on PR #3133.
- Added a new `providers.memoryModel` setting with tiny memory model options and `ONLINE_MEMORY_MODEL_KEY` default in settings.
- Updated Mnemosyne provider resolution so a configured local tiny model overrode remote completion and used new memory extraction and consolidation prompts.
- Expanded the tiny-model CLI registry to download and report all local tiny models (title plus memory) through a unified list.
- Updated hashline streaming preview tests to generate snapshot-tagged section headers and use an in-memory snapshot store.
- Replaced untagged file markers in multi-section preview inputs with `formatHashlineHeader` values derived from recorded file contents.
- Changed the bash command error test to expect a returned `isError` result with exit code 1 instead of a rejected promise.
- Added tiny-title protocol contracts, including progress-state unions, message payloads, and transport interfaces.
- Added title text utilities to truncate long inputs, wrap `<user-message>` blocks, and normalize generated titles.
- Added tiny-title model registry and helpers with type-safe keys and runtime optional loading via optionalDependencies.
- Added client-side worker orchestration with spawn fallback, request queueing, progress/error routing, and smoke-test APIs.
- Added worker runtime for model resolution, prompt-based inference, lock-based install retries, and close-time cache clear.