Extended ONNX Runtime CUDA sidecar repair to PI_TINY_DEVICE=auto and gpu on Linux x64, matching the CUDA-capable diagnostics path.
Added focused regression coverage for both generic accelerated settings.
Refs #4475
Deferred ONNX Runtime CUDA sidecar repair failure into the runtime metadata so loadTransformersRuntime keeps loading and loadPipelineWithDeviceFallback still gets its CUDA→CPU retry when NuGet is offline or the ort install script is unavailable. The failure surfaces through the CUDA diagnostics helper instead of hard-erroring the tiny worker.
Refs #4475
Downloaded missing ONNX Runtime CUDA provider sidecars when the compiled tiny-model side runtime is used with PI_TINY_DEVICE=cuda.
Preserved actionable CUDA worker diagnostics in tiny-models text output and added focused regression coverage.
Fixes#4475
- Established shared `subprocess` infrastructure to standardize worker IPC, environment resolution, and error handling.
- Migrated mnemopi, speech-to-text, tiny-model, and TTS clients to utilize consolidated worker-client utilities.
- Implemented common runtime helpers for ONNX model loading, logging, and process readiness probing.
- Eliminated redundant local spawn logic and environment mapping across all inference worker clients.