Extended ONNX Runtime CUDA sidecar repair to PI_TINY_DEVICE=auto and gpu on Linux x64, matching the CUDA-capable diagnostics path.
Added focused regression coverage for both generic accelerated settings.
Refs #4475
Deferred ONNX Runtime CUDA sidecar repair failure into the runtime metadata so loadTransformersRuntime keeps loading and loadPipelineWithDeviceFallback still gets its CUDA→CPU retry when NuGet is offline or the ort install script is unavailable. The failure surfaces through the CUDA diagnostics helper instead of hard-erroring the tiny worker.
Refs #4475
Downloaded missing ONNX Runtime CUDA provider sidecars when the compiled tiny-model side runtime is used with PI_TINY_DEVICE=cuda.
Preserved actionable CUDA worker diagnostics in tiny-models text output and added focused regression coverage.
Fixes#4475
Bun keeps the parent event loop alive when an unref'd child has a piped
stderr stream. The first #4324 fix started with stderr: "pipe", so a
long-lived idle TTS/STT/tiny/mnemopi worker could keep short CLI commands
alive even after proc.unref().
Switch worker stderr capture to a temp-file fd target instead of a Bun
ReadableStream pipe. The parent does not start any JS read while the
worker is alive; after onExit it reads the bounded tail from the file,
logs captured lines, appends the tail to the surfaced worker Error, then
closes and removes the temp capture.
Add a regression that spawns a non-test wrapper process with an idle
unref'd worker and asserts the wrapper exits immediately. Existing stderr
capture, truncation, and intentional SIGKILL behavior remain covered.
Fixes#4324
Inference worker subprocesses (TTS, STT, tiny-model, mnemopi embeddings)
were spawned with stderr: "ignore", so a native crash inside the child
was completely discarded. The parent only ever logged the bare exit code
(e.g. Kokoro TTS's recurring "tts subprocess exited with code 7"),
leaving the recurring crash loop undiagnosable.
createWorkerSubprocess now pipes stderr and drains it in the parent:
- Each decoded stderr line is forwarded to logger.debug under
"<exitLabel> stderr" so operators get live visibility on chatty native
runtimes without touching the chat scrollback.
- A bounded 16 KiB ring keeps the tail of stderr so the eventual exit
Error carries the actual crash reason (ONNX Runtime traceback, glibc
assertion, etc.) instead of "code 7" alone. The prefix is preserved so
existing log grepping keeps working.
- The exit event and the stderr pipe are independent, so a synchronous
read in onExit would race the drain. SpawnedSubprocess grew a
stderrDrained: Promise<void>, and onExit chains the error surface off
it so callers see the whole tail. Tests can await stderrDrained
deterministically instead of racing wall-clock timers.
- Intentional terminate() SIGKILLs still stay silent — signal-exit
gating on intentionalExit is unchanged.
Fixes#4324
Rejoined split Windows extension module paths before launch parsing finishes and stripped extended-length Win32 prefixes before Bun import and worker spawn APIs see them.
Fixes#3804
- Established shared `subprocess` infrastructure to standardize worker IPC, environment resolution, and error handling.
- Migrated mnemopi, speech-to-text, tiny-model, and TTS clients to utilize consolidated worker-client utilities.
- Implemented common runtime helpers for ONNX model loading, logging, and process readiness probing.
- Eliminated redundant local spawn logic and environment mapping across all inference worker clients.