- Switched from `miniaudio` to `maudio` Rust crate and added `AudioCapture` and `AudioPlayback` native classes.
- Removed browser-side audio infrastructure including Web Audio API, audio worklet processor, and WebRTC runtime.
- Migrated STT recorder and transcriber modules to use native `AudioCapture` with callback-based streaming.
- Replaced streaming audio player with native `AudioPlayback` that writes PCM directly without TypeScript intermediaries.
- Removed ffmpeg, wav, and platform-specific playback commands from the audio toolchain.
Streamed PCM kept for the nonzero-exit replay is now dropped once the
utterance exceeds 60s (~5.8 MB at 24 kHz mono f32): the recovered failure
is a short clip that fits the pipe buffer before a broken backend dies,
while replaying a long already-played utterance would duplicate audio and
unbounded retention would defeat streaming for long input.
`omp say` for a short single-segment clip printed `spoke` but produced no
sound on hosts where the first streaming backend (bundled ffmpeg without a
pulse/alsa output device) spawns then exits nonzero. The single pipe write
succeeds into the OS buffer before the backend dies, so the drain's
broken-pipe replay never fires, and `player.end()` sets `#inputClosed`
before the exit lands, so the early-exit handler short-circuits and never
advances to paplay/aplay. The end-of-stream cleanup awaited the exit but
ignored its code.
StreamingAudioPlayer now retains the utterance PCM and, when the streaming
backend exits nonzero, replays it through per-file playback so short clips
still reach the speakers. Added injectable backend/file-playback seams and a
regression test exercising the nonzero-exit, clean-exit, and no-backend
paths.
Fixes#5875
- Implemented streaming synthesis in the `say` command to allow processing of arbitrarily long text without hitting model phoneme limits.
- Added file input support via the `--file` flag and updated the CLI to prevent conflicting arguments.
- Refined `SpeakableStream` segmentation logic to prioritize valid sentence and clause breaks within the maximum segment length when processing large text buffers.
- Added comprehensive tests for stream segmentation behavior under long-form input.
- Added an enhanced speech pipeline that utilizes small models to rewrite text for natural language synthesis.
- Implemented `SpeechEnhancer` and `BlockAccumulator` to manage fence-aware text streaming and paragraph splitting.
- Configured a new `speech.enhanced` setting to toggle between mechanical and enhanced vocalization modes.
- Resolved `EPIPE` rejections during speech playback by ensuring stream flushes and suppressing stop-related errors.
- Added `SpeakableStream` to strip markdown noise, silence code blocks and tables, normalize links and paths, and emit sentence/clause segments.
- Reworked `Vocalizer` to segment assistant deltas in the parent process, lazily open TTS streams, idle-flush partial thoughts, and chain playback sessions.
- Added gapless streaming playback with ffmpeg/sox backends, ducking-aware pacing, fallback file playback, and immediate stop handling.
- Added IPC `sendAndFlush` support and used it in the TTS worker so audio chunks drain before blocking ONNX inference resumes.
- Added speakable-stream coverage for markdown filtering, segmentation latency, idle flushing, and forced long-segment splits.