- Loaded the nearest wrapper first and limited ancestor fallback to missing-addon failures.
- Covered nested version-conflict installs alongside the broken-hoist layout.
- Selected the sherpa wrapper colocated with its platform-native package in source workspaces.
- Added a regression test for Bun workspace hoisting.
Fixes#6690
- Switched from `miniaudio` to `maudio` Rust crate and added `AudioCapture` and `AudioPlayback` native classes.
- Removed browser-side audio infrastructure including Web Audio API, audio worklet processor, and WebRTC runtime.
- Migrated STT recorder and transcriber modules to use native `AudioCapture` with callback-based streaming.
- Replaced streaming audio player with native `AudioPlayback` that writes PCM directly without TypeScript intermediaries.
- Removed ffmpeg, wav, and platform-specific playback commands from the audio toolchain.
- Probed Linux ffmpeg demuxers and fell back to ALSA when PulseAudio input is unavailable.
- Preserved recorder stderr so immediate capture failures report their real cause.
- Added regressions for ALSA selection and stderr diagnostics.
Fixes#5907
Kept the STT subprocess referenced while download and stream requests are pending so setup cannot exit before the worker answers.
Propagated worker download errors to setup callers and verified completed downloads leave the expected cache files.
Fixes#3939
- Added comprehensive integration tests verifying Speech-to-Text submit triggers in STTController under different configuration settings.
- Cleaned up unused imports and types from the stt-controller source file.
- Documented Speech-to-Text submit trigger updates and programmatic editor submission changes in package changelogs.
- Added `stt.submitTrigger` setting to control automatic dictation submission.
- Introduced `evaluateSubmitTrigger` to process sentence punctuation and spoken cues.
- Integrated trigger evaluation into batch and streaming `STTController` pipelines.
- Implemented trailing word trimming to strip trigger words like "submit" before sending.
- Added comprehensive unit tests for all trigger evaluation behaviors.
- Established shared `subprocess` infrastructure to standardize worker IPC, environment resolution, and error handling.
- Migrated mnemopi, speech-to-text, tiny-model, and TTS clients to utilize consolidated worker-client utilities.
- Implemented common runtime helpers for ONNX model loading, logging, and process readiness probing.
- Eliminated redundant local spawn logic and environment mapping across all inference worker clients.
- Added `safeSend` helper wrapping `Subprocess.send()` so sync throws and async EPIPE rejections cannot escape.
- Replaced inline try/catch send wrappers in STT, TTS, and tiny-title clients with shared `safeSend`.
- Added `isIpcSendEpipe` predicate and made matching rejections non-fatal in the `unhandledRejection` handler.
- Added contract tests for `safeSend` and `isIpcSendEpipe` covering sync throws, async rejections, and edge cases.
- Added `WorkerInbox` and `installWorkerInbox(port)` to queue worker messages before bind.
- Added `consumeWorkerInbox()` to replay buffered messages and clear one active inbox.
- Added buffered inbox consumption in JS and tab worker transports before direct message handlers.
- Normalized worker selector arguments to the `__omp_worker_*` naming across workers and tests.
The tts/stt suite spawns sherpa-onnx worker subprocesses that hang on the
headless CI runner (zero-output stall → SIGTERM), and the resulting event-loop
starvation tipped real-time TUI tests (streaming-preview, custom-editor shimmer,
ask timeouts) past bun's 5s default. Remove the tts/stt tests and give the
timing-sensitive tests an explicit 30s timeout.
Hardcoding device index `:0` grabbed whatever avfoundation enumerated first
(often a camera or the wrong input), so recording could capture silence or the
wrong source. Both the single-shot and streaming ffmpeg recorder paths now
request the system default input device.
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.
- Replaced all Bun.which() calls with $which() utility from @oh-my-pi/pi-utils across 22 files.
- Removed findBashOnPath() wrapper function from procmgr.ts, consolidating binary path resolution.
- Updated AGENTS.md documentation to reflect new $which() API usage pattern.
- Centralized binary detection logic through shared utility, reducing code duplication.
- Added animated microphone icon with color cycling during voice recording and transcription states.
- Added support for discovering skills via symbolic links in the skills directory.
- Changed STT status messages to display via state change callbacks instead of dedicated status line segment.
- Removed dedicated STT status line segment in favor of animated cursor-based feedback.
- Added cursor override feature to TUI editor for customizing end-of-text cursor glyph with ANSI-styled strings.
Cross-platform audio recording with automatic tool detection and
fallback chain (SoX > FFmpeg > arecord > PowerShell mciSendString).
Transcription via Python openai-whisper with automatic pip install
on first use.
Recording:
- SoX with explicit waveaudio device on Windows (-t waveaudio 0)
- FFmpeg with auto-detected dshow device name on Windows
- arecord for ALSA on Linux
- PowerShell mciSendString as zero-dependency Windows fallback
- Process health check after spawn, graceful stop for FFmpeg/PS
- Fallback chain: if one tool fails, tries the next automatically
Transcription:
- Python openai-whisper as primary backend (pip install openai-whisper)
- Custom WAV loader in transcribe.py using Python wave module
- Resamples to 16kHz mono via numpy (no ffmpeg dependency)
- Passes float32 numpy array directly to whisper.transcribe()
- 120s timeout with process kill guard
TUI integration:
- Alt+H keybinding to toggle recording (configurable in keybindings)
- /stt command with on|off|status|setup subcommands
- Status line segment showing REC/STT state
- Transcribed text inserted directly into editor prompt
- Lazy controller instantiation on first use
Settings:
- stt.enabled (default: false)
- stt.language (default: en)
- stt.modelName (default: base.en)
New files: packages/coding-agent/src/stt/{downloader,index,recorder,
setup,stt-controller,transcribe.py,transcriber}.ts