Commit Graph
23 Commits
Author SHA1 Message Date
roboomp 765a54305e fix(stt): preserved nearest sherpa runtime
- Loaded the nearest wrapper first and limited ancestor fallback to missing-addon failures.

- Covered nested version-conflict installs alongside the broken-hoist layout.
2026-07-26 11:18:31 +00:00
roboomp 838fad7e5c fix(stt): resolved workspace sherpa addon loading
- Selected the sherpa wrapper colocated with its platform-native package in source workspaces.

- Added a regression test for Bun workspace hoisting.

Fixes #6690
2026-07-26 11:08:06 +00:00
can1357 c9c0882724 feat(audio): replaced Chromium browser audio with native audio stack
- Switched from `miniaudio` to `maudio` Rust crate and added `AudioCapture` and `AudioPlayback` native classes.
- Removed browser-side audio infrastructure including Web Audio API, audio worklet processor, and WebRTC runtime.
- Migrated STT recorder and transcriber modules to use native `AudioCapture` with callback-based streaming.
- Replaced streaming audio player with native `AudioPlayback` that writes PCM directly without TypeScript intermediaries.
- Removed ffmpeg, wav, and platform-specific playback commands from the audio toolchain.
2026-07-24 08:54:16 +02:00
roboomp 26e75b251c fix(stt): selected supported linux ffmpeg input
- Probed Linux ffmpeg demuxers and fell back to ALSA when PulseAudio input is unavailable.
- Preserved recorder stderr so immediate capture failures report their real cause.
- Added regressions for ALSA selection and stderr diagnostics.

Fixes #5907
2026-07-17 20:25:21 +00:00
roboomp 5e5c88979f style: bun run fix 2026-07-01 00:07:30 +00:00
roboomp efcd6cafd6 fix(stt): kept whisper setup downloads alive
Kept the STT subprocess referenced while download and stream requests are pending so setup cannot exit before the worker answers.

Propagated worker download errors to setup callers and verified completed downloads leave the expected cache files.

Fixes #3939
2026-07-01 00:07:14 +00:00
can1357 f50caede0b test(coding-agent): added integration tests for STT submit trigger behaviors
- Added comprehensive integration tests verifying Speech-to-Text submit triggers in STTController under different configuration settings.
- Cleaned up unused imports and types from the stt-controller source file.
- Documented Speech-to-Text submit trigger updates and programmatic editor submission changes in package changelogs.
2026-06-30 16:25:34 +02:00
can1357 168bdae5da feat(coding-agent): introduced automatic dictation submission triggers
- Added `stt.submitTrigger` setting to control automatic dictation submission.
- Introduced `evaluateSubmitTrigger` to process sentence punctuation and spoken cues.
- Integrated trigger evaluation into batch and streaming `STTController` pipelines.
- Implemented trailing word trimming to strip trigger words like "submit" before sending.
- Added comprehensive unit tests for all trigger evaluation behaviors.
2026-06-30 16:16:39 +02:00
can1357 1750973503 refactor(coding-agent): established shared subprocess infrastructure to
- Established shared `subprocess` infrastructure to standardize worker IPC, environment resolution, and error handling.
- Migrated mnemopi, speech-to-text, tiny-model, and TTS clients to utilize consolidated worker-client utilities.
- Implemented common runtime helpers for ONNX model loading, logging, and process readiness probing.
- Eliminated redundant local spawn logic and environment mapping across all inference worker clients.
2026-06-22 05:50:07 +02:00
oldschoola 23565e3e58 fix(ipc): hardened worker IPC send against async EPIPE rejections
- Added `safeSend` helper wrapping `Subprocess.send()` so sync throws and async EPIPE rejections cannot escape.
- Replaced inline try/catch send wrappers in STT, TTS, and tiny-title clients with shared `safeSend`.
- Added `isIpcSendEpipe` predicate and made matching rejections non-fatal in the `unhandledRejection` handler.
- Added contract tests for `safeSend` and `isIpcSendEpipe` covering sync throws, async rejections, and edge cases.
2026-06-18 21:42:02 -07:00
can1357 37646a4a14 fix(coding-agent): refined STT warmup and mechanical space-hold handling
- Reworked space-bar hold detection to require sustained mechanical gaps.
- Implemented per-model STT dependency tracking with background warmup and deferred setup.
- Tightened Whisper cache validation to require both encoder and decoder models.
- Caught STT request promise rejections to avoid unhandled rejection noise.
2026-06-17 01:57:14 +02:00
can1357 3ab675a83b fix: added buffered worker inboxing and standardized worker selectors
- Added `WorkerInbox` and `installWorkerInbox(port)` to queue worker messages before bind.
- Added `consumeWorkerInbox()` to replay buffered messages and clear one active inbox.
- Added buffered inbox consumption in JS and tab worker transports before direct message handlers.
- Normalized worker selector arguments to the `__omp_worker_*` naming across workers and tests.
2026-06-15 11:59:44 +02:00
can1357 687487385b test(coding-agent): drop tts/stt worker tests; harden timing-bound tests
The tts/stt suite spawns sherpa-onnx worker subprocesses that hang on the
headless CI runner (zero-output stall → SIGTERM), and the resulting event-loop
starvation tipped real-time TUI tests (streaming-preview, custom-editor shimmer,
ask timeouts) past bun's 5s default. Remove the tts/stt tests and give the
timing-sensitive tests an explicit 30s timeout.
2026-06-14 20:18:09 +02:00
can1357 9fbd8288ab fix(stt): use macOS default audio input (:default) for ffmpeg avfoundation capture
Hardcoding device index `:0` grabbed whatever avfoundation enumerated first
(often a camera or the wrong input), so recording could capture silence or the
wrong source. Both the single-shot and streaming ffmpeg recorder paths now
request the system default input device.
2026-06-14 17:14:37 +02:00
can1357 b830f7912b feat: unified speech setup and introduced local STT/TTS capabilities
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.
2026-06-14 16:07:44 +02:00
can1357 6d07944654 refactor: migrated binary detection to $which() utility across codebase
- Replaced all Bun.which() calls with $which() utility from @oh-my-pi/pi-utils across 22 files.
- Removed findBashOnPath() wrapper function from procmgr.ts, consolidating binary path resolution.
- Updated AGENTS.md documentation to reflect new $which() API usage pattern.
- Centralized binary detection logic through shared utility, reducing code duplication.
2026-04-08 05:28:22 +02:00
can1357 85c5fad2bc refactor: replace named re-exports with star re-exports in barrel files
Convert export { A, B, ... } from and export type { ... } from blocks
to export * from across all packages (ai, coding-agent, natives, tui, utils).
2026-02-28 22:34:02 +01:00
can1357 f20a964986 feat(coding-agent): added animated microphone icon and skill discovery via symlinks, improved STT feedback
- Added animated microphone icon with color cycling during voice recording and transcription states.
- Added support for discovering skills via symbolic links in the skills directory.
- Changed STT status messages to display via state change callbacks instead of dedicated status line segment.
- Removed dedicated STT status line segment in favor of animated cursor-based feedback.
- Added cursor override feature to TUI editor for customizing end-of-text cursor glyph with ANSI-styled strings.
2026-02-15 13:37:43 +01:00
can1357 d1f41ae553 merge: PR #51 — speech-to-text with Alt+H (closes #51)
- Moved /stt slash command to 'omp setup stt' CLI command
- Fixed pip install blocking TUI (sync -> async)
- Fixed 'Transcribing...' status not clearing after success
2026-02-15 10:30:47 +01:00
can1357 9e32219a21 fix(stt): address review feedback 2026-02-15 10:23:48 +01:00
can1357 629f0a1d60 fix(stt): added concurrency guard and dispose cancellation for STT controller 2026-02-15 10:23:48 +01:00
can1357 b56e8a72fc fix(stt): set stderr to ignore on recording spawns and removed exists() anti-pattern 2026-02-15 10:23:48 +01:00
daaximusandcan1357 6969076386 feat(coding-agent): add speech-to-text with Alt+H keybinding
Cross-platform audio recording with automatic tool detection and
fallback chain (SoX > FFmpeg > arecord > PowerShell mciSendString).
Transcription via Python openai-whisper with automatic pip install
on first use.

Recording:
- SoX with explicit waveaudio device on Windows (-t waveaudio 0)
- FFmpeg with auto-detected dshow device name on Windows
- arecord for ALSA on Linux
- PowerShell mciSendString as zero-dependency Windows fallback
- Process health check after spawn, graceful stop for FFmpeg/PS
- Fallback chain: if one tool fails, tries the next automatically

Transcription:
- Python openai-whisper as primary backend (pip install openai-whisper)
- Custom WAV loader in transcribe.py using Python wave module
- Resamples to 16kHz mono via numpy (no ffmpeg dependency)
- Passes float32 numpy array directly to whisper.transcribe()
- 120s timeout with process kill guard

TUI integration:
- Alt+H keybinding to toggle recording (configurable in keybindings)
- /stt command with on|off|status|setup subcommands
- Status line segment showing REC/STT state
- Transcribed text inserted directly into editor prompt
- Lazy controller instantiation on first use

Settings:
- stt.enabled (default: false)
- stt.language (default: en)
- stt.modelName (default: base.en)

New files: packages/coding-agent/src/stt/{downloader,index,recorder,
setup,stt-controller,transcribe.py,transcriber}.ts
2026-02-15 10:23:48 +01:00