Commit Graph
12 Commits
Author SHA1 Message Date
can1357 79ca069d9c feat(coding-agent): supported streaming TTS for long text inputs
- Implemented streaming synthesis in the `say` command to allow processing of arbitrarily long text without hitting model phoneme limits.
- Added file input support via the `--file` flag and updated the CLI to prevent conflicting arguments.
- Refined `SpeakableStream` segmentation logic to prioritize valid sentence and clause breaks within the maximum segment length when processing large text buffers.
- Added comprehensive tests for stream segmentation behavior under long-form input.
2026-07-06 04:12:39 +02:00
can1357 600191b6f5 Merge PR #4356: fix(coding-agent): size title budget for backends that ignore disableReasoning (@roboomp) 2026-07-05 13:25:25 +02:00
roboomp 8cce6f637c fix(tts): stopped queued TTS playback on Esc after stream end
Once the assistant reply stops streaming, `vocalizer.clear()` was only invoked from the aborted-stream cascade in EventController. Escaping after the model finished fell through InputController to the empty-editor double-Esc gesture while StreamingAudioPlayer kept draining buffered Kokoro PCM.

Add `Vocalizer.isSpeaking()` (true while any live player, stream handle, or in-flight abort is around) and consult it in the Esc handler before the double-Esc branch: if speech is still audible, a single Esc calls `vocalizer.clear()` and resets `lastEscapeTime` so tree/branch stays reachable via the next press.

Fixes #4521
2026-07-04 14:42:24 +00:00
roboomp 8886a528dc fix(coding-agent): size title/commit/speech/classifier budgets for backends that ignore disableReasoning
Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.

Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.

Sites fixed:
- utils/title-generator.ts       (30   -> 1024)
- utils/commit-message-generator (60   -> 1024)
- tts/speech-enhancer            (512  -> 1536)
- auto-thinking/classifier       online path (8  -> 1024); classifyLocal
                                  keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
                                  classifyLocal keeps ANSWER_MAX_TOKENS

Fixes #4355
2026-07-03 00:47:20 +00:00
can1357 94645752ff feat(coding-agent/tts): introduced natural speech synthesis pipeline
- Added an enhanced speech pipeline that utilizes small models to rewrite text for natural language synthesis.
- Implemented `SpeechEnhancer` and `BlockAccumulator` to manage fence-aware text streaming and paragraph splitting.
- Configured a new `speech.enhanced` setting to toggle between mechanical and enhanced vocalization modes.
- Resolved `EPIPE` rejections during speech playback by ensuring stream flushes and suppressing stop-related errors.
2026-07-02 08:59:13 +02:00
can1357 41cc57c238 feat(coding-agent/tts): redesigned streaming speech vocalization
- Added `SpeakableStream` to strip markdown noise, silence code blocks and tables, normalize links and paths, and emit sentence/clause segments.
- Reworked `Vocalizer` to segment assistant deltas in the parent process, lazily open TTS streams, idle-flush partial thoughts, and chain playback sessions.
- Added gapless streaming playback with ffmpeg/sox backends, ducking-aware pacing, fallback file playback, and immediate stop handling.
- Added IPC `sendAndFlush` support and used it in the TTS worker so audio chunks drain before blocking ONNX inference resumes.
- Added speakable-stream coverage for markdown filtering, segmentation latency, idle flushing, and forced long-segment splits.
2026-07-02 08:30:58 +02:00
can1357 1750973503 refactor(coding-agent): established shared subprocess infrastructure to
- Established shared `subprocess` infrastructure to standardize worker IPC, environment resolution, and error handling.
- Migrated mnemopi, speech-to-text, tiny-model, and TTS clients to utilize consolidated worker-client utilities.
- Implemented common runtime helpers for ONNX model loading, logging, and process readiness probing.
- Eliminated redundant local spawn logic and environment mapping across all inference worker clients.
2026-06-22 05:50:07 +02:00
oldschoola 23565e3e58 fix(ipc): hardened worker IPC send against async EPIPE rejections
- Added `safeSend` helper wrapping `Subprocess.send()` so sync throws and async EPIPE rejections cannot escape.
- Replaced inline try/catch send wrappers in STT, TTS, and tiny-title clients with shared `safeSend`.
- Added `isIpcSendEpipe` predicate and made matching rejections non-fatal in the `unhandledRejection` handler.
- Added contract tests for `safeSend` and `isIpcSendEpipe` covering sync throws, async rejections, and edge cases.
2026-06-18 21:42:02 -07:00
can1357 3ab675a83b fix: added buffered worker inboxing and standardized worker selectors
- Added `WorkerInbox` and `installWorkerInbox(port)` to queue worker messages before bind.
- Added `consumeWorkerInbox()` to replay buffered messages and clear one active inbox.
- Added buffered inbox consumption in JS and tab worker transports before direct message handlers.
- Normalized worker selector arguments to the `__omp_worker_*` naming across workers and tests.
2026-06-15 11:59:44 +02:00
roboomp d7d94949fa fix(tts): isolated kokoro onnxruntime loading
Loaded Kokoro's side-installed transformers runtime by absolute path before requiring kokoro-js, avoiding host/workspace onnxruntime libraries in the worker process.

Kept runtime-cache bare module requests inside the registered runtime cache when the parent module is already inside that cache, and covered the resolver boundary with a regression test.

Fixes #2591
2026-06-14 21:56:50 +00:00
can1357 687487385b test(coding-agent): drop tts/stt worker tests; harden timing-bound tests
The tts/stt suite spawns sherpa-onnx worker subprocesses that hang on the
headless CI runner (zero-output stall → SIGTERM), and the resulting event-loop
starvation tipped real-time TUI tests (streaming-preview, custom-editor shimmer,
ask timeouts) past bun's 5s default. Remove the tts/stt tests and give the
timing-sensitive tests an explicit 30s timeout.
2026-06-14 20:18:09 +02:00
can1357 b830f7912b feat: unified speech setup and introduced local STT/TTS capabilities
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.
2026-06-14 16:07:44 +02:00