Read vLLM max_model_len and OpenAI-compatible context_length metadata during model discovery, route providers.vllm.baseUrl into built-in discovery before cached models exist, and avoid sending local placeholder bearer tokens.
Scope the vLLM model cache to the discovery base URL so endpoint changes refetch immediately, and add focused regression coverage for configured and built-in vLLM discovery.
Add a proxy-discovery regression: a `context_length: 0` upstream value
must be rejected by `toPositiveNumberOrUndefined` and fall back to the
default window. Raw `??` would have pinned it at 0 (nullish coalescing
does not skip 0), so this guards the should-fix the helper was added for.
Also apply biome formatting to the discovery change (collapse the
multiline `contextWindow:` expression, drop a trailing space) and the new
tests so `biome check` passes — the PR as submitted failed the formatter.
Addresses review feedback on #2466.
Updated model discovery to use `context_length` reported by the API when available, falling back to bundled reference data and default. Added `context_length` field to parsed model response and modified context window assignment logic.
Deferred MCP discovery wrote 'Connecting to MCP servers: …' straight to process.stderr while the TUI owned the terminal, overdrawing the chat input box border. onMCPConnecting now emits McpConnectingEvent on the mcp:connecting channel; InteractiveMode subscribes and renders it via showStatus (status container), mirroring the LSP-startup pattern. New mcp/startup-events.ts holds the channel, type, and formatMCPConnectingMessage.
Hardcoding device index `:0` grabbed whatever avfoundation enumerated first
(often a camera or the wrong input), so recording could capture silence or the
wrong source. Both the single-shot and streaming ffmpeg recorder paths now
request the system default input device.
Treated SearXNG HTTP 200 responses with no usable sources and upstream engine failures as transient provider errors. Added a generic renderable-content guard so provider fallback continues instead of returning an invisible success.\n\nFixes #2571
- Tracked snapshot-safe boundaries to retain commit-stable streamed rows in scrollback.
- Computed durableBoundary from commit and snapshot ends to prevent row drop regressions.
- Updated audit-row handling so drifting durable rows were excluded from resync checks.
- Added regression tests for commit-stable and commit-unstable relayout streaming cases.
- agent-loop: raise repetition-detection floor to 180 chars and clear thinking
replay anchors when collapsing a detected loop.
- providers/google: ignore empty text parts, retain terminal thoughtSignatures,
and stop function-call signatures clobbering the prior block.
- autolearn: capture goal-mode at the turn boundary; harden managed-skill writes
against hard-links/symlinks (O_NOFOLLOW + nlink); refuse minting managed skills
whose name an authored skill already claims.
- eager tasks: thread agentKind through the session so a custom top-level agentId
still gets always-mode delegation; split Eager Tasks prompt into hard vs soft.
- title-generator: race the online title model against a local tiny-model fallback.
- eager-todo: keep the soft reminder aligned with the todo init schema.
- mcp/stdio: keep close() detaching the read loop instead of awaiting it.
- stream loop: fix collapsing and tool-call thought-signature handling.
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.