- Removed the `fflate` dependency in favor of using `node:zlib` for ZIP operations.
- Updated documentation and internal code comments to reflect the transition to native `node:zlib` DEFLATE support.
- Replaced raw TypeScript documentation map with a lazily-inflated gzip blob.
- Reduced bundled binary/npm package size by approximately 0.9MB.
- Encapsulated index logic into `docs-index.ts` to separate header metadata from content.
- Added `postpack` and robust `try/finally` patterns in build scripts to ensure clean artifacts.
- Implemented disk-based fallback during development to maintain existing developer experience.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
Confirmed root cause of the CI test hang: bun --parallel spawns one isolated
worker per core, and each worker loads the 116MB pi-natives addon plus a large
JS heap. On the 16GB hosted runner that exceeds memory, triggering swap thrash
(100s+ event-loop stalls, transient file-read failures) that looks like a hang.
Capping to 2 workers keeps per-file memory recycling (isolation) while halving
peak memory so it fits the runner.
- Added unified `omp setup speech` flow with JSON/check modes and model picker.
- Added local STT pipeline with sherpa workers, recorder/download flow, and streaming inference.
- Added local TTS pipeline with `omp say`, backend selection, and streaming vocalization.
- Replaced legacy speech settings with unified `speech`/`speechgen` configuration keys.