- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
Downloaded missing ONNX Runtime CUDA provider sidecars when the compiled tiny-model side runtime is used with PI_TINY_DEVICE=cuda.
Preserved actionable CUDA worker diagnostics in tiny-models text output and added focused regression coverage.
Fixes#4475
- Updated `normalizeGeneratedTitle` to reconcile model-generated titles against the user's input instead of forcing title-case.
- Added logic to restore distinctive proper-noun casing (e.g., `TinyVMM`) and flatten model-generated camelCase artifacts (e.g., `dAemon`) that do not appear in the user's message.
- Ensured model-cased proper nouns that are not in the source message (e.g., `GitHub`) are preserved.
- Established shared `subprocess` infrastructure to standardize worker IPC, environment resolution, and error handling.
- Migrated mnemopi, speech-to-text, tiny-model, and TTS clients to utilize consolidated worker-client utilities.
- Implemented common runtime helpers for ONNX model loading, logging, and process readiness probing.
- Eliminated redundant local spawn logic and environment mapping across all inference worker clients.
Blocked the unsupported Qwen3 1.7B ONNX memory model before loading transformers and remembered local model execution failures so the client returns null instead of spawning another __omp_worker_tiny_inference process for the same failed model.
Fixes#3132
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.
Fixes#2808
- Moved `fastembed` and `onnxruntime-node` to optional peerDependencies.
- Fixed bundled installs that could not resolve `onnxruntime_binding.node`.
- Added shared `runtime-install` utilities for on-demand module resolution.
- Added tests for runtime-resolution parsing and exact peer-version checks.
- add discovery of `TITLE_SYSTEM.md` and pass it through interactive startup context
- route custom title prompts to online and local tiny title generators via protocol
- update session-title docs and changelog with override behavior
- add tests for prompt discovery, forwarding, and fallback to bundled title prompts
Moved the tiny title/memory worker from a Bun Worker thread into a child process spawned via Bun.spawn IPC. The agent CLI gains a hidden --tiny-worker dispatch the parent invokes through process.execPath; the parent SIGKILLs the child on dispose so onnxruntime-node's NAPI finalizer never runs in any address space the agent owns. On Windows that finalizer was segfaulting Bun at shutdown after the tiny title model loaded (issue #1606). Drops the now-dead 'close'/'closed' handshake and the unused parentPort bootstrap, and removes tiny/worker.ts from --compile worker entries in both build scripts plus the regression test that pinned them.
Fixes#1606
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
- Fixed Anthropic stream idle-timeout errors incorrectly triggering provider retries after streaming had begun.
- Fixed darwin-x64 `bun build --compile` failure by guarding `onnxruntime-node` preload behind a `process.platform === "win32"` literal for dead-code elimination.
- Added `prefill` and `stop` parameters to the tiny-model worker's `complete` message type to pin output format without biasing content.
- Added `PI_TINY_DEVICE` env var to control ONNX execution provider (`gpu` default, `cpu`, `metal`, `cuda`, `dml`, `coreml`).
- Local tiny-model inference now tries accelerated GPU provider first and retries on CPU if initialization fails.
- Added `device.ts` module with normalization, preference resolution, and load-order helpers.
- Updated docs and model descriptions to drop CPU-specific language.
- Added a new `providers.memoryModel` setting with tiny memory model options and `ONLINE_MEMORY_MODEL_KEY` default in settings.
- Updated Mnemosyne provider resolution so a configured local tiny model overrode remote completion and used new memory extraction and consolidation prompts.
- Expanded the tiny-model CLI registry to download and report all local tiny models (title plus memory) through a unified list.
- Added Mnemosyne runtime `extractionPrompt` and `consolidationPrompt` options and wired them into resolved LLM config.
- Added fact-extraction branch to call configured completion first (temp 0), then parse facts and safely fall back.
- Added tiny local model registry features for memory/title, including keys, specs, and validation helpers.
- Added `complete` protocol messages and abort-aware client/worker paths for local title completion generation.
- Added local-models documentation for tiny/memory transformer paths, defaults, and known parser caveats.
- Updated hashline streaming preview tests to generate snapshot-tagged section headers and use an in-memory snapshot store.
- Replaced untagged file markers in multi-section preview inputs with `formatHashlineHeader` values derived from recorded file contents.
- Changed the bash command error test to expect a returned `isError` result with exit code 1 instead of a rejected promise.
- Added tiny-title protocol contracts, including progress-state unions, message payloads, and transport interfaces.
- Added title text utilities to truncate long inputs, wrap `<user-message>` blocks, and normalize generated titles.
- Added tiny-title model registry and helpers with type-safe keys and runtime optional loading via optionalDependencies.
- Added client-side worker orchestration with spawn fallback, request queueing, progress/error routing, and smoke-test APIs.
- Added worker runtime for model resolution, prompt-based inference, lock-based install retries, and close-time cache clear.