StopOnTextCriteria decoded the last STOP_DECODE_WINDOW_TOKENS of the whole
sequence, so prompt tokens were eligible for matching. A prompt that itself
contains the stop string stops generation at the first generated token and
yields an empty title.
Anchor the window to the generation boundary by recording the first
generated index per batch entry. Existing local title models are
unaffected: with the assistant-prefill prompt shape, the example `</title>`
tags sit outside the 32-token window for normal messages, so no shipping
model changes behavior. The bug becomes reachable with any chat-level
few-shot prompt that places the stop string near the generation boundary.
The memory-extraction prompt concatenated its instructions, few-shot
examples, and the user message into a single user turn, so a small local
model could not distinguish instructions from input and frequently echoed
the Globex/weather examples instead of extracting facts.
Send the instructions as a real system turn and the raw text as the user
turn. The tiny worker protocol gains a systemPrompt field, and Mnemopi
completion input carries task metadata so the backend selects the right
prompt per call.
Drop the code-built MEMORY_EXTRACTION_TEMPLATE rather than porting it:
prompt text belongs in .md files, and resolveMemoryCompletionInput already
overrides that template for every extraction call, so Mnemopi rendered it
only for the result to be discarded.
Measured on ONNX q4 CPU, LFM2.5-1.2B memory extraction improved from 1/8
to 5/8 once the roles were separated.
- Set generation temperature to zero for online title generation to prevent garbled names.
- Update system prompt to instruct exact copying of technical terms and names.
- Reject generated titles containing no word characters to prevent punctuation-only sessions.
normalizeGeneratedTitle only guarded emptiness and the none sentinel, so
when the tiny title model ignored the titling task and answered the first
user message, its full one-line reply became the session title verbatim.
Bound accepted titles to 80 chars / 12 words and return null past that,
deferring titling to the next user turn. Both the online and local-worker
paths funnel through this normalizer, so both are covered.
Fixes#7303
The first interactive submit fired session.generateTitle() before
startPendingSubmission() painted the optimistic user row, and
tinyTitleClient.generate() spawned the local tiny-title subprocess
synchronously in #ensureWorker() before its first await. With a local
providers.tinyModel configured, subprocess-spawn latency therefore landed
ahead of the first frame, stalling the first prompt.
Paint the pending row before starting titling, and prewarm an idle,
unref'd worker at TUI startup via a no-op ping (no model load) so the
first submit reuses a live subprocess.
Fixes#6462
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Recognized preformatted chat context in message-preproc and bypassed paired-tag stripping that consumed the entire envelope.
- Integrated scaffolding-tag removal in low-signal title prefilter.
- Added corresponding tests for formatTitleUserMessage and isLowSignalTitleInput.
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
Downloaded missing ONNX Runtime CUDA provider sidecars when the compiled tiny-model side runtime is used with PI_TINY_DEVICE=cuda.
Preserved actionable CUDA worker diagnostics in tiny-models text output and added focused regression coverage.
Fixes#4475
isAllCapsWord matched any multi-letter token without a lowercase letter,
so CJK tokens registered as ALL-CAPS words: two adjacent ones marked the
whole source shouty and silently disabled acronym restoration for every
non-Latin-script message (e.g. '修复 CNPG 集群故障' kept 'Cnpg'). Require an
actual uppercase letter; cased-script shout detection is unchanged.
The PR's changelog, doc comment, and prompt examples all name ETL as a
restored acronym, but the review-response narrowing (vowel heuristic +
allowlist) silently dropped it: ETL bears a vowel and was not listed.
Add it to COMMON_TITLE_ACRONYMS and pin it in the allowlist test.
Narrow acronym restoration so plain all-caps English words such as FIX
and WORK do not get restored when the model naturally capitalizes the
first title word. Restorable all-caps source tokens now need a stronger
acronym signal: a common technical acronym allowlist, digits, or a
consonant-only shape.
This keeps CNPG, ETL, JWT, SQL, and API restoration while preserving the
anti-shout behavior for single emphatic words.
Fixes#4220
reconcileTitleCasing now maps ALL-CAPS source tokens (CNPG, API, ETL,
JWT) into an acronyms table and restores them when the model produces a
plain title-cased artifact (Cnpg). Restoration is disabled when the
source is shouty (>=2 consecutive multi-letter ALL-CAPS tokens like
"FIX the BUG NOW" or "ALL ERROR HANDLING"), and lowercase model output
is left alone so isolated single-word emphasis (WORK -> work) is never
re-shouted.
The three title system prompts (title-system.md, title-system-marker.md,
tiny-title-system.md) also gained an explicit instruction to preserve
ALL-CAPS acronyms verbatim, so a competent model short-circuits via the
verbatim set before post-processing kicks in.
Fixes#4220
- Added the `llama3.2:3b` model configuration pointing to the quantized `onnx-community/Llama-3.2-3B-Instruct-ONNX` repository.
- Registered the model in both the available local models registry and list of valid memory model values.
- Documented the new option as a shipped local model choice in the documentation and changelog.
- Added a fixed-width title slot system to serialize and persist session titles across physical files and backend storage.
- Integrated automated session title updates triggered by todo replan operations using conversation history context.
- Extended storage interfaces across memory, file-system, Redis, and SQL backends to support independent title metadata updates.
- Implemented title metadata parsing within session loaders and list utilities to ensure accurate retrieval and display.
- Updated `reconcileTitleCasing` to ignore purely uppercase words when restoring source casing.
- Restricted restoration logic to mixed-case identifiers (e.g., `TinyVMM`, `iOS`) to avoid overriding the model's clean sentence case with emphatic user shouting.
- Added regression tests to ensure all-caps input does not trigger title re-shouting.
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
- Updated `normalizeGeneratedTitle` to reconcile model-generated titles against the user's input instead of forcing title-case.
- Added logic to restore distinctive proper-noun casing (e.g., `TinyVMM`) and flatten model-generated camelCase artifacts (e.g., `dAemon`) that do not appear in the user's message.
- Ensured model-cased proper nouns that are not in the source message (e.g., `GitHub`) are preserved.
Referenced the tiny-model worker while requests are pending so standalone downloads cannot exit before worker IPC resolves. Added regression coverage for the download lifetime contract.
Fixes#3291
- Established shared `subprocess` infrastructure to standardize worker IPC, environment resolution, and error handling.
- Migrated mnemopi, speech-to-text, tiny-model, and TTS clients to utilize consolidated worker-client utilities.
- Implemented common runtime helpers for ONNX model loading, logging, and process readiness probing.
- Eliminated redundant local spawn logic and environment mapping across all inference worker clients.
Stopped treating subprocess crash notifications as model-specific inference failures. Added regression coverage proving an unrelated queued title model can spawn a replacement worker after the crashed worker faults all pending requests.
Fixes#3132
Blocked the unsupported Qwen3 1.7B ONNX memory model before loading transformers and remembered local model execution failures so the client returns null instead of spawning another __omp_worker_tiny_inference process for the same failed model.
Fixes#3132
- Added `safeSend` helper wrapping `Subprocess.send()` so sync throws and async EPIPE rejections cannot escape.
- Replaced inline try/catch send wrappers in STT, TTS, and tiny-title clients with shared `safeSend`.
- Added `isIpcSendEpipe` predicate and made matching rejections non-fatal in the `unhandledRejection` handler.
- Added contract tests for `safeSend` and `isIpcSendEpipe` covering sync throws, async rejections, and edge cases.
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.
Fixes#2808
- Added `WorkerInbox` and `installWorkerInbox(port)` to queue worker messages before bind.
- Added `consumeWorkerInbox()` to replay buffered messages and clear one active inbox.
- Added buffered inbox consumption in JS and tab worker transports before direct message handlers.
- Normalized worker selector arguments to the `__omp_worker_*` naming across workers and tests.
- Applied title-casing to `normalizeGeneratedTitle` outputs using a new internal helper.
- Adjusted tiny text and title generator tests to assert the new title-cased results.
- Moved `fastembed` and `onnxruntime-node` to optional peerDependencies.
- Fixed bundled installs that could not resolve `onnxruntime_binding.node`.
- Added shared `runtime-install` utilities for on-demand module resolution.
- Added tests for runtime-resolution parsing and exact peer-version checks.
- add discovery of `TITLE_SYSTEM.md` and pass it through interactive startup context
- route custom title prompts to online and local tiny title generators via protocol
- update session-title docs and changelog with override behavior
- add tests for prompt discovery, forwarding, and fallback to bundled title prompts
- Rerouted sync, tab, js-eval, and tiny workers to re-enter CLI modes via `__omp_*` selectors.
- Adjusted `cli.ts` startup to dispatch worker entrypoints before parsing and exit 1 on uncaught errors.
- Bundled CLI as `dist/cli.js` in prepack, switching `omp` binary and published files.
- Removed explicit Bun `--compile` worker entrypoints from build/release scripts in favor of host-entry dispatch.
- Added `declareWorkerHostEntry()` and `workerHostEntry()` environment helpers and `PI_COMPILED` binary detection.
Stopped the tiny-title subprocess from inheriting stdout and stderr so native model runtime output cannot corrupt the interactive scrollback. Added a regression test for worker stdio configuration.\n\nFixes #2206
The release_binary smoke probe spawns the tiny-title worker as a cold subprocess
of the compiled binary and pings it. On the contended macos-15-intel runner the
cold start (decompress + module-graph load, with a cold bun cache) blew past the
5s bound and failed the 15.10.3 release, while arm64/linux/win passed. The probe
only needs to prove the worker spawns and ponges at all, so the timeout is raised
to 30s — a dead worker still never ponges, so the check is unchanged in substance.
Hard-killed the tiny-model subprocess after worker-reported execution errors so ONNX native allocations are released before retries.
Added a fake-worker regression for the unknown-failure path and queued local completions.
Fixes#1940
- Adjusted session and dashboard renderers to derive heights from live terminal rows.
- Propagated terminal-row callbacks through picker/controller wiring and session selectors.
- Reworked list visibility and page navigation to honor row-based budgets and footer lines.
- Added DECRQM 2031 startup probing and stopped OSC11 polling once support was confirmed.
- Skipped titling for greeting/filler/empty first messages deterministically, retrying on later user messages.
- Let capable title models decline taskless input via a "none" sentinel.
- Guarded against clobbering a name set by a concurrent attempt.
The first cut at the subprocess isolation swallowed every signal exit (`exitCode === null`) on the assumption it was the intentional SIGKILL from `terminate()`. That misclassifies real worker deaths — SIGSEGV from a native crash, SIGKILL from the OOM killer, an operator `kill -9` — so any in-flight title/completion/download promise would await forever while `#worker` still pointed at a dead process.
Added an `intentionalExit` flag flipped by `wrapSubprocess.terminate()` right before its SIGKILL. `onExit` swallows only the flagged exit; every other signal exit now fires the `errors` channel with a "signal SIGFOO" message so `TinyTitleClient.#handleWorkerError` clears `#pending` and dumps the dead worker handle. Added two regression tests pinning both branches.
Reported by chatgpt-codex-connector on #1607.
Moved the tiny title/memory worker from a Bun Worker thread into a child process spawned via Bun.spawn IPC. The agent CLI gains a hidden --tiny-worker dispatch the parent invokes through process.execPath; the parent SIGKILLs the child on dispose so onnxruntime-node's NAPI finalizer never runs in any address space the agent owns. On Windows that finalizer was segfaulting Bun at shutdown after the tiny title model loaded (issue #1606). Drops the now-dead 'close'/'closed' handshake and the unused parentPort bootstrap, and removes tiny/worker.ts from --compile worker entries in both build scripts plus the regression test that pinned them.
Fixes#1606
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
- Fixed Anthropic stream idle-timeout errors incorrectly triggering provider retries after streaming had begun.
- Fixed darwin-x64 `bun build --compile` failure by guarding `onnxruntime-node` preload behind a `process.platform === "win32"` literal for dead-code elimination.
- Added `prefill` and `stop` parameters to the tiny-model worker's `complete` message type to pin output format without biasing content.
- Changed tiny-device preference resolution to always default to CPU instead of platform-specific DirectML/CUDA heuristics.
- Updated settings schema, documentation, and changelog text to describe the CPU default while keeping accelerated providers behind explicit `providers.tinyModelDevice`/`PI_TINY_DEVICE` choices.
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.
- Added persistent settings in the Providers tab for ONNX execution provider and quantization/precision, replacing env-var-only configuration.
- PI_TINY_DEVICE and PI_TINY_DTYPE env vars still override the matching setting at spawn time.
- Added tinyWorkerEnvOverlay to map settings onto worker env without clobbering explicit env vars.
- Updated docs and tests to reflect the new setting-first resolution order.
- Added `PI_TINY_DEVICE` env var to control ONNX execution provider (`gpu` default, `cpu`, `metal`, `cuda`, `dml`, `coreml`).
- Local tiny-model inference now tries accelerated GPU provider first and retries on CPU if initialization fails.
- Added `device.ts` module with normalization, preference resolution, and load-order helpers.
- Updated docs and model descriptions to drop CPU-specific language.