StopOnTextCriteria decoded the last STOP_DECODE_WINDOW_TOKENS of the whole
sequence, so prompt tokens were eligible for matching. A prompt that itself
contains the stop string stops generation at the first generated token and
yields an empty title.
Anchor the window to the generation boundary by recording the first
generated index per batch entry. Existing local title models are
unaffected: with the assistant-prefill prompt shape, the example `</title>`
tags sit outside the 32-token window for normal messages, so no shipping
model changes behavior. The bug becomes reachable with any chat-level
few-shot prompt that places the stop string near the generation boundary.
The memory-extraction prompt concatenated its instructions, few-shot
examples, and the user message into a single user turn, so a small local
model could not distinguish instructions from input and frequently echoed
the Globex/weather examples instead of extracting facts.
Send the instructions as a real system turn and the raw text as the user
turn. The tiny worker protocol gains a systemPrompt field, and Mnemopi
completion input carries task metadata so the backend selects the right
prompt per call.
Drop the code-built MEMORY_EXTRACTION_TEMPLATE rather than porting it:
prompt text belongs in .md files, and resolveMemoryCompletionInput already
overrides that template for every extraction call, so Mnemopi rendered it
only for the result to be discarded.
Measured on ONNX q4 CPU, LFM2.5-1.2B memory extraction improved from 1/8
to 5/8 once the roles were separated.
The first interactive submit fired session.generateTitle() before
startPendingSubmission() painted the optimistic user row, and
tinyTitleClient.generate() spawned the local tiny-title subprocess
synchronously in #ensureWorker() before its first await. With a local
providers.tinyModel configured, subprocess-spawn latency therefore landed
ahead of the first frame, stalling the first prompt.
Paint the pending row before starting titling, and prewarm an idle,
unref'd worker at TUI startup via a no-op ping (no model load) so the
first submit reuses a live subprocess.
Fixes#6462
- Replaced tool-based `set_title` invocation with XML-style `<title>` marker tags for session title discovery.
- Implemented robust JSON-unwrapping logic to handle and sanitize title generation outputs.
- Updated model registry in catalog with new model support, provider prefixes, and metadata adjustments.
- Synchronized system prompt documentation and test suites to reflect the new marker-based generation flow.
When the user set `providers.tinyModel` to a local key, `generateSessionTitle`
still raced local against the online `smol` path with a 10 s timeout and
silently fired the online request whenever the local worker returned `null`
(unknown key, model not downloaded, transformers.js failure). The online path
resolves the `smol` role through `priority.json` (haiku → flash → mini → …);
with an `OPENROUTER_API_KEY` picked up from env, that silently billed
OpenRouter without consent.
Drop the race entirely for local choices: honor the user's setting, log a
warning on local failure, leave the session untitled. The `raceFirstNonNull`
helper and `TITLE_LOCAL_FALLBACK_DELAY_MS` had no other consumer and are
removed; the obsolete \"silently bills online when local fails\" tests are
flipped into regressions that lock the no-fallback contract, including the
unknown-key path (e.g. \"ollama:gpt-oss\") which previously also leaked
straight through to the online billing path.
Fixes#3187
- add discovery of `TITLE_SYSTEM.md` and pass it through interactive startup context
- route custom title prompts to online and local tiny title generators via protocol
- update session-title docs and changelog with override behavior
- add tests for prompt discovery, forwarding, and fallback to bundled title prompts
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
Stopped the tiny-title subprocess from inheriting stdout and stderr so native model runtime output cannot corrupt the interactive scrollback. Added a regression test for worker stdio configuration.\n\nFixes #2206
- Updated the initial render intent to carry a `clearScrollback` flag.
- Adjusted first-frame rendering logic to preserve existing scrollback by default and clear only when requested.
- Added `resolver` stubs to coding-agent and web-search test registries for interface compatibility.
- Added persistent settings in the Providers tab for ONNX execution provider and quantization/precision, replacing env-var-only configuration.
- PI_TINY_DEVICE and PI_TINY_DTYPE env vars still override the matching setting at spawn time.
- Added tinyWorkerEnvOverlay to map settings onto worker env without clobbering explicit env vars.
- Updated docs and tests to reflect the new setting-first resolution order.
- Updated hashline streaming preview tests to generate snapshot-tagged section headers and use an in-memory snapshot store.
- Replaced untagged file markers in multi-section preview inputs with `formatHashlineHeader` values derived from recorded file contents.
- Changed the bash command error test to expect a returned `isError` result with exit code 1 instead of a rejected promise.
- Added tiny-title protocol contracts, including progress-state unions, message payloads, and transport interfaces.
- Added title text utilities to truncate long inputs, wrap `<user-message>` blocks, and normalize generated titles.
- Added tiny-title model registry and helpers with type-safe keys and runtime optional loading via optionalDependencies.
- Added client-side worker orchestration with spawn fallback, request queueing, progress/error routing, and smoke-test APIs.
- Added worker runtime for model resolution, prompt-based inference, lock-based install retries, and close-time cache clear.
- Added `get(path)` support in test `createSettings` to return tiny model setting for `providers.tinyModel`.
- Updated `createSettings` in tests to accept optional `tinyModel` with default `"online"`.
- Added `flushMicrotasks()` and helper factories for test settings, registry stubs, and online model mocks.
- Added tests for `raceFirstNonNull`, `generateSessionTitle` routing, and providers tinyModel enum validation.