The prompt-inventory test sliced on the '# Inventory'/'ENV' markers from the
PR's merge-base prompt.md. On current main (chore: prompt reorder) the heading
is '# Tool Inventory' and the 'ENV' marker is gone, which is why the file was
deleted there. Accept either layout so the resurrected tests (incl. the SDK
'render provided tools' contract) pass after the cherry-pick.
fork() reset mnemopi conversation tracking directly but skipped the shared new-transcript reset, so the folded/promoted first-turn memory stayed in #baseSystemPrompt. The next turn re-recalled and the change-detection saw no diff, taking the fallback promotion path and injecting the <memories> block twice into the forked session prompt. Route fork() through #resetMemoryContextForNewTranscript() like the other reset paths and add a regression test asserting the forked prompt contains recalled memory exactly once.
The PR placed the entry under the released [16.0.7] section with a
duplicate ### Fixed header, mutating released notes. Move it under
[Unreleased] per the repo changelog convention.
The #3099 test omitted the startInAllScope flag, so it passed against the
buggy baseline too (empty folder defaults to folder scope regardless). Force
the removed flag via a cast to pin the real contract: even when a caller asks
for all-projects scope on an empty folder, the picker must stay folder-scoped.
Proven: passes on head src, fails on baseline src (renders '(all projects)').
Move the legacy pi-ai compat note from the released [16.0.5] section to
[Unreleased], and drop the inaccurate bare-`typebox` claim: bare
`typebox` specifier handling already exists on main (issue #2858,
TYPEBOX_SPECIFIER_FILTER / TYPEBOX_IMPORT_SPECIFIER_REGEX) and this PR's
diff does not touch it. The entry now scopes the change to the actual
delta: restored getModel/getModels aliases and StringEnum enum-object
support.
`MnemopiSessionState.retainMessages` hands `embed()` the chronological
multi-turn transcript (oldest -> newest). A naive `slice(0, max)` cap landed
on the oldest turns and dropped the most recent content, so every retained
episode past the cap collapsed onto essentially the same prefix vector and
dense recall could not match topics introduced after the first 8192 chars.
`capInputs` now routes oversized inputs through `clipToWindow`, which keeps
roughly half the cap from the head and half from the tail with a small
`[...]` elision marker between them. Short inputs and array reference pass-
through are unchanged. Falls back to a tail-only clip when `max` is too small
to fit a useful split. Locked in by a new `embedding-input-cap.test.ts` case
that pins markers at both ends of a 50k transcript and asserts both survive
the clip.
Fixes#3126
Defaulted MNEMOPI_EMBEDDING_MAX_INPUT_CHARS to 8192 so the automatic guard matches bge-m3 and OpenAI text-embedding context limits by default. Larger local embedding servers such as Qwen3-Embedding with 32k ctx can still raise the cap, and 0 still disables truncation.
Fixes#3126
`MnemopiSessionState.retainMessages` (`packages/coding-agent/src/mnemopi/state.ts:352`)
always calls `prepareRetentionTranscript(messages, true)` and hands the whole
multi-turn transcript to `embed([transcript])`. Long sessions (especially CJK
content) routinely outgrow the embedding model's context window, and llama.cpp's
`/embeddings` server rejects oversized requests with
`request (N tokens) exceeds the available context size` — every retain after
that point silently lost its vector row, leaving recall on FTS-only.
Capped per-input length inside `embed()` (the single chokepoint every retain /
query / consolidate flow funnels through) at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`
(default 32000 chars ≈ 8k English tokens / 16–32k CJK tokens, override via env
or `embeddings.maxInputChars` runtime option; `0` disables). The new array
is allocated only when at least one input is oversized, so the typical short-
query path through `embedQuery` still passes the original array through; the
truncation also emits a debug-or-warn log so the resize is no longer silent.
Fixes#3126
The chunked-welcome refactor moved welcome-timer arming into socket.onOpen,
leaving the connect phase uncovered: if the relay blackholes the WebSocket
handshake (no onOpen and no onClose), the timer never arms and /join hangs
forever. Baseline armed the 30s timeout right after connect(); restore that
so a stalled handshake still rejects the join. onOpen continues to re-arm
(resetting the budget) once the socket opens.
resolveModels("all") expanded the full TINY_LOCAL_MODELS registry, which now
includes the qwen3-1.7b entry marked unsupportedReason. loadPipeline() throws
for such specs, so the download worker reported it as failed and the bulk
command exited with "One or more tiny title models failed to download" even
when every usable model downloaded. Filter unsupported specs out of the `all`
prefetch path; explicit single-model requests are unchanged. Addresses the
unaddressed Codex P2 on PR #3133.
Run threshold compaction maintenance when an active goal turn ends through a successful yield, while preserving the final-yield skip for non-goal completions.
Fixes#3146
The chunked welcome path cleared its snapshot progress timer before
writing the replica file and switching sessions. If that apply work
failed, the frame-apply catch only logged the error, leaving the
initial join promise pending with no welcome/progress timer left to
settle it.
Reject the pending initial join when a welcome or snapshot-chunk apply
fails before the join has completed, preserving reconnect-time logging
for already-joined guests. Add a regression test that forces the replica
write to fail and asserts /join rejects instead of hanging.
Fixes#3144
The host used to ship the entire transcript inside a single welcome
frame, so a multi-MB session spent the guest's 30s first-welcome
timeout on the relay transfer itself: ~1.3 MB took ~3s, ~4.2 MB took
~12s, and ~13.6 MB never arrived before the guest gave up with
'timed out waiting for the host's welcome'.
Bump COLLAB_PROTO to 2 and split the welcome:
- welcome carries metadata only (header, state, agents, entryCount,
readOnly) and lands in well under one second.
- a train of snapshot-chunk frames (SNAPSHOT_CHUNK_BYTES = 512 KB,
oversize entries ship alone) carries the transcript. Last chunk
flips final: true; an empty snapshot still emits one final chunk.
- the host queues welcome + chunks synchronously inside #handleHello,
preserving the host comment's ordering invariant (later broadcast
frames cannot interleave between them).
- the TUI guest accumulates chunks under a SNAPSHOT_PROGRESS_TIMEOUT_MS
that resets per chunk; only after final does it write the replica
jsonl, switchSession, and render. The first-welcome timeout still
guards arrival of the small welcome.
- the collab-web GuestClient streams entries into the snapshot as
chunks arrive and flips phase to 'live' on final.
Includes a contract test (in-process relay) asserting the welcome is
metadata-only, the chunk train fans the 1.5 MB synthetic transcript
across multiple frames with only the last marked final, and the
flattened entries match the source snapshot.
Fixes#3144
Stopped treating subprocess crash notifications as model-specific inference failures. Added regression coverage proving an unrelated queued title model can spawn a replacement worker after the crashed worker faults all pending requests.
Fixes#3132
Blocked the unsupported Qwen3 1.7B ONNX memory model before loading transformers and remembered local model execution failures so the client returns null instead of spawning another __omp_worker_tiny_inference process for the same failed model.
Fixes#3132
Stopped routing internal URL directory reads through the filesystem tree renderer so vault:// and '/data/workspaces/can1357__oh-my-pi__3116/.omp-session/2026-06-20T09-33-10-396Z_019ee460-817c-7000-8139-f8e2927809dd/local' keep their custom navigable listings.
Allow file-backed internal URL handlers to return existing directories as resources so read can list them and search/find can walk their source paths.
Fixes#3116
Restored and refreshed promoted memory prompts through the shared new-transcript reset path used by new sessions, handoff, branch, /btw, and cross-session switches.\n\nAdded coverage for newSession after first-turn memory recall so stale recalled memories cannot leak into the next transcript when recall returns no context.\n\nFixes #3111
Reset memory recall state before rebuilding the prompt during session switches, and clear fallback-promoted memory prompts when moving to another session.\n\nAdded a regression test that switches sessions after first-turn memory recall and verifies the next session does not receive stale memories.\n\nFixes #3111
Promoted first-turn memory recall into the stable base prompt so append-only sessions do not drop the memory block on the next turn and rebuild the provider prefix.\n\nAdded a regression test covering a memory backend that recalls once before the first model request.\n\nFixes #3111
The ASCII table renderer in `sqlite-reader.ts` shrank columns down to
`MIN_COLUMN_WIDTH=1` to fit the 120-cell budget. With ~20+ columns
(the reporter had 33) every multi-char cell collapsed to a lone `…`
and the final per-line `truncateToWidth(..., MAX_RENDER_WIDTH)` then
chopped the right edge — so the read tool returned a table of nothing
but ellipses with the rightmost cells missing entirely.
Bump the per-column floor to 3 (so cells always show at least two real
glyphs alongside the ellipsis) and, when the column count alone forces
the floor over budget, fall back to a per-row vertical block layout —
mirroring `psql`'s expanded display mode. Each row becomes a
`column: value` group with column names padded so colons align and
the value line truncated to the same 120-cell budget.
Fixes#3107
- Updated `buildOpenAICompat` to override the `qwen` thinking format for Fireworks-hosted models, ensuring they use `openai` thinking parameters instead.
- Prevented invalid `enable_thinking` payload errors by ensuring Fireworks-hosted Qwen requests conform to their strict schema.
- Updated `AgentSession` to allow Fireworks fast-fallback logic to execute even when standard retries are disabled.