fix(mnemopi): kept transcript tail when clipping oversized embedding inputs
`MnemopiSessionState.retainMessages` hands `embed()` the chronological multi-turn transcript (oldest -> newest). A naive `slice(0, max)` cap landed on the oldest turns and dropped the most recent content, so every retained episode past the cap collapsed onto essentially the same prefix vector and dense recall could not match topics introduced after the first 8192 chars. `capInputs` now routes oversized inputs through `clipToWindow`, which keeps roughly half the cap from the head and half from the tail with a small `[...]` elision marker between them. Short inputs and array reference pass- through are unchanged. Falls back to a tail-only clip when `max` is too small to fit a useful split. Locked in by a new `embedding-input-cap.test.ts` case that pins markers at both ends of a 50k transcript and asserts both survive the clip. Fixes #3126
This commit is contained in:
@@ -26,7 +26,7 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Stopped Mnemopi retention from overflowing the embedding model's context window: `embed()` now caps each input at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 8192 chars) so a long multi-turn `MnemopiSessionState.retainMessages` transcript can't make llama.cpp's `/embeddings` server reject the request with `request (N tokens) exceeds the available context size` and silently drop vector recall for that memory ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
|
||||
- Stopped Mnemopi retention from overflowing the embedding model's context window: `embed()` now caps each input at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 8192 chars) and clips with a head/tail split so a long multi-turn `MnemopiSessionState.retainMessages` transcript can't make llama.cpp's `/embeddings` server reject the request with `request (N tokens) exceeds the available context size` and silently drop vector recall for that memory. The head/tail clip keeps both the opening setup and the most recent turns so later episodes don't collapse onto the same prefix vector ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
|
||||
|
||||
## [16.1.7] - 2026-06-20
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Capped per-input length in `embed()` at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 8192 chars, override via the env var or `embeddings.maxInputChars` runtime option; `0` disables) so a long retention transcript can no longer overflow the embedding model's context window. llama.cpp's `/embeddings` server used to reject the request with `request (N tokens) exceeds the available context size`, silently dropping vector recall for that memory ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
|
||||
- Capped per-input length in `embed()` at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 8192 chars, override via the env var or `embeddings.maxInputChars` runtime option; `0` disables) so a long retention transcript can no longer overflow the embedding model's context window. Oversized inputs are clipped with a head/tail split so chronological transcripts keep both the opening setup and the most recent turns instead of losing the latest content under a naive prefix slice. llama.cpp's `/embeddings` server used to reject the request with `request (N tokens) exceeds the available context size`, silently dropping vector recall for that memory ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
|
||||
|
||||
## [16.1.3] - 2026-06-19
|
||||
|
||||
|
||||
@@ -135,14 +135,39 @@ function effectiveMaxInputChars(): number {
|
||||
return 8192;
|
||||
}
|
||||
|
||||
/** Elision marker injected between the retained head and tail of an oversized input. */
|
||||
const EMBEDDING_ELISION_MARKER = "\n\n[...]\n\n";
|
||||
|
||||
/**
|
||||
* Right-truncate every input to {@link effectiveMaxInputChars} so a runaway
|
||||
* retention transcript can't blow past the embedding model's context window.
|
||||
* Returns the original array when no input needs trimming (the common case);
|
||||
* the new array is allocated only when at least one input is oversized so
|
||||
* we don't churn arrays for the typical short-query path through `embedQuery`.
|
||||
* Emits one debug-or-warn log per call summarizing how many inputs were
|
||||
* trimmed and by how much — silent truncation was the original bug (#3126).
|
||||
* Right-clip a single oversized input to {@link max} chars while preserving
|
||||
* both ends. Retention transcripts are chronological (oldest → newest), so a
|
||||
* naive `slice(0, max)` would drop the most recent — and most semantically
|
||||
* loaded — turns once a session passed the cap, leaving every later retained
|
||||
* episode with essentially the same prefix vector. Keeping a head/tail split
|
||||
* lets the embedding capture the topic setup at the start AND the latest
|
||||
* exchanges at the end. Falls back to a tail-only clip when `max` is too
|
||||
* small to fit the elision marker plus a useful slice on either side.
|
||||
*/
|
||||
function clipToWindow(text: string, max: number): string {
|
||||
if (text.length <= max) return text;
|
||||
if (max <= EMBEDDING_ELISION_MARKER.length + 16) return text.slice(text.length - max);
|
||||
const budget = max - EMBEDDING_ELISION_MARKER.length;
|
||||
const headLen = budget >>> 1;
|
||||
const tailLen = budget - headLen;
|
||||
return text.slice(0, headLen) + EMBEDDING_ELISION_MARKER + text.slice(text.length - tailLen);
|
||||
}
|
||||
|
||||
/**
|
||||
* Clip every input to {@link effectiveMaxInputChars} so a runaway retention
|
||||
* transcript can't blow past the embedding model's context window. Uses a
|
||||
* head/tail split via {@link clipToWindow} so the embedding still sees the
|
||||
* tail of the conversation (where the latest topic shifts live) and not just
|
||||
* the stale prefix. Returns the original array when no input needs trimming
|
||||
* (the common case); the new array is allocated only when at least one input
|
||||
* is oversized so we don't churn arrays for the typical short-query path
|
||||
* through `embedQuery`. Emits one debug-or-warn log per call summarizing how
|
||||
* many inputs were trimmed and by how much — silent truncation was the
|
||||
* original bug (#3126).
|
||||
*/
|
||||
function capInputs(texts: readonly string[]): readonly string[] {
|
||||
const max = effectiveMaxInputChars();
|
||||
@@ -154,7 +179,7 @@ function capInputs(texts: readonly string[]): readonly string[] {
|
||||
const text = texts[i] ?? "";
|
||||
if (text.length <= max) continue;
|
||||
if (trimmed === null) trimmed = texts.slice() as string[];
|
||||
trimmed[i] = text.slice(0, max);
|
||||
trimmed[i] = clipToWindow(text, max);
|
||||
trimmedCount++;
|
||||
if (text.length > maxOriginalLen) maxOriginalLen = text.length;
|
||||
}
|
||||
|
||||
@@ -90,6 +90,27 @@ describe("embed() input cap (#3126)", () => {
|
||||
expect(provider.calls[0]?.[0]?.length).toBe(256);
|
||||
});
|
||||
|
||||
it("preserves both ends of a chronological transcript via the head/tail clip", async () => {
|
||||
const provider = captureProvider();
|
||||
setEmbeddingProviderForTests(provider);
|
||||
|
||||
// `MnemopiSessionState.retainMessages` hands `embed()` the full
|
||||
// chronological transcript. Before the head/tail clip, a `slice(0, max)`
|
||||
// would land on the oldest turns and drop the most recent (and most
|
||||
// semantically loaded) content. Verify both ends survive.
|
||||
const earliest = "OPENING_TURN_MARKER";
|
||||
const latest = "FINAL_TURN_MARKER";
|
||||
const transcript = `${earliest}${"x".repeat(50_000)}${latest}`;
|
||||
|
||||
await withEnvValue(undefined, () => embed([transcript]));
|
||||
|
||||
const seen = provider.calls[0]?.[0] ?? "";
|
||||
expect(seen.length).toBe(8192);
|
||||
expect(seen.startsWith(earliest)).toBe(true);
|
||||
expect(seen.endsWith(latest)).toBe(true);
|
||||
expect(seen).toContain("[...]");
|
||||
});
|
||||
|
||||
it("returns the original array reference when no input needs trimming", async () => {
|
||||
const provider = captureProvider();
|
||||
setEmbeddingProviderForTests(provider);
|
||||
|
||||
Reference in New Issue
Block a user