fix(mnemopi): capped embed() input to keep retention transcripts under the embedding context

`MnemopiSessionState.retainMessages` (`packages/coding-agent/src/mnemopi/state.ts:352`)
always calls `prepareRetentionTranscript(messages, true)` and hands the whole
multi-turn transcript to `embed([transcript])`. Long sessions (especially CJK
content) routinely outgrow the embedding model's context window, and llama.cpp's
`/embeddings` server rejects oversized requests with
`request (N tokens) exceeds the available context size` — every retain after
that point silently lost its vector row, leaving recall on FTS-only.

Capped per-input length inside `embed()` (the single chokepoint every retain /
query / consolidate flow funnels through) at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`
(default 32000 chars ≈ 8k English tokens / 16–32k CJK tokens, override via env
or `embeddings.maxInputChars` runtime option; `0` disables). The new array
is allocated only when at least one input is oversized, so the typical short-
query path through `embedQuery` still passes the original array through; the
truncation also emits a debug-or-warn log so the resize is no longer silent.

Fixes #3126
This commit is contained in:
roboomp
2026-06-20 11:58:56 +00:00
committed by can1357
parent c6c32242e9
commit 6666dec657
7 changed files with 185 additions and 1 deletions
+4
View File
@@ -24,6 +24,10 @@
- Fixed `/join` failing with `timed out waiting for the host's welcome` on collab sessions whose existing transcript was more than a few MB. The host now sends a small `welcome` frame (header + state + agents + `entryCount`) followed by a train of `snapshot-chunk` frames (`SNAPSHOT_CHUNK_BYTES = 512 KB`), and the guest accumulates them under a per-chunk progress timeout that resets on each chunk arrival. The first welcome lands well under one second on the default relay, so the guest's 30s first-welcome budget is no longer spent transferring the snapshot. Requires the new `COLLAB_PROTO = 2` on both sides; older hosts/guests are rejected with the existing protocol-mismatch error. ([#3144](https://github.com/can1357/oh-my-pi/issues/3144))
### Fixed
- Stopped Mnemopi retention from overflowing the embedding model's context window: `embed()` now caps each input at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 32000 chars) so a long multi-turn `MnemopiSessionState.retainMessages` transcript can't make llama.cpp's `/embeddings` server reject the request with `request (N tokens) exceeds the available context size` and silently drop vector recall for that memory ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
## [16.1.7] - 2026-06-20
### Fixed
+4
View File
@@ -2,6 +2,10 @@
## [Unreleased]
### Fixed
- Capped per-input length in `embed()` at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 32000 chars, override via the env var or `embeddings.maxInputChars` runtime option; `0` disables) so a long retention transcript can no longer overflow the embedding model's context window. llama.cpp's `/embeddings` server used to reject the request with `request (N tokens) exceeds the available context size`, silently dropping vector recall for that memory ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
## [16.1.3] - 2026-06-19
### Added
+20
View File
@@ -99,6 +99,26 @@ export function embeddingsDisabled(env: Env = process.env): boolean {
return envString("MNEMOPI_NO_EMBEDDINGS", "", env) !== "";
}
/**
* Per-input character cap applied inside `embed()` before any provider sees the text.
*
* Long retention transcripts (full multi-turn session windows) routinely outgrow
* embedding model context windows: BGE/E5 defaults are 512 tokens, bge-m3 is
* 8192, and OpenAI's text-embedding-3-* is 8192. llama.cpp's `/embeddings`
* server rejects oversized requests with `request (N tokens) exceeds the
* available context size`; OpenAI silently right-truncates. Capping at the
* source gives both backends deterministic behavior and prevents the silent
* recall degradation we saw in issue #3126.
*
* Default `32000` ≈ 8k English tokens (4 chars/token) or ~16k–32k CJK tokens —
* fits 8k-context models (bge-m3, text-embedding-3) for English and most CJK
* content. Lower it for 512-token models like `BAAI/bge-small-en-v1.5`, raise
* it for larger contexts. `0` disables the cap.
*/
export function embeddingMaxInputChars(env: Env = process.env): number {
return Math.max(0, envInt("MNEMOPI_EMBEDDING_MAX_INPUT_CHARS", 32000, env));
}
export function isApiEmbeddingModel(model = embeddingModel(), env: Env = process.env): boolean {
if (model.startsWith("openai/") || model.includes("text-embedding") || model.startsWith("text-embedding"))
return true;
+49
View File
@@ -120,6 +120,54 @@ export function embeddingsDisabled(): boolean {
return $flag("MNEMOPI_NO_EMBEDDINGS");
}
/**
* Resolved per-input character cap for {@link embed}.
*
* Reads (in order): the active runtime scope's `embeddings.maxInputChars`, then
* `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`, then the bundled `32000` default. `0`
* disables the cap entirely.
*/
function effectiveMaxInputChars(): number {
const override = activeEmbeddingOptions()?.maxInputChars;
if (override !== undefined) return Math.max(0, Math.trunc(override));
const envValue = Number.parseInt($env.MNEMOPI_EMBEDDING_MAX_INPUT_CHARS ?? "", 10);
if (Number.isFinite(envValue) && envValue >= 0) return envValue;
return 32000;
}
/**
* Right-truncate every input to {@link effectiveMaxInputChars} so a runaway
* retention transcript can't blow past the embedding model's context window.
* Returns the original array when no input needs trimming (the common case);
* the new array is allocated only when at least one input is oversized so
* we don't churn arrays for the typical short-query path through `embedQuery`.
* Emits one debug-or-warn log per call summarizing how many inputs were
* trimmed and by how much — silent truncation was the original bug (#3126).
*/
function capInputs(texts: readonly string[]): readonly string[] {
const max = effectiveMaxInputChars();
if (max === 0) return texts;
let trimmed: string[] | null = null;
let trimmedCount = 0;
let maxOriginalLen = 0;
for (let i = 0; i < texts.length; i++) {
const text = texts[i] ?? "";
if (text.length <= max) continue;
if (trimmed === null) trimmed = texts.slice() as string[];
trimmed[i] = text.slice(0, max);
trimmedCount++;
if (text.length > maxOriginalLen) maxOriginalLen = text.length;
}
if (trimmed === null) return texts;
logger[mnemopiDebugEnabled() ? "warn" : "debug"]("mnemopi: embedding input truncated", {
inputCount: texts.length,
trimmedCount,
maxOriginalLen,
maxInputChars: max,
});
return trimmed;
}
function embeddingApiKey(): ApiKey {
const active = activeEmbeddingOptions();
if (active?.apiKey !== undefined) {
@@ -408,6 +456,7 @@ export async function embed(texts: readonly string[]): Promise<EmbeddingMatrix |
if (texts.length === 0 || embeddingsDisabled()) {
return null;
}
texts = capInputs(texts);
const activeProvider = resolveEmbeddingProvider(activeEmbeddingOptions()?.provider);
if (activeProvider !== undefined) {
try {
+4 -1
View File
@@ -161,19 +161,22 @@ function resolveRuntimeOptions(options: MnemopiOptions): ResolvedMnemopiRuntimeO
const embeddingApiUrl = options.embeddingApiUrl ?? nestedEmbeddings?.apiUrl;
const embeddingApiKey = options.embeddingApiKey ?? nestedEmbeddings?.apiKey;
const embeddingProvider = resolveEmbeddingProvider(nestedEmbeddings?.provider);
const embeddingMaxInputChars = nestedEmbeddings?.maxInputChars;
const embeddings =
embeddingDisabled !== undefined ||
embeddingModel !== undefined ||
embeddingApiUrl !== undefined ||
embeddingApiKey !== undefined ||
embeddingProvider !== undefined
embeddingProvider !== undefined ||
embeddingMaxInputChars !== undefined
? {
disabled: embeddingDisabled,
model: embeddingModel,
apiUrl: embeddingApiUrl,
apiKey: embeddingApiKey,
provider: embeddingProvider,
maxInputChars: embeddingMaxInputChars,
}
: undefined;
@@ -33,6 +33,8 @@ export interface MnemopiEmbeddingRuntimeOptions {
apiUrl?: string;
apiKey?: ApiKey;
provider?: MnemopiEmbeddingProvider | ((texts: readonly string[]) => EmbeddingOutput | Promise<EmbeddingOutput>);
/** Override `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`. `0` disables the cap. See `config.embeddingMaxInputChars`. */
maxInputChars?: number;
}
export interface MnemopiLlmRuntimeOptions {
@@ -61,6 +63,7 @@ export interface ResolvedMnemopiEmbeddingRuntimeOptions {
apiUrl?: string;
apiKey?: ApiKey;
provider?: MnemopiEmbeddingProvider;
maxInputChars?: number;
}
export interface ResolvedMnemopiLlmRuntimeOptions {
@@ -0,0 +1,101 @@
import { afterEach, describe, expect, it } from "bun:test";
import "./setup";
import {
embed,
resetEmbeddingProviderForTests,
setEmbeddingProviderForTests,
} from "@oh-my-pi/pi-mnemopi/core/embeddings";
import { withMnemopiRuntimeOptions } from "@oh-my-pi/pi-mnemopi/core/runtime-options";
/**
* Regression coverage for issue #3126: `MnemopiSessionState.retainMessages`
* passes the entire multi-turn transcript to `embed()`. Long sessions used to
* overflow whatever ctx the embedding server was started with — llama.cpp
* rejects oversized requests with `request (N tokens) exceeds the available
* context size`, OpenAI silently right-truncates. `embed()` now caps each
* input to `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 32000) before the
* provider sees it.
*/
function captureProvider(): {
embed: (texts: readonly string[]) => AsyncGenerator<number[][]>;
calls: string[][];
} {
const calls: string[][] = [];
return {
calls,
async *embed(texts) {
calls.push([...texts]);
yield texts.map(text => [text.length]);
},
};
}
const ENV_KEY = "MNEMOPI_EMBEDDING_MAX_INPUT_CHARS";
function withEnvValue<T>(value: string | undefined, fn: () => Promise<T>): Promise<T> {
const previous = process.env[ENV_KEY];
if (value === undefined) delete process.env[ENV_KEY];
else process.env[ENV_KEY] = value;
return fn().finally(() => {
if (previous === undefined) delete process.env[ENV_KEY];
else process.env[ENV_KEY] = previous;
});
}
afterEach(() => {
resetEmbeddingProviderForTests();
});
describe("embed() input cap (#3126)", () => {
it("truncates oversized inputs to the default cap before reaching the provider", async () => {
const provider = captureProvider();
setEmbeddingProviderForTests(provider);
const huge = "x".repeat(40_000);
await withEnvValue(undefined, () => embed(["short", huge]));
expect(provider.calls).toHaveLength(1);
const [seenShort, seenHuge] = provider.calls[0] ?? [];
expect(seenShort).toBe("short");
expect(seenHuge?.length).toBe(32_000);
});
it("honors MNEMOPI_EMBEDDING_MAX_INPUT_CHARS env override", async () => {
const provider = captureProvider();
setEmbeddingProviderForTests(provider);
await withEnvValue("1024", () => embed(["y".repeat(5000)]));
expect(provider.calls[0]?.[0]?.length).toBe(1024);
});
it("disables the cap when the env override is 0", async () => {
const provider = captureProvider();
setEmbeddingProviderForTests(provider);
const huge = "z".repeat(50_000);
await withEnvValue("0", () => embed([huge]));
expect(provider.calls[0]?.[0]).toBe(huge);
});
it("respects a constructor-scoped maxInputChars override", async () => {
const provider = captureProvider();
setEmbeddingProviderForTests(provider);
await withEnvValue(undefined, () =>
withMnemopiRuntimeOptions({ embeddings: { maxInputChars: 256 } }, () => embed(["w".repeat(10_000)])),
);
expect(provider.calls[0]?.[0]?.length).toBe(256);
});
it("returns the original array reference when no input needs trimming", async () => {
const provider = captureProvider();
setEmbeddingProviderForTests(provider);
await withEnvValue("1024", () => embed(["fits", "still fits"]));
expect(provider.calls[0]).toEqual(["fits", "still fits"]);
});
});