fix(mnemopi): capped embed() input to keep retention transcripts under the embedding context
`MnemopiSessionState.retainMessages` (`packages/coding-agent/src/mnemopi/state.ts:352`) always calls `prepareRetentionTranscript(messages, true)` and hands the whole multi-turn transcript to `embed([transcript])`. Long sessions (especially CJK content) routinely outgrow the embedding model's context window, and llama.cpp's `/embeddings` server rejects oversized requests with `request (N tokens) exceeds the available context size` — every retain after that point silently lost its vector row, leaving recall on FTS-only. Capped per-input length inside `embed()` (the single chokepoint every retain / query / consolidate flow funnels through) at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 32000 chars ≈ 8k English tokens / 16–32k CJK tokens, override via env or `embeddings.maxInputChars` runtime option; `0` disables). The new array is allocated only when at least one input is oversized, so the typical short- query path through `embedQuery` still passes the original array through; the truncation also emits a debug-or-warn log so the resize is no longer silent. Fixes #3126
This commit is contained in:
@@ -24,6 +24,10 @@
|
||||
|
||||
- Fixed `/join` failing with `timed out waiting for the host's welcome` on collab sessions whose existing transcript was more than a few MB. The host now sends a small `welcome` frame (header + state + agents + `entryCount`) followed by a train of `snapshot-chunk` frames (`SNAPSHOT_CHUNK_BYTES = 512 KB`), and the guest accumulates them under a per-chunk progress timeout that resets on each chunk arrival. The first welcome lands well under one second on the default relay, so the guest's 30s first-welcome budget is no longer spent transferring the snapshot. Requires the new `COLLAB_PROTO = 2` on both sides; older hosts/guests are rejected with the existing protocol-mismatch error. ([#3144](https://github.com/can1357/oh-my-pi/issues/3144))
|
||||
|
||||
### Fixed
|
||||
|
||||
- Stopped Mnemopi retention from overflowing the embedding model's context window: `embed()` now caps each input at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 32000 chars) so a long multi-turn `MnemopiSessionState.retainMessages` transcript can't make llama.cpp's `/embeddings` server reject the request with `request (N tokens) exceeds the available context size` and silently drop vector recall for that memory ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
|
||||
|
||||
## [16.1.7] - 2026-06-20
|
||||
|
||||
### Fixed
|
||||
|
||||
@@ -2,6 +2,10 @@
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### Fixed
|
||||
|
||||
- Capped per-input length in `embed()` at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 32000 chars, override via the env var or `embeddings.maxInputChars` runtime option; `0` disables) so a long retention transcript can no longer overflow the embedding model's context window. llama.cpp's `/embeddings` server used to reject the request with `request (N tokens) exceeds the available context size`, silently dropping vector recall for that memory ([#3126](https://github.com/can1357/oh-my-pi/issues/3126)).
|
||||
|
||||
## [16.1.3] - 2026-06-19
|
||||
|
||||
### Added
|
||||
|
||||
@@ -99,6 +99,26 @@ export function embeddingsDisabled(env: Env = process.env): boolean {
|
||||
return envString("MNEMOPI_NO_EMBEDDINGS", "", env) !== "";
|
||||
}
|
||||
|
||||
/**
|
||||
* Per-input character cap applied inside `embed()` before any provider sees the text.
|
||||
*
|
||||
* Long retention transcripts (full multi-turn session windows) routinely outgrow
|
||||
* embedding model context windows: BGE/E5 defaults are 512 tokens, bge-m3 is
|
||||
* 8192, and OpenAI's text-embedding-3-* is 8192. llama.cpp's `/embeddings`
|
||||
* server rejects oversized requests with `request (N tokens) exceeds the
|
||||
* available context size`; OpenAI silently right-truncates. Capping at the
|
||||
* source gives both backends deterministic behavior and prevents the silent
|
||||
* recall degradation we saw in issue #3126.
|
||||
*
|
||||
* Default `32000` ≈ 8k English tokens (4 chars/token) or ~16k–32k CJK tokens —
|
||||
* fits 8k-context models (bge-m3, text-embedding-3) for English and most CJK
|
||||
* content. Lower it for 512-token models like `BAAI/bge-small-en-v1.5`, raise
|
||||
* it for larger contexts. `0` disables the cap.
|
||||
*/
|
||||
export function embeddingMaxInputChars(env: Env = process.env): number {
|
||||
return Math.max(0, envInt("MNEMOPI_EMBEDDING_MAX_INPUT_CHARS", 32000, env));
|
||||
}
|
||||
|
||||
export function isApiEmbeddingModel(model = embeddingModel(), env: Env = process.env): boolean {
|
||||
if (model.startsWith("openai/") || model.includes("text-embedding") || model.startsWith("text-embedding"))
|
||||
return true;
|
||||
|
||||
@@ -120,6 +120,54 @@ export function embeddingsDisabled(): boolean {
|
||||
return $flag("MNEMOPI_NO_EMBEDDINGS");
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolved per-input character cap for {@link embed}.
|
||||
*
|
||||
* Reads (in order): the active runtime scope's `embeddings.maxInputChars`, then
|
||||
* `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`, then the bundled `32000` default. `0`
|
||||
* disables the cap entirely.
|
||||
*/
|
||||
function effectiveMaxInputChars(): number {
|
||||
const override = activeEmbeddingOptions()?.maxInputChars;
|
||||
if (override !== undefined) return Math.max(0, Math.trunc(override));
|
||||
const envValue = Number.parseInt($env.MNEMOPI_EMBEDDING_MAX_INPUT_CHARS ?? "", 10);
|
||||
if (Number.isFinite(envValue) && envValue >= 0) return envValue;
|
||||
return 32000;
|
||||
}
|
||||
|
||||
/**
|
||||
* Right-truncate every input to {@link effectiveMaxInputChars} so a runaway
|
||||
* retention transcript can't blow past the embedding model's context window.
|
||||
* Returns the original array when no input needs trimming (the common case);
|
||||
* the new array is allocated only when at least one input is oversized so
|
||||
* we don't churn arrays for the typical short-query path through `embedQuery`.
|
||||
* Emits one debug-or-warn log per call summarizing how many inputs were
|
||||
* trimmed and by how much — silent truncation was the original bug (#3126).
|
||||
*/
|
||||
function capInputs(texts: readonly string[]): readonly string[] {
|
||||
const max = effectiveMaxInputChars();
|
||||
if (max === 0) return texts;
|
||||
let trimmed: string[] | null = null;
|
||||
let trimmedCount = 0;
|
||||
let maxOriginalLen = 0;
|
||||
for (let i = 0; i < texts.length; i++) {
|
||||
const text = texts[i] ?? "";
|
||||
if (text.length <= max) continue;
|
||||
if (trimmed === null) trimmed = texts.slice() as string[];
|
||||
trimmed[i] = text.slice(0, max);
|
||||
trimmedCount++;
|
||||
if (text.length > maxOriginalLen) maxOriginalLen = text.length;
|
||||
}
|
||||
if (trimmed === null) return texts;
|
||||
logger[mnemopiDebugEnabled() ? "warn" : "debug"]("mnemopi: embedding input truncated", {
|
||||
inputCount: texts.length,
|
||||
trimmedCount,
|
||||
maxOriginalLen,
|
||||
maxInputChars: max,
|
||||
});
|
||||
return trimmed;
|
||||
}
|
||||
|
||||
function embeddingApiKey(): ApiKey {
|
||||
const active = activeEmbeddingOptions();
|
||||
if (active?.apiKey !== undefined) {
|
||||
@@ -408,6 +456,7 @@ export async function embed(texts: readonly string[]): Promise<EmbeddingMatrix |
|
||||
if (texts.length === 0 || embeddingsDisabled()) {
|
||||
return null;
|
||||
}
|
||||
texts = capInputs(texts);
|
||||
const activeProvider = resolveEmbeddingProvider(activeEmbeddingOptions()?.provider);
|
||||
if (activeProvider !== undefined) {
|
||||
try {
|
||||
|
||||
@@ -161,19 +161,22 @@ function resolveRuntimeOptions(options: MnemopiOptions): ResolvedMnemopiRuntimeO
|
||||
const embeddingApiUrl = options.embeddingApiUrl ?? nestedEmbeddings?.apiUrl;
|
||||
const embeddingApiKey = options.embeddingApiKey ?? nestedEmbeddings?.apiKey;
|
||||
const embeddingProvider = resolveEmbeddingProvider(nestedEmbeddings?.provider);
|
||||
const embeddingMaxInputChars = nestedEmbeddings?.maxInputChars;
|
||||
|
||||
const embeddings =
|
||||
embeddingDisabled !== undefined ||
|
||||
embeddingModel !== undefined ||
|
||||
embeddingApiUrl !== undefined ||
|
||||
embeddingApiKey !== undefined ||
|
||||
embeddingProvider !== undefined
|
||||
embeddingProvider !== undefined ||
|
||||
embeddingMaxInputChars !== undefined
|
||||
? {
|
||||
disabled: embeddingDisabled,
|
||||
model: embeddingModel,
|
||||
apiUrl: embeddingApiUrl,
|
||||
apiKey: embeddingApiKey,
|
||||
provider: embeddingProvider,
|
||||
maxInputChars: embeddingMaxInputChars,
|
||||
}
|
||||
: undefined;
|
||||
|
||||
|
||||
@@ -33,6 +33,8 @@ export interface MnemopiEmbeddingRuntimeOptions {
|
||||
apiUrl?: string;
|
||||
apiKey?: ApiKey;
|
||||
provider?: MnemopiEmbeddingProvider | ((texts: readonly string[]) => EmbeddingOutput | Promise<EmbeddingOutput>);
|
||||
/** Override `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`. `0` disables the cap. See `config.embeddingMaxInputChars`. */
|
||||
maxInputChars?: number;
|
||||
}
|
||||
|
||||
export interface MnemopiLlmRuntimeOptions {
|
||||
@@ -61,6 +63,7 @@ export interface ResolvedMnemopiEmbeddingRuntimeOptions {
|
||||
apiUrl?: string;
|
||||
apiKey?: ApiKey;
|
||||
provider?: MnemopiEmbeddingProvider;
|
||||
maxInputChars?: number;
|
||||
}
|
||||
|
||||
export interface ResolvedMnemopiLlmRuntimeOptions {
|
||||
|
||||
@@ -0,0 +1,101 @@
|
||||
import { afterEach, describe, expect, it } from "bun:test";
|
||||
import "./setup";
|
||||
import {
|
||||
embed,
|
||||
resetEmbeddingProviderForTests,
|
||||
setEmbeddingProviderForTests,
|
||||
} from "@oh-my-pi/pi-mnemopi/core/embeddings";
|
||||
import { withMnemopiRuntimeOptions } from "@oh-my-pi/pi-mnemopi/core/runtime-options";
|
||||
|
||||
/**
|
||||
* Regression coverage for issue #3126: `MnemopiSessionState.retainMessages`
|
||||
* passes the entire multi-turn transcript to `embed()`. Long sessions used to
|
||||
* overflow whatever ctx the embedding server was started with — llama.cpp
|
||||
* rejects oversized requests with `request (N tokens) exceeds the available
|
||||
* context size`, OpenAI silently right-truncates. `embed()` now caps each
|
||||
* input to `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS` (default 32000) before the
|
||||
* provider sees it.
|
||||
*/
|
||||
function captureProvider(): {
|
||||
embed: (texts: readonly string[]) => AsyncGenerator<number[][]>;
|
||||
calls: string[][];
|
||||
} {
|
||||
const calls: string[][] = [];
|
||||
return {
|
||||
calls,
|
||||
async *embed(texts) {
|
||||
calls.push([...texts]);
|
||||
yield texts.map(text => [text.length]);
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
const ENV_KEY = "MNEMOPI_EMBEDDING_MAX_INPUT_CHARS";
|
||||
|
||||
function withEnvValue<T>(value: string | undefined, fn: () => Promise<T>): Promise<T> {
|
||||
const previous = process.env[ENV_KEY];
|
||||
if (value === undefined) delete process.env[ENV_KEY];
|
||||
else process.env[ENV_KEY] = value;
|
||||
return fn().finally(() => {
|
||||
if (previous === undefined) delete process.env[ENV_KEY];
|
||||
else process.env[ENV_KEY] = previous;
|
||||
});
|
||||
}
|
||||
|
||||
afterEach(() => {
|
||||
resetEmbeddingProviderForTests();
|
||||
});
|
||||
|
||||
describe("embed() input cap (#3126)", () => {
|
||||
it("truncates oversized inputs to the default cap before reaching the provider", async () => {
|
||||
const provider = captureProvider();
|
||||
setEmbeddingProviderForTests(provider);
|
||||
|
||||
const huge = "x".repeat(40_000);
|
||||
await withEnvValue(undefined, () => embed(["short", huge]));
|
||||
|
||||
expect(provider.calls).toHaveLength(1);
|
||||
const [seenShort, seenHuge] = provider.calls[0] ?? [];
|
||||
expect(seenShort).toBe("short");
|
||||
expect(seenHuge?.length).toBe(32_000);
|
||||
});
|
||||
|
||||
it("honors MNEMOPI_EMBEDDING_MAX_INPUT_CHARS env override", async () => {
|
||||
const provider = captureProvider();
|
||||
setEmbeddingProviderForTests(provider);
|
||||
|
||||
await withEnvValue("1024", () => embed(["y".repeat(5000)]));
|
||||
|
||||
expect(provider.calls[0]?.[0]?.length).toBe(1024);
|
||||
});
|
||||
|
||||
it("disables the cap when the env override is 0", async () => {
|
||||
const provider = captureProvider();
|
||||
setEmbeddingProviderForTests(provider);
|
||||
|
||||
const huge = "z".repeat(50_000);
|
||||
await withEnvValue("0", () => embed([huge]));
|
||||
|
||||
expect(provider.calls[0]?.[0]).toBe(huge);
|
||||
});
|
||||
|
||||
it("respects a constructor-scoped maxInputChars override", async () => {
|
||||
const provider = captureProvider();
|
||||
setEmbeddingProviderForTests(provider);
|
||||
|
||||
await withEnvValue(undefined, () =>
|
||||
withMnemopiRuntimeOptions({ embeddings: { maxInputChars: 256 } }, () => embed(["w".repeat(10_000)])),
|
||||
);
|
||||
|
||||
expect(provider.calls[0]?.[0]?.length).toBe(256);
|
||||
});
|
||||
|
||||
it("returns the original array reference when no input needs trimming", async () => {
|
||||
const provider = captureProvider();
|
||||
setEmbeddingProviderForTests(provider);
|
||||
|
||||
await withEnvValue("1024", () => embed(["fits", "still fits"]));
|
||||
|
||||
expect(provider.calls[0]).toEqual(["fits", "still fits"]);
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user