fix(coding-agent): include llama.cpp in append-only allowlist

ModelRegistry registers a built-in llama.cpp discovery as `provider:
"llama.cpp"` (model-registry.ts:1063), so a reverse-proxied or
public-DNS llama.cpp endpoint never trips the loopback heuristic and
was still falling through to no append-only context.

Add "llama.cpp" to LOCAL_INFERENCE_PROVIDERS, refresh the docstring
on `hasLocalLoopbackBaseUrl` (it covers user-defined providers; built-
in local ids are caught by the allowlist), and extend the test to
exercise a public-host llama.cpp baseUrl so the allowlist path is
covered independently of the loopback heuristic.

Refs #3033
This commit is contained in:
roboomp
2026-06-19 08:26:14 +00:00
parent 9a07533924
commit 4353920b95
3 changed files with 16 additions and 6 deletions
+1 -1
View File
@@ -4,7 +4,7 @@
### Fixed
- Fixed Ollama (and LM Studio / loopback llama.cpp / vLLM servers) reprocessing the full prompt on every turn because `provider.appendOnlyContext: auto` only recognized DeepSeek and Xiaomi as prefix-cache providers. The auto-detect now enables append-only mode for `ollama`, `ollama-cloud`, `lm-studio`, and any baseUrl resolving to a loopback/RFC1918/`.local` host, so the system prompt + tool catalogue + prior-turn message bytes stay byte-stable across turns and llama.cpp's KV-cache prefix reuse can hit ([#3033](https://github.com/can1357/oh-my-pi/issues/3033)).
- Fixed Ollama, LM Studio, and llama.cpp (plus loopback vLLM / sglang servers) reprocessing the full prompt on every turn because `provider.appendOnlyContext: auto` only recognized DeepSeek and Xiaomi as prefix-cache providers. The auto-detect now enables append-only mode for `ollama`, `ollama-cloud`, `lm-studio`, `llama.cpp`, and any baseUrl resolving to a loopback/RFC1918/`.local` host, so the system prompt + tool catalogue + prior-turn message bytes stay byte-stable across turns and llama.cpp's KV-cache prefix reuse can hit ([#3033](https://github.com/can1357/oh-my-pi/issues/3033)).
## [16.1.1] - 2026-06-19
@@ -16,12 +16,14 @@ export interface AppendOnlyContextModel {
* tool catalogue, and message log all flow through fresh allocations every
* step (see `agent-loop.ts` `streamAssistantResponse` fallback path).
*/
const LOCAL_INFERENCE_PROVIDERS = new Set(["ollama", "ollama-cloud", "lm-studio"]);
const LOCAL_INFERENCE_PROVIDERS = new Set(["ollama", "ollama-cloud", "lm-studio", "llama.cpp"]);
/** True when `baseUrl` resolves to a loopback or RFC1918 host — the heuristic
* for user-defined providers (llama.cpp / vLLM via `models.yaml`) that omp
* has no provider id for. Substring match on parsed hostname only; ports,
* paths, and unparseable URLs return false.
/** True when `baseUrl` resolves to a loopback or RFC1918 host — covers
* llama.cpp/vLLM/sglang servers registered under a user-defined provider id
* via `models.yaml`. Built-in local provider ids (`ollama`, `lm-studio`,
* `llama.cpp`) are already handled by `LOCAL_INFERENCE_PROVIDERS`.
* Substring match on the parsed hostname only; ports, paths, and unparseable
* URLs return false.
*/
function hasLocalLoopbackBaseUrl(baseUrl: string | undefined): boolean {
if (!baseUrl) return false;
@@ -55,6 +55,14 @@ describe("shouldEnableAppendOnlyContext", () => {
expect(
shouldEnableAppendOnlyContext("auto", { provider: "lm-studio", baseUrl: "http://127.0.0.1:1234/v1" }),
).toBe(true);
// `llama.cpp` is a built-in provider id (ModelRegistry registers it for keyless local discovery);
// the allowlist must catch it even when the user reverse-proxies the server through a public host.
expect(
shouldEnableAppendOnlyContext("auto", { provider: "llama.cpp", baseUrl: "https://llamacpp.example.com/v1" }),
).toBe(true);
expect(shouldEnableAppendOnlyContext("auto", { provider: "llama.cpp", baseUrl: "http://127.0.0.1:8080" })).toBe(
true,
);
});
test("auto enables for loopback and private baseUrls (user-defined llama.cpp/vLLM)", () => {