ModelRegistry registers a built-in llama.cpp discovery as `provider:
"llama.cpp"` (model-registry.ts:1063), so a reverse-proxied or
public-DNS llama.cpp endpoint never trips the loopback heuristic and
was still falling through to no append-only context.
Add "llama.cpp" to LOCAL_INFERENCE_PROVIDERS, refresh the docstring
on `hasLocalLoopbackBaseUrl` (it covers user-defined providers; built-
in local ids are caught by the allowlist), and extend the test to
exercise a public-host llama.cpp baseUrl so the allowlist path is
covered independently of the loopback heuristic.
Refs #3033
Ollama, LM Studio, and llama.cpp / vLLM all do byte-prefix KV cache reuse
on the model server side. The agent loop's non-append-only path rebuilds
the system prompt and tool catalogue on every turn through fresh
allocations (`normalizeTools`, `convertToLlm`, optional memory-backend
`beforeAgentStartPrompt` injection), which dirties enough leading bytes
to invalidate the cache and force a full prompt re-evaluation.
`shouldAutoEnableAppendOnlyContext` previously only recognized DeepSeek
and Xiaomi Token Plan. Extend the auto-detect:
- Allowlist the known local-server provider ids (`ollama`,
`ollama-cloud`, `lm-studio`).
- Detect user-defined local servers by parsing `baseUrl`: loopback,
RFC1918 private IPv4, and `.local` mDNS hostnames.
- Keep the existing `compat.supportsStore` opt-in escape hatch and the
explicit `provider.appendOnlyContext: on`/`off` override paths.
Regression coverage in `append-only-context-mode.test.ts` exercises
each new positive case plus negative samples (172.15/172.32 just outside
RFC1918, public hosts, malformed URLs).
Fixes#3033
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
Centralized append-only context auto-mode resolution so SDK sessions, interactive sessions, and the status display use the same provider checks. Auto mode now enables for Xiaomi/SGLang hosts and explicit stored-request compat signals while preserving DeepSeek behavior.
Added focused regression coverage for Xiaomi Token Plan SGLang HiCache endpoints and explicit on/off behavior.
Fixes#1851