The matcher compares selector ids case-insensitively, but the lock's
bundled-catalog lookup was exact-case: Anthropic/Claude-Opus-5 exact-
matched OpenRouter's flat id while getBundledModel("Anthropic", ...)
missed, silently re-enabling the aggregator shadow the lock exists to
prevent. Scan the named provider's bundled ids case-insensitively.
- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
The idle watchdog aborts the request signal and cursor.ts closes
the Connect stream, so there is no in-flight server exec to race.
Unmarked MCP/todo blocks can continue once every emitted call has
a matching result, same as HTTP/2 RST.
The exact-echo check used accent-insensitive compare, but the collision
suffix path was a case-sensitive startsWith. AuthLoader-3 vs authloader
therefore leaked through as a real description.
The first HUD commit hid Name: Name. The cause was earlier: task
name was copied into identity.label, which became progress.description
and skipped generateTaskLabel. Keep the handle for id allocation, but
only treat eval label as a real UI description so the tiny-model
summary can run.
The anchored Subagents list printed only `Id: description` and treated a
label that repeated the spawn handle as a real description. Show the same
⟨role⟩ badge as inline task rows and omit descriptions that only echo the id.
StopOnTextCriteria decoded the last STOP_DECODE_WINDOW_TOKENS of the whole
sequence, so prompt tokens were eligible for matching. A prompt that itself
contains the stop string stops generation at the first generated token and
yields an empty title.
Anchor the window to the generation boundary by recording the first
generated index per batch entry. Existing local title models are
unaffected: with the assistant-prefill prompt shape, the example `</title>`
tags sit outside the 32-token window for normal messages, so no shipping
model changes behavior. The bug becomes reachable with any chat-level
few-shot prompt that places the stop string near the generation boundary.
The memory-extraction prompt concatenated its instructions, few-shot
examples, and the user message into a single user turn, so a small local
model could not distinguish instructions from input and frequently echoed
the Globex/weather examples instead of extracting facts.
Send the instructions as a real system turn and the raw text as the user
turn. The tiny worker protocol gains a systemPrompt field, and Mnemopi
completion input carries task metadata so the backend selects the right
prompt per call.
Drop the code-built MEMORY_EXTRACTION_TEMPLATE rather than porting it:
prompt text belongs in .md files, and resolveMemoryCompletionInput already
overrides that template for every extraction call, so Mnemopi rendered it
only for the result to be discarded.
Measured on ONNX q4 CPU, LFM2.5-1.2B memory extraction improved from 1/8
to 5/8 once the roles were separated.
The claude-plugins provider read installed_plugins.json but never the
`enabledPlugins` map Claude Code keeps in ~/.claude/settings.json and
<project>/.claude/settings(.local).json. Two consequences: a plugin the
user switched off for a project still loaded there, and a local-scope
install enabled for a project never loaded unless the project directory
matched the install's recorded projectPath exactly.
Merge enabledPlugins across the same layers Claude Code consults (user
settings, then the active project root's and cwd's .claude/settings.json
and settings.local.json; later wins) and apply it in
listClaudePluginRoots: `false` hides the plugin, `true` opts a
local-scope install in regardless of projectPath. Untouched ids keep the
existing behavior. Contributing settings files join the cache key.
Solves: Claude marketplace plugins loading in the wrong projects
Tests: claude-plugins.test.ts — per-project off switch (local wins over
settings.json, other projects unaffected) and enabledPlugins:true
opt-in for a local install recorded under a parent directory
The local text read path opened the same file for every consumer. A ranged
read of a file within the snapshot cap cost four opens and three decodes:
an 8KiB binary sniff, a streaming scan for the rendered window, a whole-file
read for bracket context, and another whole-file read to hash the snapshot.
Whole-file reads under the structural summarizer paid a fifth. Two of those
readers also ran normalizeToLF over the same bytes.
Read the bytes once at or below SNAPSHOT_MAX_BYTES and derive every view
from them: sniff the leading 8KiB of the buffer, slice the rendered window
out of it under the identical line and byte budgets, index bracket context
into its addressable lines, and hand the normalized text to the snapshot
store and the summarizer. Past the cap nothing wants the whole file, so the
streaming reader stays.
Line byte lengths are walked out of the buffer rather than measured on the
decoded strings, so reported byte counts and the truncation boundary stay
exact for content that is not valid UTF-8.
The buffered text is BOM-stripped for hashing, matching the decoder the
patcher's live read uses. A whole-file read of a BOM file previously hashed
its tag from BOM-bearing text, so the following edit only applied through
stale-hash recovery and told the model the file had changed externally when
it had not.
Also stop rejoining lines into a fresh whole-file string on the way to
tree-sitter when the caller still holds that text, and drop the unread
selectedBytesTotal accounting from the streaming reader.
Measured on 966 differential cases across CRLF, BOM, lone-CR, invalid-UTF-8,
oversized-line, empty, no-trailing-newline, multi-range and raw shapes: the
spurious recovery warning is the only behavioral difference. Raw reads, which
skip the tree-sitter parse that dominates everything else, get 30-45% faster
(2.7MB: 10.3ms -> 5.7ms); non-raw reads 1-3%.
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.
`lsp regressions > detects pyright and pylsp in Windows virtualenv Scripts for
Python-only roots` fails on any machine that has ~/.omp/agent/lsp.json:
expect(config.servers[server]?.resolvedCommand).toBe(localBin)
Expected: ".../.venv/Scripts/pyright-langserver.exe"
Received: undefined
The test was not asserting against the packaged defaults at all. loadConfig
walks the user config dirs (~/.omp/agent, ~/.pi/agent, ~/.claude) via
getConfigDirPaths, which resolves from os.homedir(). Any user lsp.json with a
`servers` block sets hasOverrides, which takes loadConfig off its auto-detect
branch and onto the override branch, where the user's rootMarkers replace the
packaged ones.
On this machine that file overrides pyright with
rootMarkers: ["pyproject.toml", "uv.lock", "requirements.txt", "setup.py"]
which contains neither `pyrightconfig.json` nor `setup.cfg` — precisely the
two markers the test creates. loadConfig therefore returned zero servers.
`ruff` is absent from that file, keeps the packaged rootMarkers, and its
sibling tests pass, which is why only this one failed.
Confirmed by probing inside the test: hasRootMarkers() true and
resolveCommand() returning the .exe, while loadConfig().servers was {} — the
detection helpers were fine, the branch was not.
Point os.homedir() at an empty directory for every test in the file so they
see a pristine environment. Bun's os.homedir() reads the passwd entry rather
than $HOME, so the env var alone does not redirect the walk; both are set.
Verification: 76 pass / 0 fail (was 75 / 1) on a machine with a user lsp.json.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A selector like anthropic/claude-opus-5 is both the anthropic provider's
canonical selector and, verbatim, an OpenRouter model id. When the named
provider is out of the candidate set (disabled, missing creds at boot), the
exact provider-scoped reference misses and the raw-id flat match re-binds the
request onto OpenRouter's same-named model. Requests that should fail closed
instead bill the aggregator at per-token Claude prices.
Raw-id fallback is only legitimate when the named provider does not carry the
id (openai/gpt-4o:extended). Lock the reference when it does by checking the
bundled catalog, then fail rather than shadow. openrouter/anthropic/claude-opus-5
still selects OpenRouter explicitly.
isMultiplexerSession() treats TMUX, STY, ZELLIJ, HERDR_ENV=1, the CMUX_*
markers and a tmux/screen TERM as authoritative, and routes rendering down the
path that cannot rebuild scrollback. Nine tests across four files assert the
destructive full-paint behavior and fail for anyone running the suite inside
tmux, screen, Zellij or CMUX.
component-render.test.ts already cleared HERDR_ENV for exactly this reason but
covered only one of the nine signals. Replace it with a shared helper that
clears the whole family and restores it afterwards, and call it from the other
three files.
Three tests fail on a clean checkout depending on where the repo lives and
which terminal runs the suite:
- status-line-path builds its fake home inside the checkout, so a clone under
/tmp lands in a SCRATCH_ROOTS prefix and renders the scratch icon.
- status-line/component renders a fixed 120 columns, so a long checkout path
or branch name pushes the cost segment out of the assertion.
- bash-executor runs an interactive login zsh, which loads the system
/etc/zshrc; under Apple Terminal that appends session-save lines to the
captured output. HOME does not isolate a system-level file.