- Verified that side-channel turns maintain stable prompt cache parity with main turns.
- Applied secret and tool obfuscation to ephemeral assistant side-channel data to prevent leakage.
- Updated `Agent` and `AgentSession` test suites to validate consistency under native dialect environments.
Mirror of the provider-anchored handoff test: a ~20k-token stored conversation
whose provider usage was deflated to 1k (as a before_provider_request compressor
would) now triggers pre-prompt compaction, proving #estimateStoredContextTokens
floors the decision through the real call-site (not just the helper math).
CI caught that flooring by the raw local estimate falsely triggers compaction on
thinking-heavy turns: estimateTokens counts the opaque thinkingSignature /
redactedThinking payloads (providers bill them on replay, #2275), but their local
byte size diverges wildly from what the provider actually charges — so a turn
with a large encrypted-reasoning blob but small provider usage would trip the
floor (broke agent-session-handoff 'provider-anchored usage' test).
estimateTokens now takes { excludeEncryptedReasoning } and the compaction floor
(#estimateStoredContextTokens) uses it: the floor counts only reliably-countable,
on-wire-compressible content (text, tool results, tool calls), while the provider
usage arm of compactionContextTokens still accounts for encrypted reasoning. This
keeps the encrypted-reasoning case provider-anchored while still flooring upward
when a before_provider_request hook compresses tool results.
A before_provider_request extension (a context-compression proxy like Headroom,
an obfuscator, or inline snapcompact) can shrink the outgoing request below the
real stored conversation. The provider then reports deflated prompt tokens, so
the auto-compaction threshold never fires and the stored history grows unbounded
until it overflows the context window and can no longer be compacted at all.
Add compactionContextTokens(provider, storedEstimate) = max(provider, estimate)
and apply it to both the pre-prompt and post-response compaction decisions,
flooring the provider-reported tokens by the agent's own estimate of the stored
conversation. Display and cost accounting still use exact provider usage; only
the compaction trigger takes the floor.
- Added `buildSideRequestContext` to the `Agent` class to generate prompt-cache-friendly provider contexts.
- Updated ephemeral side-channel turns to forward the full tool catalog to maintain prompt cache hit rates.
- Injected a `developer` role reminder into ephemeral turns to instruct the model to suppress tool calls.
- Implemented automatic post-processing to strip any tool calls from ephemeral turn responses.
- Exported message and dialect helper functions in `agent-loop.ts` to support context construction.
- Added a set to track tool result IDs that have been rewound to ensure they are not re-appended to the session history.
- Updated message handling logic to conditionally skip persistence for tool results associated with active rewind operations.
- Implemented message deduplication in `AdvisorRuntime` to collapse verbatim re-injected primary context (plan rules/approved plans) between turns.
- Reduced token consumption by replacing identical primary context segments with a status marker.
- Updated history formatter to allow expansion of load-bearing context types specifically, while retaining one-line summaries for other custom messages.
- Fixed session history desynchronization during `rewind` by performing a full context rebuild after applying branch changes.
- Prevented `rewind` tool output and assistant side-channel data from polluting the prompt cache by flushing and sanitizing session state.
- Added explicit test coverage for context reconstruction and assistant message sanitization after rewind events.
- Remove the static "pending" hourglass icon from edit and write tool headers to reduce visual noise.
- Update multi-file status lines to use the active spinner icon directly instead of replacing a static icon, ensuring consistent liveness cues.
The prompt-inventory test sliced on the '# Inventory'/'ENV' markers from the
PR's merge-base prompt.md. On current main (chore: prompt reorder) the heading
is '# Tool Inventory' and the 'ENV' marker is gone, which is why the file was
deleted there. Accept either layout so the resurrected tests (incl. the SDK
'render provided tools' contract) pass after the cherry-pick.
fork() reset mnemopi conversation tracking directly but skipped the shared new-transcript reset, so the folded/promoted first-turn memory stayed in #baseSystemPrompt. The next turn re-recalled and the change-detection saw no diff, taking the fallback promotion path and injecting the <memories> block twice into the forked session prompt. Route fork() through #resetMemoryContextForNewTranscript() like the other reset paths and add a regression test asserting the forked prompt contains recalled memory exactly once.
The PR placed the entry under the released [16.0.7] section with a
duplicate ### Fixed header, mutating released notes. Move it under
[Unreleased] per the repo changelog convention.
The #3099 test omitted the startInAllScope flag, so it passed against the
buggy baseline too (empty folder defaults to folder scope regardless). Force
the removed flag via a cast to pin the real contract: even when a caller asks
for all-projects scope on an empty folder, the picker must stay folder-scoped.
Proven: passes on head src, fails on baseline src (renders '(all projects)').
Move the legacy pi-ai compat note from the released [16.0.5] section to
[Unreleased], and drop the inaccurate bare-`typebox` claim: bare
`typebox` specifier handling already exists on main (issue #2858,
TYPEBOX_SPECIFIER_FILTER / TYPEBOX_IMPORT_SPECIFIER_REGEX) and this PR's
diff does not touch it. The entry now scopes the change to the actual
delta: restored getModel/getModels aliases and StringEnum enum-object
support.
The parallel streaming walker accumulated match_count toward the
first-page stop budget, but match_count can exceed the matches actually
returned (collected) by one whenever a file overflows its per-file cap.
With the production config (max_count=2000, max_count_per_file=21) that
over-count could quit the walk a few percent short of the requested
page. Budget on collected instead, matching run_sequential_grep and
aggregate_parallel_results.
`MnemopiSessionState.retainMessages` hands `embed()` the chronological
multi-turn transcript (oldest -> newest). A naive `slice(0, max)` cap landed
on the oldest turns and dropped the most recent content, so every retained
episode past the cap collapsed onto essentially the same prefix vector and
dense recall could not match topics introduced after the first 8192 chars.
`capInputs` now routes oversized inputs through `clipToWindow`, which keeps
roughly half the cap from the head and half from the tail with a small
`[...]` elision marker between them. Short inputs and array reference pass-
through are unchanged. Falls back to a tail-only clip when `max` is too small
to fit a useful split. Locked in by a new `embedding-input-cap.test.ts` case
that pins markers at both ends of a 50k transcript and asserts both survive
the clip.
Fixes#3126
Defaulted MNEMOPI_EMBEDDING_MAX_INPUT_CHARS to 8192 so the automatic guard matches bge-m3 and OpenAI text-embedding context limits by default. Larger local embedding servers such as Qwen3-Embedding with 32k ctx can still raise the cap, and 0 still disables truncation.
Fixes#3126
`MnemopiSessionState.retainMessages` (`packages/coding-agent/src/mnemopi/state.ts:352`)
always calls `prepareRetentionTranscript(messages, true)` and hands the whole
multi-turn transcript to `embed([transcript])`. Long sessions (especially CJK
content) routinely outgrow the embedding model's context window, and llama.cpp's
`/embeddings` server rejects oversized requests with
`request (N tokens) exceeds the available context size` — every retain after
that point silently lost its vector row, leaving recall on FTS-only.
Capped per-input length inside `embed()` (the single chokepoint every retain /
query / consolidate flow funnels through) at `MNEMOPI_EMBEDDING_MAX_INPUT_CHARS`
(default 32000 chars ≈ 8k English tokens / 16–32k CJK tokens, override via env
or `embeddings.maxInputChars` runtime option; `0` disables). The new array
is allocated only when at least one input is oversized, so the typical short-
query path through `embedQuery` still passes the original array through; the
truncation also emits a debug-or-warn log so the resize is no longer silent.
Fixes#3126
The chunked-welcome refactor moved welcome-timer arming into socket.onOpen,
leaving the connect phase uncovered: if the relay blackholes the WebSocket
handshake (no onOpen and no onClose), the timer never arms and /join hangs
forever. Baseline armed the 30s timeout right after connect(); restore that
so a stalled handshake still rejects the join. onOpen continues to re-arm
(resetting the budget) once the socket opens.
resolveModels("all") expanded the full TINY_LOCAL_MODELS registry, which now
includes the qwen3-1.7b entry marked unsupportedReason. loadPipeline() throws
for such specs, so the download worker reported it as failed and the bulk
command exited with "One or more tiny title models failed to download" even
when every usable model downloaded. Filter unsupported specs out of the `all`
prefetch path; explicit single-model requests are unchanged. Addresses the
unaddressed Codex P2 on PR #3133.