ImmutablePrefix caches system prompt + tool specs after first build()
so subsequent turns reuse identical byte sequences. AppendOnlyLog
converts messages once via syncMessages() and only appends deltas
on further turns — prior-turn bytes stay stable.
- New module: packages/agent/src/append-only-context.ts
StablePrefix, AppendOnlyLog, AppendOnlyContextManager
- AppendOnlyContextManager added to AgentLoopConfig
- Wired into streamAssistantResponse in agent-loop.ts
- Toggleable via provider.appendOnlyContext setting (auto/on/off)
- Default auto enables for deepseek provider
- 38 tests covering prefix, log, sync, compaction handling
- /session info surfaces current active state
Replaced overwrite-style session rewrites with an EPERM fallback that moves the old session file aside before retrying and restores it if the retry fails.
Added regression coverage for active-session rewrite recovery so the session remains writable after the fallback.
Fixes#1337
Previously, any AbortError (including the 15s timeout from
AbortSignal.timeout) was immediately rethrown, which prevented the
SGP→AMS fallback loop from reaching AMS during regional SGP outages.
Now only re-throws AbortError when the caller explicitly cancelled
(signal?.aborted), letting timeout-induced aborts continue to the
next endpoint via the existing lastError/continue path.
Xiaomi MiMo supports /v1/chat/completions in addition to the previously
used /anthropic/v1/messages. The Anthropic protocol's thinking-signature
mechanism, which Xiaomi does not implement, causes unsigned reasoning
blocks to be converted to visible text — leaking the model's internal
thinking into the conversation history and triggering repeated reasoning
loops.
OpenAI Chat Completions avoids this entirely:
- No signature-based thinking chain verification
- Standard reasoning_effort parameter instead of thinking blocks
- Already used by Kilo, NanoGPT, OpenRouter, and ZenMux for Xiaomi
Changes:
- xiaomiModelManagerOptions: anthropic-messages → openai-completions
- xiaomi.ts validation: /v1/messages → /v1/chat/completions
- tp- keys: try SGP first, fall back to AMS
- models.json: all 5 Xiaomi models updated to openai-completions + /v1
- Revert AnthropicHeaderOptions.usesApiKeyAuth (no longer needed)
- Remove AnthropicCompat.usesApiKeyAuth from types
- Update test expectations for SGP-first routing and OpenAI endpoints
Root cause: getProviderBaseUrl('xiaomi') returned the standard
endpoint baseUrl from the bundled catalog (api.xiaomimimo.com/anthropic),
overriding the token-plan URL (token-plan-ams.xiaomimimo.com/anthropic)
that should be used for tp- prefixed keys.
Fix:
- xiaomiModelManagerOptions: for tp- keys, compute baseUrl from the
key prefix alone, ignoring config?.baseUrl from the bundled catalog.
- Add usesApiKeyAuth flag to AnthropicCompat / AnthropicHeaderOptions.
- Xiaomi models set usesApiKeyAuth: true so the anthropic provider sends
X-Api-Key header (matching what the Xiaomi validation code expects)
instead of Authorization: Bearer.
Gated WebP encoding behind OMP_NO_WEBP environment variable so that local
llama.cpp vision models (which use the STB library that lacks WebP
decoding support) can accept browser snapshots without returning HTTP 400.
Upstream reference: cline/cline PR #9837
(https://github.com/cline/cline/pull/9837) implements a similar workaround
for the same llama.cpp STB WebP incompatibility.
When /plan <text> or /goal set <text> is typed while plan/goal mode is
already active, the handler shows a warning and clears the editor, but
the typed text was silently discarded — the user could not recover it
with Up Arrow.
Fix: save the text to editor history before clearing, so it remains
recoverable via input history.
Fixes#1287
When session.prompt() returns, idle-flush tasks for async-job result
deliveries are scheduled via #schedulePostPromptTask (1ms delay) and
added to #postPromptTasks immediately. The 800ms loop timer could fire
in that window before isStreaming became true, causing the loop to
submit the next prompt while the delivery turn was still pending. The
delivery then hit AgentBusyError and the job result was silently dropped.
Add AgentSession.hasPostPromptWork (= #postPromptTasks.size > 0) and
include it in #isLoopAutoSubmitBlocked() alongside isStreaming and
isCompacting. Add a regression test that verifies the loop defers when
hasPostPromptWork is true and fires once it becomes false.
Fixes#1294
WSLg exposes WAYLAND_DISPLAY, so readImageFromClipboard() took the
native arboard path on WSL2. arboard::Clipboard::get_image() returns
ContentNotAvailable on WSLg because the Wayland clipboard does not
carry image payloads from the Windows clipboard, and that surfaced as
silent 'No image in clipboard' on Ctrl+V.
Detect WSL via WSL_DISTRO_NAME / WSL_INTEROP and read the image with a
PowerShell one-liner that emits base64-encoded PNG bytes from
[System.Windows.Forms.Clipboard]::GetImage(). Fall back to the native
bridge when PowerShell returns nothing, exits non-zero, or is missing,
so non-WSLg Wayland setups continue working unchanged.
Fixes#1280
- Extended the abort scenario shell command from a 5-second sleep to 60 seconds.
- Raised the corresponding test timeout to 15,000 ms for the abort test case.
- Changed hashline format to canonical `§` section headers and `"`/`"`/`≔` operations across grammar, parser, and docs.
- Reworked range and op parsing so single anchors are valid, legacy `-`/`-=` ops error, and empty `≔` payloads now delete ranges.
- Removed legacy `HL_EDIT_SEP`, `$HSEP$`, and `hsep` plumbing, adopting raw payload lines.
- Updated execution, input, renderer, and streaming flows to use `HL_FILE_PREFIX` and `HL_OP_CHARS` helpers.
- Updated session-stats parsing to `§`/`≔`/`"`/`"` format, patch envelopes, and bumped parser versions.
- Removed the hashline-separator benchmark script and its PI_HL_SEP job orchestration.
- Added a shared `interruptHint()` utility in the modes shared module to generate the interrupt suffix with themed bracket glyphs.
- Replaced hardcoded working-message interrupt text in interactive and event controllers with calls to `interruptHint()`.
- Updated working-message rendering to recognize and strip the new themed hint when applying shimmer styling.
- Updated OpenAI completions compatibility logic to translate a disabled-reasoning `minimal` effort to `none` for Fireworks endpoints.
- Added a Fireworks-specific test confirming the disableReasoning payload uses `reasoning_effort: "none"`.
- Refactored the disable-reasoning tests to reuse a shared request-payload capture helper.
- Added constants for OAuth and metadata token endpoints in the Vertex AI repro test.
- Expanded the fetch mock in the issue #1270 test to return tokens for either token endpoint, accommodating different ADC environments.
- Updated OpenAI completions parameterization to honor `disableReasoning` on effort-based compatible models by sending the minimum supported effort.
- Expanded commit and title generation token budgets so reasoning models can return output after internal thinking while non-reasoning calls keep existing limits.
- Switched title generation to request a `set_title` tool call, added extraction from tool-call arguments, and updated tests for the new behavior.
- Updated interactive mode to render the loader spinner with the current working message accent when available.
- Added a fallback to theme accent and a reset color code when no working-message accent is provided.
- Loading message rendering now derived session-specific accent colors and applied them to shimmer output when a session name was available.
- Shimmer palettes were updated to accept raw ANSI color values as well as theme color names during compilation.
- A unit test was added to confirm shimmer text rendered with a supplied ANSI crest color.
- OpenAICodexOAuthFlow now passes a fixed redirectUri option to
OAuthCallbackFlow, preventing the silent fallback to a random port.
OpenAI's /oauth/authorize accepts any localhost redirect URI (loose
validation), so the browser callback could succeed even on a non-1455
port; but /oauth/token rejects it with 403 because the port is not
in the registered allowlist. The fix surfaces port conflicts as an
immediate actionable error instead.
- exchangeCodeForToken now reads the 403 response body (error /
error_description) and includes it in the thrown message, matching
the existing refreshOpenAICodexToken error handling.
- Added loginOpenAICodexDevice (openai-codex-device provider): device-
code flow via /api/accounts/deviceauth/usercode + polling, avoiding a
local callback server entirely. Credentials are stored under the
existing 'openai-codex' key so no model or tooling reconfiguration is
needed. Useful when port 1455 is unavailable or the browser flow is
blocked by network/proxy.
Fixes#1277
- Added SettingsList#setItems to replace items and clamp selection to a valid index after updates.
- Updated SettingsSelector to rebuild active memory items on backend changes and skip refresh when appropriate.
- Switched MCP wizard and command spinners to theme frames with themed initial frame and 80ms updates.
- Reworked welcome intro animation for a 3-second eased sweep with optional shine blending.
- Added memory backend refresh tests and aligned package changelogs with the updated behavior.
- Added optional `onBeforeYield` configuration and `setOnBeforeYield` in Agent, executed before follow-up checks.
- Added `YieldQueue` to `AgentSession`, with setup/teardown and streaming/idle flush via `setOnBeforeYield`.
- Replaced immediate async-result follow-up dispatch with queued batch entries, including stale-state suppression.
- Added MCP follow-up queueing in SDK, deduplicating updates by `serverName` and `uri`.
- Added changelog entries for `onBeforeYield`, async-result batching, MCP dedupe, and `display.shimmer` modes.
- Added yield queue unit tests for streaming emission, debounced idle batches, stale filtering, and error isolation.
- Added `display.shimmer` setting (`classic`, `kitt`, `disabled`) with default `classic` and UI metadata.
- Added shimmer mode resolution with defaulting plus classic/KITT intensity tier profiles and thresholds.
- Added `getFgAnsi` handling across shimmer compilation and progress-bar themes, with fallback ANSI output defaults.
- Reworked shimmer rendering to coalesce same-tier segments, cache palettes by symbol, and map disabled mode to mid-tier.
- Tuned loader timing to a 16ms render interval and 80ms spinner stepping for smoother ~60fps updates.