- Added `SpeakableStream` to strip markdown noise, silence code blocks and tables, normalize links and paths, and emit sentence/clause segments.
- Reworked `Vocalizer` to segment assistant deltas in the parent process, lazily open TTS streams, idle-flush partial thoughts, and chain playback sessions.
- Added gapless streaming playback with ffmpeg/sox backends, ducking-aware pacing, fallback file playback, and immediate stop handling.
- Added IPC `sendAndFlush` support and used it in the TTS worker so audio chunks drain before blocking ONNX inference resumes.
- Added speakable-stream coverage for markdown filtering, segmentation latency, idle flushing, and forced long-segment splits.
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
- Added support for parsing unquoted bareword strings in object and array value positions.
- Implemented safety checks to prevent bareword recovery from masking structure, consuming non-finite atoms, or swallowing valid JSON delimiters.
- Included logic to preserve URL-style and Windows-style paths containing colons while rejecting invalid or ambiguous syntax.
- Removed the `requiresReasoningSuppressionPrompt` compatibility flag and associated logic.
- Simplified `buildOpenAIResponsesChainedParams` by removing support for trailing input scaffolding.
- Cleaned up parameter builders and test suites that handled the suppressed developer role messages.
- Simplified the check for empty Anthropic thinking signatures during same-model replays by leveraging optional chaining.
- Preserved the existing behavior of returning an empty array for empty signature strings while cleaning up unnecessary logical branches.
- Extracted the evaluation of `bareModelId(modelId)` into a `canonicalId` variable to avoid duplicate calls.
- Simplified assertions in the prior-turn thinking tests to check for an empty text blocks array directly.
- Stopped same-model official Anthropic replays from falling back to text demotion when thinking blocks lack signatures.
- Ensured unsigned thinking blocks are stripped instead of converted to textual dialect fallbacks when replayed on the same model.
- Removed an unused variable assignment in `renderDemotedThinking`.
ZenMux's `anthropic-messages` route (`zenmux.ai/api/anthropic`) forwards to
signature-enforcing Anthropic and returns full thinking signatures, but the
compat builder classified it as a non-signing reasoning endpoint via the
generic `reasoning && !official` default (`replayUnsignedThinking: true`).
Same failure class as GitHub Copilot #2851: when a checkpoint/branch-return
turn is an abandoned tool-use turn (adaptive Sonnet 5 emits a tool call then
ends on `stop`/`end_turn`), `transformMessages` correctly strips its
end_turn-bound, unreplayable signature. On a `replayUnsignedThinking`
endpoint the encoder then re-emitted that block as
`{ type: "thinking", signature: "" }`. An empty signature is rejected by
the signature-enforcing backend with
`400 messages.1.content.0: Invalid signature in thinking`.
Exclude ZenMux from `replayUnsignedThinking` (via a new `zenmux` host
classifier covering the `zenmux` provider id and the `zenmux.ai` url marker)
so unsigned/stripped thinking degrades to text exactly like the official
Anthropic API — wire-valid and lossless of the tool_use pairing. Z.AI /
DeepSeek / other 3p reasoning endpoints (#2005) and cross-model preservation
(#2257/#2265) are unaffected.
Tests:
- packages/catalog/test/anthropic-zenmux-signing-compat.test.ts: zenmux
(provider id and url marker paths) -> replayUnsignedThinking false; generic
3p reasoning -> true; official -> false. Fails before / passes after.
- packages/ai/test/anthropic-zenmux-checkpoint-thinking-signature.test.ts: a
derived-compat zenmux sonnet 5 model never emits an empty-signature
thinking block for a historical checkpoint turn (demotes to text, keeps
tool_use), and still replays a clean signed historical thinking block
natively.
Fixes#4192
- Persist signed message blocks (`text`, `thinking`, `toolCall`) and encrypted reasoning payloads verbatim during session serialization instead of clearing or truncating them.
- Preserve signature keys instead of replacing them with empty strings when they exceed persistence size limits.
- Exempt official first-party OpenAI and Anthropic API endpoints from the leaked-thinking stream healing wrapper to prevent misfires on legitimate visible text fences.
- Exempts official Anthropic, OpenAI, and OpenAI-Codex base URLs from stream-healing logic.
- Adds `isLeakedThinkingHealExempt` to detect official API hosts and avoid processing structured thinking blocks.
- Incorporates Anthropic Foundry toggle detection to accurately route enterprise gateway URLs.
- Refactors stream dispatch exits to conditionally apply `healLeakedThinking`.
- Adds unit tests verifying that leaked fences are preserved for official endpoints but healed on third-party gateways.
- Excluded signed `thinking` blocks and `redactedThinking` blobs from size-based persistence truncation.
- Preserved signature-bound reasoning verbatim to prevent provider validation failures on session replay.
- Maintained normal truncation behavior for unsigned thinking and standard text blocks.
- Stopped calling the consuming `getSteeringMessages` getter during mid-batch interrupt polls to prevent stranding or dropping messages before they reach the injection boundary.
- Skip subsequent steering checks in the poll loop once an interrupt has already triggered.
- Added a regression test to ensure legacy steering remains queued until the injection boundary when no non-consuming peek exists.
The branch-scan rehydrator only rebuilt `#lastCompletedRewind` and wiped
`#checkpointState` unconditionally at entry — so a branch whose latest
checkpoint had not yet been rewound came back with neither an active
checkpoint nor completed-rewind guidance. Reloading such a session (or
`switchSession()` on the same file) made the next `rewind` fail with
"No active checkpoint" even though the checkpoint entry was still the
branch leaf.
Extended the walker to also track the last unresolved checkpoint entry
and, when the branch ends without a rewind-report, seed `#checkpointState`
from that entry (id, `details.startedAt`) so `rewind` can complete
normally. Renamed the method to `#rehydrateCheckpointRewindState` to
reflect the widened responsibility and added a regression test that
truncates the branch to the checkpoint entry, resumes into a fresh
`AgentSession`, and calls `rewind` end-to-end.
Fixes#4187
- Refactored `renderSubagentHudLines` to use `renderTreeList` with dim connectors and a single-space indentation shift.
- Adjusted budget limits to account for the new layout wrapping and tree-list padding.
Cleared checkpoint rewind runtime state when starting new sessions or creating branch sessions so stale completed-rewind guidance cannot leak into unrelated contexts.
Added regression coverage for /new and branch reset paths.
Fixes#4187
Reconstructed the completed rewind marker from the active branch so resumed sessions keep repeat-rewind recovery guidance.
Covered resume rehydration with the checkpoint rewind branch regression test.
Fixes#4187
llama-server in router/preset mode advertises each preset via /v1/models,
but meta.n_ctx / n_ctx_train are only merged in after the preset's child
instance loads. The router-level /props returns a dummy n_ctx: 0. As a
result every unloaded preset fell through to DISCOVERY_DEFAULT_CONTEXT_WINDOW
(128000), and picking a preset from /model kept surfacing 128k in the
status bar regardless of the configured --ctx-size — a restart didn't
help because discovery repopulated the cache from the same broken chain.
Parse each entry's status.args (rendered CLI vector) for --ctx-size or
-c, and fall back to ctx-size = N in status.preset (INI). Positive values
slot between runtimeContextWindow and serverMetadata in the resolution
chain so a running child's live n_ctx still wins; --ctx-size 0 ("loaded
from model") is correctly skipped so we don't publish 0.
The same fallback wires through discoverLlamaCppModelRuntimeMetadata so
the refresh triggered by /model uses the configured window even before
the child spawns.
Fixes#4190
Reconstructed the completed rewind marker from the active branch so resumed sessions keep repeat-rewind recovery guidance.
Covered resume rehydration with the checkpoint rewind branch regression test.
Fixes#4187
- Refactored the mid-run todo nudge to trigger on mutating tools (bash, eval, edit, write, ast_edit) rather than overall tool turns.
- Simplified the nudge prompt template to a concise, non-escalating reminder.
- Migrated nudge messages from "developer" role with public events to a hidden "custom" role that is excluded from the TUI and transcript.
- Introduced a separate per-cycle reminder cap of 2 to decouple mid-run hints from the user-visible stop-time escalation budget.
- Avoided triggering the todo nudge when read-only exploration tools (e.g. grep, read, glob, lsp) or errored results are returned.
Wrapped retained rewind reports with completion guidance so the post-rewind turn knows the checkpoint is closed.
Added repeat-rewind recovery errors and regression coverage for both the retained context and no-active-checkpoint path.
Fixes#4187
- Reverted to flushing only on the last file write or explicitly on early failure paths within `apply_patch` multi-file operations.
- Refactored error counting logic within single path entries to use clean booleans instead of numeric counters.
- Replaced custom preview capping logic in task progress rendering with `capPreviewLines` and added an option to hide the expand hint.
- Instructed the tester agent to never write assertions or tests for default values, configurations, or fallback properties.
- Allowed the tester agent to skip writing tests entirely if the changes are trivial, already covered, or if any new tests would be worthless.
- Required the deletion of existing default-testing assertions or entire tests when modifying code that currently contains them.
- Bypasses frame array allocation and merging during render except when hosted under ConPTY.
- Reuses the active window array slice for DECCARA calculations to avoid redundant array copies.
- Simplifies the `stdin-buffer` CSI sequence scanner by removing unnecessary resume search logic and redundant SGR mouse validation regexes.
- Extracted `AskEditor` into a standalone, key-bound component keyed by `reqId`.
- Preserved user-typed draft state across duplicate incoming host requests of the same ID.
- Reset the editor draft state only when a genuinely new request ID arrives.
- Simplified `Composer` autosizing logic and state orchestration by isolating editor-specific hooks.
- Adjusted integration tests to match HTML structure changes of the submit action.
- Introduced the `CollabGuestUiResult` type to distinguish between answers and unavailable states during guest UI requests.
- Updated the remote dialog race logic to ignore unavailable guest results and fallback to the local host dialog.
- Centralized the `FakeWebSocket` and `InMemoryRelay` test infrastructure into a shared test helpers file.
- Cleaned up duplicate mock implementations and updated existing test suites to utilize the shared in-memory relay helpers.
- Projected full tool-call arguments down to a compact summary containing only `command` and `path`.
- Truncated summarized argument fields to 200 characters to prevent inflating session log sizes.
- Replaced routine clean session disposal warnings with debug logs to reduce noise.
- Streamlined debug context in assistant message removal and agent continuation skip paths.
- Extracted duplicate user-facing compaction warning strings into a helper function.
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
- Introduced `RpcShutdownCoordinator` to track background tasks and manage deferred shutdowns safely.
- Guaranteed all background bash task response frames are fully written before the process exits.
- Re-checked shutdown requests automatically as each tracked background task settles.
- Latched the shutdown sequence to prevent concurrent execution from duplicate triggers.
- Updated `mock-rpc-agent` to consume stdin via an async iterator to match standard behavior.
- Added comprehensive unit tests in `rpc-input-frame.test.ts` covering background task coordination.
- Introduced JSON depth tracking to ensure string keys are only extracted at the top level.
- Prevented nested keys matching designated streaming keys from being prematurely or incorrectly captured.
- Added comprehensive unit tests validating nested key exclusion and correct top-level extraction order.
- Introduces `sanitizeErrorLine` to collapse newlines, replace tabs, and shorten absolute paths.
- Truncates remote error and notice text to the available terminal width to prevent layout breaking.
- Corrects a potential runtime exception in status-line by safely accessing JSON stringified length.
- Introduced the `allowCreateOverwrite` option to permit `op: "create"` to replace existing files.
- Enabled `allowCreateOverwrite` specifically for the JSON-based `patch` edit mode to support full-file restructures.
- Maintained the strict non-overwriting behavior for Codex `apply_patch` envelope-based file additions.
- Configured patch diff previews to respect the configured overwrite permission during streaming.
- Fixed an issue where stopping a multi-file patch application early skipped flushing the active LSP writethrough batch.
- Introduced the `task.softRequestBudgetNotice` boolean setting to opt into budget steering notices.
- Disabled the wrap-up steering notice by default when a subagent crosses its soft request budget.
- Maintained the 1.5x graceful abort safety guard regardless of whether the steering notice option is enabled.
- Updated the settings schema to document the conditional steering notice behavior.
- Added `discoveryFetch` utility to wrap global fetch with `NODE_EXTRA_CA_CERTS` support.
- Consolidated SSL-stable fetch overrides across all catalog discovery models.
- Replaced direct `wrapFetchForExtraCa` calls with the unified `discoveryFetch` helper.
- Patched models.dev metadata and Ollama native probes to support private CA gateways.
- Renamed `requiresJuiceZeroHack` to `requiresReasoningSuppressionPrompt` across the catalog codebase.
- Dropped legacy non-msg string signature IDs during historical replay rebuilding when reasoning items are missing.
- Maintained legacy signature IDs in rebuilding fallback history when paired with matching reasoning items.
- Cleaned up obsolete GPT-5 reasoning-disable assertions from the test suite.