The consume-once source map kept entries for graph modules the initial
import never loaded (modules only reached via lazy dynamic imports).
Their first import - possibly long after load, and after an on-disk
edit - was served the boot-time snapshot instead of current file
content, and the unconsumed sources stayed in the plugin closure for
the process lifetime.
Clear the map once the entry import settles: everything Bun loaded at
startup was already consumed (keeping the read-once win), and anything
left must be read at its actual import time, matching pre-dedup
behavior for lazy modules. The new regression test passes on the
pre-dedup baseline and fails on the unfixed dedup.
isAllCapsWord matched any multi-letter token without a lowercase letter,
so CJK tokens registered as ALL-CAPS words: two adjacent ones marked the
whole source shouty and silently disabled acronym restoration for every
non-Latin-script message (e.g. '修复 CNPG 集群故障' kept 'Cnpg'). Require an
actual uppercase letter; cased-script shout detection is unchanged.
The PR's changelog, doc comment, and prompt examples all name ETL as a
restored acronym, but the review-response narrowing (vowel heuristic +
allowlist) silently dropped it: ETL bears a vowel and was not listed.
Add it to COMMON_TITLE_ACRONYMS and pin it in the allowlist test.
Narrow acronym restoration so plain all-caps English words such as FIX
and WORK do not get restored when the model naturally capitalizes the
first title word. Restorable all-caps source tokens now need a stronger
acronym signal: a common technical acronym allowlist, digits, or a
consonant-only shape.
This keeps CNPG, ETL, JWT, SQL, and API restoration while preserving the
anti-shout behavior for single emphatic words.
Fixes#4220
- Added an enhanced speech pipeline that utilizes small models to rewrite text for natural language synthesis.
- Implemented `SpeechEnhancer` and `BlockAccumulator` to manage fence-aware text streaming and paragraph splitting.
- Configured a new `speech.enhanced` setting to toggle between mechanical and enhanced vocalization modes.
- Resolved `EPIPE` rejections during speech playback by ensuring stream flushes and suppressing stop-related errors.
reconcileTitleCasing now maps ALL-CAPS source tokens (CNPG, API, ETL,
JWT) into an acronyms table and restores them when the model produces a
plain title-cased artifact (Cnpg). Restoration is disabled when the
source is shouty (>=2 consecutive multi-letter ALL-CAPS tokens like
"FIX the BUG NOW" or "ALL ERROR HANDLING"), and lowercase model output
is left alone so isolated single-word emphasis (WORK -> work) is never
re-shouted.
The three title system prompts (title-system.md, title-system-marker.md,
tiny-title-system.md) also gained an explicit instruction to preserve
ALL-CAPS acronyms verbatim, so a competent model short-circuits via the
verbatim set before post-processing kicks in.
Fixes#4220
Removed the live-context guard that let default model selection persist a new role without changing the active session model. The next prompt's compaction path now owns oversized-context recovery after a switch.
Fixes#4219
- Added `SpeakableStream` to strip markdown noise, silence code blocks and tables, normalize links and paths, and emit sentence/clause segments.
- Reworked `Vocalizer` to segment assistant deltas in the parent process, lazily open TTS streams, idle-flush partial thoughts, and chain playback sessions.
- Added gapless streaming playback with ffmpeg/sox backends, ducking-aware pacing, fallback file playback, and immediate stop handling.
- Added IPC `sendAndFlush` support and used it in the TTS worker so audio chunks drain before blocking ONNX inference resumes.
- Added speakable-stream coverage for markdown filtering, segmentation latency, idle flushing, and forced long-segment splits.
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
Loaded cached runtime extension provider catalogs before deferred model resolution so dynamic-only providers can satisfy cold-start --model and session resume selection from models.db.\n\nFixes #4216
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.
Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).
Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.
Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
`discoverExtensionPaths` loaded the extension-module capability across all
registered providers (native, claude, codex, gemini, opencode) and then
discarded every item whose `_source.provider !== "native"`. Four foreign
directory walks ran on every session startup only for their results to be
dropped — worst on Windows.
Scope the load to the native provider via the existing `LoadOptions.providers`
filter: `loadCapability(extensionModuleCapability.id, { ...loadOptions,
providers: ["native"] })`. The hook-capability load in the same function is
unchanged (hooks legitimately span providers). Output is identical.
Regression test in `extensions-discovery.test.ts` spies on each extension-module
provider's `load()` and asserts only "native" is invoked.
Fixes#4198
collectExtensionModules already reads every own-source module's text to scan imports; the onLoad rewrite hook then re-read each file. Return a Map<path,source> from the collector and serve it consume-once in onLoad (delete on first hit, disk fallback on miss), so transitive modules are read once and the entry still re-reads on its ?mtime re-import. Adds a regression test asserting exactly one read per graph module.
Closes#4196
- Persist signed message blocks (`text`, `thinking`, `toolCall`) and encrypted reasoning payloads verbatim during session serialization instead of clearing or truncating them.
- Preserve signature keys instead of replacing them with empty strings when they exceed persistence size limits.
- Exempt official first-party OpenAI and Anthropic API endpoints from the leaked-thinking stream healing wrapper to prevent misfires on legitimate visible text fences.
- Excluded signed `thinking` blocks and `redactedThinking` blobs from size-based persistence truncation.
- Preserved signature-bound reasoning verbatim to prevent provider validation failures on session replay.
- Maintained normal truncation behavior for unsigned thinking and standard text blocks.
- Stopped calling the consuming `getSteeringMessages` getter during mid-batch interrupt polls to prevent stranding or dropping messages before they reach the injection boundary.
- Skip subsequent steering checks in the poll loop once an interrupt has already triggered.
- Added a regression test to ensure legacy steering remains queued until the injection boundary when no non-consuming peek exists.
The branch-scan rehydrator only rebuilt `#lastCompletedRewind` and wiped
`#checkpointState` unconditionally at entry — so a branch whose latest
checkpoint had not yet been rewound came back with neither an active
checkpoint nor completed-rewind guidance. Reloading such a session (or
`switchSession()` on the same file) made the next `rewind` fail with
"No active checkpoint" even though the checkpoint entry was still the
branch leaf.
Extended the walker to also track the last unresolved checkpoint entry
and, when the branch ends without a rewind-report, seed `#checkpointState`
from that entry (id, `details.startedAt`) so `rewind` can complete
normally. Renamed the method to `#rehydrateCheckpointRewindState` to
reflect the widened responsibility and added a regression test that
truncates the branch to the checkpoint entry, resumes into a fresh
`AgentSession`, and calls `rewind` end-to-end.
Fixes#4187
- Refactored `renderSubagentHudLines` to use `renderTreeList` with dim connectors and a single-space indentation shift.
- Adjusted budget limits to account for the new layout wrapping and tree-list padding.
Cleared checkpoint rewind runtime state when starting new sessions or creating branch sessions so stale completed-rewind guidance cannot leak into unrelated contexts.
Added regression coverage for /new and branch reset paths.
Fixes#4187
Reconstructed the completed rewind marker from the active branch so resumed sessions keep repeat-rewind recovery guidance.
Covered resume rehydration with the checkpoint rewind branch regression test.
Fixes#4187
llama-server in router/preset mode advertises each preset via /v1/models,
but meta.n_ctx / n_ctx_train are only merged in after the preset's child
instance loads. The router-level /props returns a dummy n_ctx: 0. As a
result every unloaded preset fell through to DISCOVERY_DEFAULT_CONTEXT_WINDOW
(128000), and picking a preset from /model kept surfacing 128k in the
status bar regardless of the configured --ctx-size — a restart didn't
help because discovery repopulated the cache from the same broken chain.
Parse each entry's status.args (rendered CLI vector) for --ctx-size or
-c, and fall back to ctx-size = N in status.preset (INI). Positive values
slot between runtimeContextWindow and serverMetadata in the resolution
chain so a running child's live n_ctx still wins; --ctx-size 0 ("loaded
from model") is correctly skipped so we don't publish 0.
The same fallback wires through discoverLlamaCppModelRuntimeMetadata so
the refresh triggered by /model uses the configured window even before
the child spawns.
Fixes#4190
Reconstructed the completed rewind marker from the active branch so resumed sessions keep repeat-rewind recovery guidance.
Covered resume rehydration with the checkpoint rewind branch regression test.
Fixes#4187
- Refactored the mid-run todo nudge to trigger on mutating tools (bash, eval, edit, write, ast_edit) rather than overall tool turns.
- Simplified the nudge prompt template to a concise, non-escalating reminder.
- Migrated nudge messages from "developer" role with public events to a hidden "custom" role that is excluded from the TUI and transcript.
- Introduced a separate per-cycle reminder cap of 2 to decouple mid-run hints from the user-visible stop-time escalation budget.
- Avoided triggering the todo nudge when read-only exploration tools (e.g. grep, read, glob, lsp) or errored results are returned.
Wrapped retained rewind reports with completion guidance so the post-rewind turn knows the checkpoint is closed.
Added repeat-rewind recovery errors and regression coverage for both the retained context and no-active-checkpoint path.
Fixes#4187