The beam backend never invoked the embedding pipeline during normal
operation: `remember()`/`rememberBatch()`/`updateWorking()` skipped `embed()`
entirely and `recall()`/`recallEnhanced()` never called `embedQuery()` on
the query text. As a result `memory_embeddings` stayed empty in every
deployment and recall silently degraded to FTS-only regardless of the
configured provider (fastembed, OpenAI-compatible API, custom).
- Added `scheduleEmbedding` on `beam.pendingExtractions` (mirroring
`scheduleFactExtraction`) and wired it from `remember`, `rememberBatch`,
`updateWorking`, and `consolidateToEpisodic`. Writes
`INSERT OR REPLACE INTO memory_embeddings(memory_id, embedding_json, model)`
with the active runtime-options model, captured before the AsyncLocalStorage
scope exits and re-entered inside the task.
- Auto-derived `queryEmbedding` inside `recall()` via `embedQuery(query)` when
the caller did not pass one. `queryEmbedding: null` is preserved as the
explicit FTS-only opt-out; `undefined` triggers auto-derive.
- Propagated `queryEmbedding` through `Mnemopi`'s `toRecallOptions` so the
facade no longer strips the override on the way to the beam layer.
- Made `Mnemopi.recall`/`recallEnhanced`/`search`/`query`, the module-level
exports, `BeamMemory.recall`/`recallEnhanced`, the free `recall`/`recallEnhanced`,
and `orchestrateRecall` async. MCP `handleToolCall`/`callToolJson`/`handleJsonRpc`
follow suit so the recall handler can await.
- Fixed `withBeam`/`withSharedBeam` to defer `beam.close()` until the async
handler resolves; otherwise the new async recall hit
`RangeError: Cannot use a closed database`.
- Updated CLI, MCP entrypoints, coding-agent `MnemopiSessionState`, and every
affected test to await the new shapes.
Verified with a new regression suite (`test/issue-1832-embedding-population.test.ts`)
exercising both ends of the bug: empty `memory_embeddings` and zero
`dense_score`.
Fixes#1832
The SSH tool renderer fed the raw command into renderStatusLine's
description, so any newline in the remote command expanded the
single-line tool header — the bordered output block then opened
mid-command and the rest of the SSH cell rendered against a broken
frame.
The renderer now keeps only [host] in the header and renders the
full command (with a dim $ prefix and tab sanitization) as its own
framed section above Output, matching the bash renderer's shape.
renderStatusLine also flattens CR/LF in title, description, meta,
and badge labels so no future caller can accidentally produce a
multiline tool header.
Covered by:
- test/tools/ssh-render.test.ts: multiline command stays out of
the header and every command line is present in the body, for
both renderCall and renderResult.
- test/tui/status-line-newline-guard.test.ts: embedded LF, CRLF,
and lone CR in description/meta are flattened to spaces.
Fixes#1828
- Added `loadOverallPlanReference` to resolve a session plan reference from local storage and skip empty or missing files.
- Updated task execution to read the active plan reference (except in plan mode) and pass it into each spawned subagent.
- Extended the subagent system prompt and session SDK/tools plumbing so subagents receive and render the approved plan path and contents.
- Added `selectionMarker`, `checkedIndices`, and `markableCount` options to render radio/checkbox glyphs per row.
- Moved checkbox rendering from inline label prefixes into the selector component.
- Kept trailing control rows like "Other"/"Done" on the plain cursor.
The `search` tool required `paths` via a strict `union([string,
array.min(1)])`, so any call that omitted `paths` (or passed an empty
array) was rejected at schema validation with `paths: Invalid input`
and never ran — turning a recoverable "no scope given" into a fatal
tool-call failure.
Make `paths` optional and default an omitted/empty value to the
workspace root (`.`), matching the documented prompt contract. Behavior
is unchanged when paths are supplied. Adds regression coverage for the
omitted and empty-array cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Renamed `TodoWriteTool` to `TodoTool` and its source/prompt files.
- Updated tool registration, schema, renderers, and gating to `todo`.
- Adjusted cursor provider native tool names and tests to match.
- Renamed strike-animation constants and `todo-error-reminder` type.
- Extended line selector parsing in path-utils to accept `..` as a forgiving alias for `-`, with `..` normalized during chunk parsing.
- Updated selector-matching regexes for file selectors and internal URL selectors to recognize alias-style ranges in single chunks and comma-separated lists.
- Added tests covering `N..M`, `N..`, mixed separators, inverted-range errors, and path-splitting behavior with `:N..M` selectors and `foo:../bar.ts` paths.
- Added a shared `renderAnswerOptions` helper that redrew answered options with markers, custom input, and cancellation state in-place.
- Set `mergeCallAndResult` to true so ask prompts now keep their question form visible while showing the final selection.
- Kept selection rendering consistent by preserving markers and option ordering while reusing the same layout for completed answers.
- Updated completion and ask notifications to pass structured payloads.
- Updated completion and ask notifications to include explicit message metadata fields.
- Adjusted abort-guard and retry-capability tests for the new notification/options behavior.
- Rendered single-choice questions with circular radio glyphs instead of checkboxes.
- Kept rectangular checkboxes for multi-select questions.
- Added radio.selected/unselected symbols to unicode, nerd-font, and ASCII presets.
- Added optional `AgentTool.matcherDigest(args)` hook so tools can expose plain source text instead of wire-encoded arguments to TTSR rule matchers.
- Implemented `matcherDigest` on edit (all modes: hashline, patch, apply_patch, replace) and write tools, stripping patch prefixes and JSON escaping.
- Added `TtsrManager.checkSnapshot()` to replace the scoped buffer with a tool digest rather than appending raw deltas.
- Fixed TTSR conditions never matching streamed edit/write calls whose wire format obscured real content.
Kept uncached sort-by-mtime glob traversal bounded to maxResults and emitted onMatch callbacks only for returned matches so broad find scans cannot grow parent memory independently of the limit.
Fixes#1761
Surface the hidden discoverable built-in tool names (write, find, search, lsp, task, ...) in the search_tool_bm25 description when tools.discoveryMode is "all", so a model can form a targeted BM25 query by name instead of guessing or falling back to shell. mcp-only mode is unchanged (no built-ins advertised) and the total-tools count still includes them.
- Expanded path parsing to split top-level comma, semicolon, and whitespace entries.
- Updated find/search and scope resolution to apply delimiter expansion before path-spec validation.
- Updated read tool fallback to try split path parts before raising missing-path errors.
- Coerced bare string arguments into singleton arrays for array-typed schemas.
- Changed `AgentOutputManager` to use requested names verbatim, adding `-2`/`-3` suffixes only on repeats (e.g. `Anna`, `Anna-2`).
- Renamed main agent id from `0-Main` to `Main`; nested ids now use dot notation without numeric prefix (e.g. `Parent.Child`).
- Updated task widget to render dotted hierarchy as `Parent>Child` breadcrumb without leading index.
- Resume scan now tracks seen names instead of a counter to avoid clobbering prior outputs.
- Added ROW_COUNT_PROBE_CAP to limit rows scanned when counting tables, preventing JS thread freezes on large databases.
- Used sqlite_stat1 estimates for tables exceeding the cap; exact counts only for provably small tables.
- Introduced TableRowCount type with exact/estimate/atLeast variants reflected in rendered output.
- Computed screenshot destination extensions from the MIME type of bytes being written.
- Updated auto-generated paths in screenshot and temp directories to use the matching extension.
- Extracted TUI rendering from `eval.ts` into a dependency-light `eval-render.ts` to break the circular initialization chain.
- `renderers.ts` now imports `evalToolRenderer` from `eval-render` directly, avoiding re-entry into the root barrel while `eval.ts` is still initializing.
- `eval.ts` re-exports `evalToolRenderer` and `EVAL_DEFAULT_PREVIEW_LINES` for backward compatibility.
- Increased first-event timeout test budget from 50ms to 5000ms to prevent CI scheduler jitter from tripping the watchdog on success cases.
- Moved EvalBackendsAllowance and related functions to a dedicated eval-backends.ts module.
- Replaced ad-hoc brokenShellSessions tracking with a quarantineShellSession helper that also awaits the abort cleanup promise.
- Applied quarantine on timeout and cancellation paths, not just errors.
- Added `DiagnosticsLedger` to track diagnostics already surfaced per file, suppressing repeats within a session.
- Wired dedup into both edit and write tools via a `transformDiagnostics` hook on the writethrough pipeline.
- Added `lsp.diagnosticsDeduplicate` setting (default: true) to control the behavior.
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.
Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.
- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
- Extended ResolveContext / WriteContext with localProtocolOptions so the
internal-URL router can thread the calling session's local-root mapping
through to handlers.
- LocalProtocolHandler.resolveOptions now prefers context.localProtocolOptions
before consulting the process-global override or the first main-kind session
in AgentRegistry, fixing multi-session ACP hosts (cmux) where reads of
local://PLAN.md were routing to a sibling session's artifacts dir even
though plan-mode writes succeeded against the calling session.
- read, find, ast_grep, ast_edit, and search now thread
this.session.localProtocolOptions into the router so local://, memory://,
agent://, and other handlers see the right caller.
- Added regression tests covering the override-vs-context priority and the
ENOENT-against-caller-root path.
Fixes#1608
Guarded find renderer path summaries so raw pre-validation string paths render instead of throwing. Added coverage for pending, fallback, empty, and detailed result render paths.\n\nFixes #1622
Included the reviews field in comments-enabled PR view fetches so pr:// output can show formal review submissions and approvals.
Added protocol coverage that emulates gh --json field selection before asserting rendered approval output.
Fixes#1600
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
- Added +Nk/+Nm turn-budget parsing with whitespace-boundary matching, multipliers, and hard `!` indicator.
- Added per-turn budget lifecycle plus APIs (`getTurnBudget`, `recordEvalSubagentUsage`) and hard-cap checks in eval runs.
- Added hard budget observability in eval preludes and docs by exposing `budget.hard` and documenting ceiling modes.
- Fixed streaming preview stutter with max-row tracking and padding, with tests for preview height and budget parsing.
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
- Stored mcpManager and localProtocolOptions on ToolSession so nested subagents inherit them without relying on process-global singletons.
- TaskTool now uses the session's localProtocolOptions and mcpManager when spawning sub-tasks, falling back to defaults if absent.
F1 (coding-agent): close the withFileLock mkdir-vs-writeLockInfo race that
let a losing contender wipe the winner's freshly-created lock directory.
Every lock now carries a per-process UUID token; releaseLock verifies the
token before fs.rm, and isLockStale no longer treats an info-less but
fresh dir (or a dir that vanished mid-check) as stale.
F4 (coding-agent): sanitize tabs and truncate oversized error strings in
formatErrorMessage so error renderings that embed file content
(apply_patch, hashline, etc.) cannot break terminal alignment or
overflow the line width.
F7 (ai): support named-tool routing on Google providers. Widens
GoogleSharedStreamOptions.toolChoice and GoogleGeminiCliOptions.toolChoice
to accept { mode: 'ANY'; allowedFunctionNames }. mapGoogleToolChoice
now converts ToolChoice { type: 'tool'|'function', name } to the wire
shape (mirroring mapAnthropicToolChoice). buildGoogleGenerateContentParams
and the gemini-cli request serializer honor the allow-list.
- Expanded search scope handling for virtual multi-file targets and executed grep only when searchable paths existed.
- Merged `searchVirtualResources` results into main output and rendered internal URL matches as accent lines.
- Updated grouped-file output to detect URL-like paths and keep full URL headers for root grouping.
- Documented URL path/range behavior and added tests for doc routing, missing-content errors, and `omp://` expansion.
- Added virtual internal URL path resolution in `SearchTool` via `InternalUrlRouter` for in-memory search.
- Added `omp://` root expansion so `search` resolves and scans each completion target.
- Fixed `SearchTool` handling of internal URLs without `sourcePath` by returning virtual matches instead of `Path not found`.
- Extended `edit-renderer.test.ts` coverage for normalizing raw streamed text in custom text renderers.
- Added clockwise sweeping dark segment animation to output block borders while bash/eval tool calls are pending/running.
- Changed bash renderCall to immediately render a full bordered block instead of a one-liner status preview, so silent commands show the framed block for their entire runtime.
- Added shimmerEnabled() helper and wired animate flag through OutputBlockOptions, CodeCellOptions, and shell/eval renderers.
- Updated read tool prompt to show alternate image description when inspect_image is disabled.
- Passed INSPECT_IMAGE_ENABLED flag into prompt rendering context.
- Added test verifying description omits inspect_image references when disabled.
- Added todo_write strike-frame animation timing and updated execution flow after completion finalizes.
- Added strike animation cancellation in cleanup to clear todo timer and reset frames when spinner is idle.
- Removed todo-closing state and timeout handling from interactive mode and simplified empty todo-list rendering.
- Reworked todo-write to compute completion transitions, track completedTasks, and render strike-through frames.
- Added test coverage for completedTasks, theme setup, and strike-through progression at hold-frame thresholds.