- Described "auto" default gating MCP tools past 40-tool threshold.
- Noted late resolution in createAgentSession after registry exists.
- Updated legacy mcp.discoveryMode mapping to MCP-only.
- Made "auto" the default, hiding MCP tools past 40-tool threshold.
- Centralized discovery mode resolution in shared mode helper.
- Activated search tool in createAgentSession once full registry exists.
- Added image-reference rendering to make `[Image #N]` placeholders clickable in chat.
- Added MIME-aware image blob materialization with extensioned sidecar paths.
- Added clickable path, line, and URL hyperlinks for read, search, and fetch outputs.
- Hardened OSC8 hyperlink emission with URI validation and control-byte/idempotency checks.
The no-selector run_watch guard was using resolveDefaultRepoMemoized via tryResolveCurrentRepo, so a long-lived process could validate against a stale cwd-to-repo cache entry after the checkout or GitHub remote at that path changed. That allowed the guard to trust the current HEAD for an explicit repo based on the old cached repository.
Add a fresh best-effort cwd repo lookup for safety checks and use it before deriving branch/HEAD. Cached lookup remains for search default scoping where stale data only affects a convenience fallback. Added a regression test that populates the cache, changes the mocked repo at the same cwd, and asserts run_watch rejects before issuing API calls.
Refs #1949#1951
resolveGitHubRepo rejected calls that supplied both an explicit repo and a full Actions run URL when the two owner/repo slugs differed only by casing. GitHub repository paths are case-insensitive, so this was the same class of false mismatch as the cwd guard fixed earlier.
Compare repo slugs through a shared ASCII case-insensitive helper and use it for both the run-URL consistency check and the cwd guard. Added a regression test for repo=cagedbird043/cxf with a run URL under CagedBird043/CXF.
Refs #1949#1951
GitHub owner/repo slugs are case-insensitive; `gh repo view` returns
the canonical casing while callers may pass any casing. The new guard
used strict equality, so a caller in the correct repo who typed
`owner/repo` while the canonical form was `Owner/Repo` was forced to
pass a redundant `branch`/`run` selector. Normalize both sides via
toLowerCase() before deciding the cwd is a different repository.
Regression test covers the casing-only match.
Refs #1949#1951
executeRunWatch passed undefined for the explicit `repo` to
resolveGitHubRepo, so a call like
`{op: "run_watch", repo: "owner/cxf", branch: "main"}` from a nested
or umbrella workspace silently fell through to `gh repo view` in cwd
and streamed `watching <sha> on <cwd-repo>` against the wrong
repository.
Route params.repo through resolveGitHubRepo so the explicit owner/repo
wins over both cwd inference and run-URL inference. When no `branch`
or `run` selector is given, refuse to derive the watched commit from
`git HEAD` unless the cwd actually points at the resolved repo —
otherwise raise a ToolError telling the caller to pass `branch` or
`run` instead of silently rebinding to an unrelated commit.
Also deduped resolveSearchRepoScope's best-effort cwd resolution into a
shared tryResolveCurrentRepo helper used by the new guard.
Fixes#1949
- Parsed zip metadata via central directory and lazy ranged reads.
- Inflated member contents only when a specific entry is read.
- Prevented large or corrupt zips from freezing directory reads.
- Added selectorLineRanges to extract ranges from raw/conflicts selectors.
- Routed internal URLs through URL-aware splitter in content search.
- Treated display-mode selectors as whole-resource searches instead of rejecting.
AgentSession now stores the same scoped AsyncJobManager reference that tools receive: owning top-level sessions use their constructed manager, subagents inherit the parent's manager, and secondary in-process top-level sessions get no manager when a singleton is already live.
getAsyncJobSnapshot and ACP delivery drains now use that scoped manager instead of AsyncJobManager.instance(), so secondary sessions cannot report or drain the primary session's background jobs. The regression test covers a secondary session created while the primary has a Main-owned running job.
Per PR review on #1926: a secondary in-process top-level createAgentSession() that exposes bash/task/job tools would still call AsyncJobManager.instance() at execute time, register on the primary's manager, and have the primary's onJobComplete enqueue results into the primary's yieldQueue — corrupting the owning session's conversation.
ToolSession now carries an asyncJobManager reference scoped to its session: the constructed manager for top-level sessions, the inherited singleton for subagents (so their bash/task completions still flow into the spawning conversation as before), and undefined for secondary in-process top-level sessions that found a singleton already installed. bash, task, and job tools resolve the manager through ToolSession instead of the process-global singleton, so a secondary session whose tools attempt async work fails fast with the standard "Async job manager unavailable" error instead of contaminating the primary.
Mid-session edits to hindsight.bankId / bankIdPrefix / scoping kept the
active HindsightSessionState pinned to the bank selected at session
start, so retain/recall/reflect calls landed in the stale bank. Settings
hooks now fire onHindsightScopeChanged; the backend rebuilds the
primary state against the recomputed scope, disposing the previous one
after flushing its queue so queued tool-initiated retains still land in
the bank they were enqueued for.
Also:
- Renamed ensureBankMission to ensureBankExists. The old version
skipped creation entirely when bankMission was blank, so the first
mental-model POST (auto-seed) could land against a never-PUT bank.
Bank creation is now idempotent and unconditional, and runs before
mental-model bootstrap.
- Fixed AgentSession.dispose to flush the retain queue BEFORE clearing
the session state pointer. Reversed, HindsightRetainQueue.#doFlush's
identity guard would see the cleared pointer and drop the spliced
batch with a 'session vanished' warning.
- Snapshotted hindsightScopeCallbacks before iterating because each
rebuild subscribes a fresh callback inside the same fire; iterating
the live Set would spin.
Fixes#1902
Handled omp://docs as an embedded documentation search root alongside omp://.
Added regression coverage for searching embedded docs through the docs alias.
Fixes#1898
Route '/data/workspaces/can1357__oh-my-pi__1863/.omp-session/2026-06-04T13-36-11-717Z_019e92d9-3fc5-7000-a66f-15cad94b7e75/local' plan artifacts through OMP's session-local storage instead of the editor writeTextFile bridge, preserving ACP bridge routing for regular editor-visible files.
Fixes#1863
The beam backend never invoked the embedding pipeline during normal
operation: `remember()`/`rememberBatch()`/`updateWorking()` skipped `embed()`
entirely and `recall()`/`recallEnhanced()` never called `embedQuery()` on
the query text. As a result `memory_embeddings` stayed empty in every
deployment and recall silently degraded to FTS-only regardless of the
configured provider (fastembed, OpenAI-compatible API, custom).
- Added `scheduleEmbedding` on `beam.pendingExtractions` (mirroring
`scheduleFactExtraction`) and wired it from `remember`, `rememberBatch`,
`updateWorking`, and `consolidateToEpisodic`. Writes
`INSERT OR REPLACE INTO memory_embeddings(memory_id, embedding_json, model)`
with the active runtime-options model, captured before the AsyncLocalStorage
scope exits and re-entered inside the task.
- Auto-derived `queryEmbedding` inside `recall()` via `embedQuery(query)` when
the caller did not pass one. `queryEmbedding: null` is preserved as the
explicit FTS-only opt-out; `undefined` triggers auto-derive.
- Propagated `queryEmbedding` through `Mnemopi`'s `toRecallOptions` so the
facade no longer strips the override on the way to the beam layer.
- Made `Mnemopi.recall`/`recallEnhanced`/`search`/`query`, the module-level
exports, `BeamMemory.recall`/`recallEnhanced`, the free `recall`/`recallEnhanced`,
and `orchestrateRecall` async. MCP `handleToolCall`/`callToolJson`/`handleJsonRpc`
follow suit so the recall handler can await.
- Fixed `withBeam`/`withSharedBeam` to defer `beam.close()` until the async
handler resolves; otherwise the new async recall hit
`RangeError: Cannot use a closed database`.
- Updated CLI, MCP entrypoints, coding-agent `MnemopiSessionState`, and every
affected test to await the new shapes.
Verified with a new regression suite (`test/issue-1832-embedding-population.test.ts`)
exercising both ends of the bug: empty `memory_embeddings` and zero
`dense_score`.
Fixes#1832
The SSH tool renderer fed the raw command into renderStatusLine's
description, so any newline in the remote command expanded the
single-line tool header — the bordered output block then opened
mid-command and the rest of the SSH cell rendered against a broken
frame.
The renderer now keeps only [host] in the header and renders the
full command (with a dim $ prefix and tab sanitization) as its own
framed section above Output, matching the bash renderer's shape.
renderStatusLine also flattens CR/LF in title, description, meta,
and badge labels so no future caller can accidentally produce a
multiline tool header.
Covered by:
- test/tools/ssh-render.test.ts: multiline command stays out of
the header and every command line is present in the body, for
both renderCall and renderResult.
- test/tui/status-line-newline-guard.test.ts: embedded LF, CRLF,
and lone CR in description/meta are flattened to spaces.
Fixes#1828
- Added `loadOverallPlanReference` to resolve a session plan reference from local storage and skip empty or missing files.
- Updated task execution to read the active plan reference (except in plan mode) and pass it into each spawned subagent.
- Extended the subagent system prompt and session SDK/tools plumbing so subagents receive and render the approved plan path and contents.
- Added `selectionMarker`, `checkedIndices`, and `markableCount` options to render radio/checkbox glyphs per row.
- Moved checkbox rendering from inline label prefixes into the selector component.
- Kept trailing control rows like "Other"/"Done" on the plain cursor.
The `search` tool required `paths` via a strict `union([string,
array.min(1)])`, so any call that omitted `paths` (or passed an empty
array) was rejected at schema validation with `paths: Invalid input`
and never ran — turning a recoverable "no scope given" into a fatal
tool-call failure.
Make `paths` optional and default an omitted/empty value to the
workspace root (`.`), matching the documented prompt contract. Behavior
is unchanged when paths are supplied. Adds regression coverage for the
omitted and empty-array cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Renamed `TodoWriteTool` to `TodoTool` and its source/prompt files.
- Updated tool registration, schema, renderers, and gating to `todo`.
- Adjusted cursor provider native tool names and tests to match.
- Renamed strike-animation constants and `todo-error-reminder` type.
- Extended line selector parsing in path-utils to accept `..` as a forgiving alias for `-`, with `..` normalized during chunk parsing.
- Updated selector-matching regexes for file selectors and internal URL selectors to recognize alias-style ranges in single chunks and comma-separated lists.
- Added tests covering `N..M`, `N..`, mixed separators, inverted-range errors, and path-splitting behavior with `:N..M` selectors and `foo:../bar.ts` paths.
- Added a shared `renderAnswerOptions` helper that redrew answered options with markers, custom input, and cancellation state in-place.
- Set `mergeCallAndResult` to true so ask prompts now keep their question form visible while showing the final selection.
- Kept selection rendering consistent by preserving markers and option ordering while reusing the same layout for completed answers.
- Updated completion and ask notifications to pass structured payloads.
- Updated completion and ask notifications to include explicit message metadata fields.
- Adjusted abort-guard and retry-capability tests for the new notification/options behavior.
- Rendered single-choice questions with circular radio glyphs instead of checkboxes.
- Kept rectangular checkboxes for multi-select questions.
- Added radio.selected/unselected symbols to unicode, nerd-font, and ASCII presets.
- Added optional `AgentTool.matcherDigest(args)` hook so tools can expose plain source text instead of wire-encoded arguments to TTSR rule matchers.
- Implemented `matcherDigest` on edit (all modes: hashline, patch, apply_patch, replace) and write tools, stripping patch prefixes and JSON escaping.
- Added `TtsrManager.checkSnapshot()` to replace the scoped buffer with a tool digest rather than appending raw deltas.
- Fixed TTSR conditions never matching streamed edit/write calls whose wire format obscured real content.
Kept uncached sort-by-mtime glob traversal bounded to maxResults and emitted onMatch callbacks only for returned matches so broad find scans cannot grow parent memory independently of the limit.
Fixes#1761
Surface the hidden discoverable built-in tool names (write, find, search, lsp, task, ...) in the search_tool_bm25 description when tools.discoveryMode is "all", so a model can form a targeted BM25 query by name instead of guessing or falling back to shell. mcp-only mode is unchanged (no built-ins advertised) and the total-tools count still includes them.
- Expanded path parsing to split top-level comma, semicolon, and whitespace entries.
- Updated find/search and scope resolution to apply delimiter expansion before path-spec validation.
- Updated read tool fallback to try split path parts before raising missing-path errors.
- Coerced bare string arguments into singleton arrays for array-typed schemas.
- Changed `AgentOutputManager` to use requested names verbatim, adding `-2`/`-3` suffixes only on repeats (e.g. `Anna`, `Anna-2`).
- Renamed main agent id from `0-Main` to `Main`; nested ids now use dot notation without numeric prefix (e.g. `Parent.Child`).
- Updated task widget to render dotted hierarchy as `Parent>Child` breadcrumb without leading index.
- Resume scan now tracks seen names instead of a counter to avoid clobbering prior outputs.
- Added ROW_COUNT_PROBE_CAP to limit rows scanned when counting tables, preventing JS thread freezes on large databases.
- Used sqlite_stat1 estimates for tables exceeding the cap; exact counts only for provably small tables.
- Introduced TableRowCount type with exact/estimate/atLeast variants reflected in rendered output.
- Computed screenshot destination extensions from the MIME type of bytes being written.
- Updated auto-generated paths in screenshot and temp directories to use the matching extension.
- Extracted TUI rendering from `eval.ts` into a dependency-light `eval-render.ts` to break the circular initialization chain.
- `renderers.ts` now imports `evalToolRenderer` from `eval-render` directly, avoiding re-entry into the root barrel while `eval.ts` is still initializing.
- `eval.ts` re-exports `evalToolRenderer` and `EVAL_DEFAULT_PREVIEW_LINES` for backward compatibility.
- Increased first-event timeout test budget from 50ms to 5000ms to prevent CI scheduler jitter from tripping the watchdog on success cases.
- Moved EvalBackendsAllowance and related functions to a dedicated eval-backends.ts module.
- Replaced ad-hoc brokenShellSessions tracking with a quarantineShellSession helper that also awaits the abort cleanup promise.
- Applied quarantine on timeout and cancellation paths, not just errors.
- Added `DiagnosticsLedger` to track diagnostics already surfaced per file, suppressing repeats within a session.
- Wired dedup into both edit and write tools via a `transformDiagnostics` hook on the writethrough pipeline.
- Added `lsp.diagnosticsDeduplicate` setting (default: true) to control the behavior.
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.
Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.
- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
- Extended ResolveContext / WriteContext with localProtocolOptions so the
internal-URL router can thread the calling session's local-root mapping
through to handlers.
- LocalProtocolHandler.resolveOptions now prefers context.localProtocolOptions
before consulting the process-global override or the first main-kind session
in AgentRegistry, fixing multi-session ACP hosts (cmux) where reads of
local://PLAN.md were routing to a sibling session's artifacts dir even
though plan-mode writes succeeded against the calling session.
- read, find, ast_grep, ast_edit, and search now thread
this.session.localProtocolOptions into the router so local://, memory://,
agent://, and other handlers see the right caller.
- Added regression tests covering the override-vs-context priority and the
ENOENT-against-caller-root path.
Fixes#1608