- Added `HindsightSessionState` to `AgentSession` and bound hindsight lifecycle hooks to session state.
- Removed global hindsight state/queue handling and replaced it with per-session `HindsightRetainQueue` batching and scoped flushing.
- Reworked recall, reflect, and retain tools to use `session.getHindsightSessionState()` instead of sessionId-based lookup.
- Updated SDK/task/backend/controller flows to pass `session`/`parentHindsightSessionState`, scope `/memory` behavior, and document it in changelog.
- Added per-session retain queues with size/time auto-flush, recursive drain, and lifecycle flushes on end/clear/enqueue.
- Changed `hindsight-retain` to validate session state, enqueue writes, and return `Memory queued.` immediately.
- Added `notice` event support and handlers that route error/warning/status messages with source-aware formatting.
- Updated client and tests with shared request mapping, `RequestOptions`, `buildMemoryItem`, and expanded batch/list/doc APIs.
- Added inline hashline parse and apply support for `<` prepend and `+` append operations with prefix+suffix edits.
- Added fail-fast behavior to reject inline modify ops combined with delete or replace on same line.
- Renamed HASHLINE_* and mode symbols to HL_* in prompt tooling, read/search checks, and prompt templates.
- Standardized separators to `PI_HL_SEP`/`HL_EDIT_SEP` and fixed `HL_BODY_SEP='|'`, updating parser formatting behavior.
- Updated benchmark subtype constants and python cleanup test setup to use HL_* values and AgentRegistry mock failure injection.
- Added a new optional `beforeAgentStartPrompt` hook to `MemoryBackend` and implemented it in the Hindsight backend to recall long-term context for the first turn.
- Updated `AgentSession` startup flow to inject the recalled context into the turn-specific system prompt before the first response is generated.
- Preserved `<hindsight_memories>` tags in Hindsight developer instructions and added tests for first-turn injection and state caching.
- Added `memory.backend` and `hindsight.*` settings schema with migration from `memories.enabled` legacy mode.
- Added Hindsight memory backend runtime modules for resolved config, client creation, bank ID derivation, and state lifecycle.
- Added off/local/hindsight backends and resolver wiring across SDK, commands, and compaction context.
- Added `hindsight_recall`, `hindsight_reflect`, and `hindsight_retain` tools with schema validation and backend gating.
- Added Memory tab metadata and symbols to expose backend selection in the settings UI.
- Added package export barrels and tests for bank ID, content formatting, and hindsight config env precedence.
- Removed title-source aware branching from session terminal-title and accent helpers, and updated callers to use session name plus cwd only.
- Dropped UUID-based recent-session naming by preferring explicit header titles or first user prompts and generating an "Untitled · <time>" fallback.
- Adjusted welcome session-row rendering for width-aware name truncation and disabled reasoning in title generation requests to keep terminal titles concise.
Manual handoff starts a fresh session and seeds it with a displayed custom handoff message, not an assistant message. Session persistence normally waits for an assistant message before creating the session file, which made the new handoff session exist only in memory until later activity.
Persist the seeded handoff session after injecting the handoff context, and record the previous session file as its parent so lineage remains discoverable.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
buildSystemPrompt injects today's date into the prompt body. The
applied-tool signature skips rebuilds when tools are byte-identical,
but did not cover the date — so a session spanning midnight with only
tool-stable MCP reconnects would keep yesterday's date indefinitely.
Append the current YYYY-MM-DD date as a suffix to the signature so any
reconnect after midnight triggers exactly one rebuild, then resumes
skipping normally for the rest of the new day.
Built-in tools whose prompt-rendered metadata depends on settings
(`TaskTool`, `SearchToolBm25Tool`, `EditTool`) expose `description`/
`label` via getters that re-evaluate on every access. The skip
optimization in `#applyActiveToolsByName` is correctness-safe for these
because `#computeAppliedToolSignature` reads `tool.description` live
each call, so a settings flip mutates the rendered string and differs
the signature on the next refresh.
This contract was implicit; a future refactor that caches per-tool
description strings would silently break it. Defending it explicitly:
- Added a regression test that wires a getter onto a CustomTool's
`description`, verifies `refreshMCPTools` skips while the underlying
state is unchanged, then mutates the state (without changing tool
object identity) and verifies the rebuild fires.
- Expanded the `#computeAppliedToolSignature` docstring to document
the getter-based coverage path and the SDK-init-time closure
constants in `sdk.ts` that genuinely cannot change at runtime
(`repeatToolDescriptions`, `eagerTasks`, `intentField`,
`mcpDiscoveryEnabled`, `secretsEnabled`).
Triggered by a review question on whether the skip breaks settings-
based prompt changes. It does not, but the property is non-obvious.
A tool's wire-visible name (`customWireName`) is rendered into the
system prompt body via `toolPromptNames`, but the applied-tool signature
only hashed name+label+description. A future tool whose wire name varied
without touching the other fields would silently produce a stale system
prompt that advertises the wrong callable name to the model — desyncing
prompt guidance from actual tool routing.
Today the only mutation path (edit-mode toggle) is also covered by a
description change and an explicit `refreshBaseSystemPrompt` from
`#syncEditToolModeAfterModelChange`, so this is a defensive fix rather
than a live bug. Including `customWireName` makes the signature a
self-consistent model of the prompt inputs.
- Extended `describeTool` in `#computeAppliedToolSignature` to include
`tool.customWireName ?? ""`. Applies to both the active-tool segment
and the (mcpDiscoveryEnabled) registry segment via the shared helper.
- Updated the docstring to call out wire-name coverage.
- Added a regression test that mutates `customWireName` between
identical-metadata refreshes and asserts the rebuild fires.
Per Codex review on #890.
Two cache-stability fixes for Anthropic prompt caching during MCP server
reconnects, which happen routinely (~5 min per server) in long sessions
due to SSE transport keepalive timeouts.
1) MCPManager: deterministic tool ordering
`#tools` is now sorted by name after every mutation. The previous
filter-out + push-to-end pattern in `#replaceServerTools` moved the
reconnecting server's tools to the end of the array, producing a new
byte order whenever the reconnect sequence differed from the initial
discovery sequence. With multiple healthy servers, each reconnect of
the non-last server flipped the order and invalidated the tools
cache breakpoint sent to Anthropic.
Sort applies in `discoverAndConnect` (initial population) and
`#replaceServerTools` (used by `reconnectServer` and
`refreshServerTools`). The comparator is character-code based,
locale-independent and deterministic. `sortMCPToolsByName` is
exported as a small generic helper and unit-tested.
2) AgentSession: skip system-prompt rebuild when inputs are unchanged
`#applyActiveToolsByName` (called from `refreshMCPTools` after every
reconnect) used to unconditionally call `rebuildSystemPrompt` and
`setSystemPrompt` even when the resulting prompt was byte-identical.
This wasted CPU on every flap and risked silent cache invalidation
if the rebuild path ever became non-deterministic.
Now `#applyActiveToolsByName` computes a stable signature of the
inputs `rebuildSystemPrompt` reads and skips the rebuild when the
signature matches the last successful one. The signature covers:
- active tool names in render order
- active tool labels and descriptions (rendered as `{{label}}:
\`{{name}}\`` in the prompt body)
- when MCP discovery is on, every registry tool's name + label +
description (the prompt summarizes discoverable-but-inactive
MCP tools)
- per-server MCP `instructions` text (embedded under "## MCP
Server Instructions" in the appended prompt; can change on
server upgrade while tool list stays identical)
Server instructions are read via a new optional
`getMcpServerInstructions` callback on `AgentSessionConfig`, wired
from the SDK as `() => mcpManager.getServerInstructions()`.
`refreshBaseSystemPrompt()` continues to rebuild unconditionally and
refreshes the cached signature, so explicit refreshes still pick up
ambient changes (edit-mode toggles, memory writes, etc.) that the
signature does not cover.
Signature inputs deliberately NOT covered: tool input schemas, memory
instructions read from disk, and other ambient state. Callers that
mutate those must call `refreshBaseSystemPrompt()` explicitly; existing
hooks (`#syncEditToolModeAfterModelChange`, memory hooks, `/clear`)
already do.
- Added `Process` class with pidfd (Linux), libproc (macOS), and handle (Windows) ownership for race-free signaling.
- Replaced `killTree`/`listDescendants` free functions with `Process.fromPid`, `fromPath`, `terminate`, and `waitForExit`.
- Added `TerminationTargets` for batching pgid+pid sets across pty and shell job teardown.
- Migrated `procmgr` and `ptree` to use the new native API, removing the `setNativeKillTree` injection pattern.
- Removed `id` fields from todo models/fixtures and switched session clones to content-based task identity.
- Replaced `/todo_write` `replace` with `init`, updated setup schemas to `list`/`phase`, and append content-only items.
- Updated `/todo` command flows to match phases and tasks by names/content (exact/prefix/substr, case-insensitive), with no ID targeting.
- Updated rendering/output labels to `# Todos`, `formatPhaseDisplayName`, and Roman-numeral phase headings across todo views.
- Aligned prompts, changelog, and todo tests/fixtures with the new init and content-based todo-write contract.
- Integrated `isUnexpectedSocketCloseMessage` into transient error detection so Bun socket-closure failures are treated as retryable.
- Added a retry fallback test that simulates a Bun socket close error and verifies the request is retried successfully with matching retry start/end events and recovered output.
- Fixed bash interceptor to check both raw and cwd-normalized commands, catching commands hidden behind leading `cd ... &&` wrappers.
- Fixed LSP client shutdown to await graceful shutdown with a 5s timeout before killing the process, and parallelized `shutdownAll` via `Promise.allSettled`.
- Fixed concurrent bash command tracking by replacing a single abort controller with a Set, preventing premature cancellation of parallel commands.
- Removed `./hooks` and `./hooks/*` export entries from the coding-agent package exports map.
- Updated pinned Rust nightly toolchain from `nightly-2026-03-27` to `nightly-2026-04-29` in `rust-toolchain.toml` and CI workflow.
- Replaced custom already-published detection in `ci-release-publish.ts` with `bun publish --tolerate-republish` flag.
Mirror anthropic.ts:disableThinkingIfToolChoiceForced for backends that 400
on combined reasoning + forced tool_choice. Kimi explicitly rejects this
combination ('tool_choice specified is incompatible with thinking enabled')
on its native API, OpenCode-Go, OpenRouter, etc. Anthropic itself enforces
the same constraint, so Claude reached through OpenAI-compat proxies
(LiteLLM, Vertex chat-completions, OpenRouter) needs the same handling.
Adds disableReasoningOnForcedToolChoice compat flag, defaulted on for any
Kimi (moonshotai/kimi*, kimi-* ids) or Anthropic (provider/baseUrl/claude*
ids) model. When tool_choice resolves to anything that forces a tool call
(required, named function), reasoning_effort and the OpenRouter-shaped
nested reasoning object are dropped for that turn. The forced tool_choice
itself stays so the agent still gets the tool call.
Replaces the previous (incorrect) approach of unconditionally dropping
tool_choice for kimi reasoning models, which broke explicit tool routing.
Fixes#827
buildSessionContext walked the entry path and unconditionally overwrote
models.default from every assistant message's reported model. Temporary
fallbacks (retry fallback, context promotion) and codex-side model
downgrades both produce assistant messages tagged with a different model
id, which clobbered the user's explicit /model pick on resume and made
the session silently revert to the older model.
Treat assistant-message inference as a legacy fallback that only fills
in models.default when no explicit `model_change` with role="default"
has been seen on the path.
Fixes#849
Add optional defaultLevel to ThinkingConfig schema/type so models.yml can
declare a preferred starting thinking level per model. On model switch
the agent session adopts model.thinking.defaultLevel when present (with
explicit caller-supplied level still winning); otherwise current behavior
is preserved. SDK initial selection prefers the model's defaultLevel
before falling back to the global defaultThinkingLevel setting.
Fixes#775
The autoContinue post-compaction prompt echoed the summary's '## Next
Steps' heading, but that section is generated only from the compacted
tail; the kept ~20k recent tokens are not fed to the summarizer. When
the user pivoted within the kept window, 'Continue if you have next
steps.' anchored the model on the now-outdated plan instead of the
latest intent.
Move the prompt to prompts/system/auto-continue.md (per AGENTS.md
no-inline-prompts rule) and rewrite it to direct the model to re-read
the kept recent messages and follow the user's most recent request,
explicitly allowing it to stop when nothing remains.
Fixes#840
- Added a shared session path resolver that maps local:// URLs through local-protocol options, skips other internal schemes, and returns an absolute filesystem path for real files.
- Updated streaming-edit pre-cache and post-edit cache invalidation to use the shared resolver, preventing internal-scheme assertions while keeping filesystem-based flow for local plan files.
- Extended streaming-edit tests to confirm local:// plan edits complete without panicking and that auto-generated checks receive resolved absolute paths.
- Added a `/context` slash command flow from registry to interactive-mode command dispatch.
- Added `handleContextCommand()` to the mode context interface and command-controller wiring.
- Added context usage breakdown utilities, cell allocation, and 20x10 usage rendering for token categories.
- Reworked compaction token estimation to use tokenizer counts, role aggregation, image token estimates, and fallback handling.
- Exported `resolveThresholdTokens()` as a public compaction helper.
- Updated `formatSessionDumpText` to skip `thinking` entries with empty or whitespace-only content.
- Prevented empty `<thinking>` sections from being emitted in session dump output.
navigateTree() was unconditionally calling buildSessionContext() twice —
once to build stateContext for agent.replaceMessages, and again after
the session_tree emit to capture any hook-driven mutations. 6 of 7
callers discard result.sessionContext, so they paid an O(N) walk for
nothing.
Gate both the emit and the post-emit rebuild behind
extensionRunner.hasHandlers("session_tree"), mirroring the
session_before_tree guard at the top of the same function. When no
handlers are registered, stateContext is returned directly (the
intermediate ops don't mutate SessionManager).
rawContext was captured before the awaited session_tree hook emit,
so extension appendEntry/setLabel mutations during hook handling were
invisible to the UI until the next full rebuild.
Rename the pre-hook build to stateContext (used only for replaceMessages)
and add a second buildSessionContext() call post-hook whose result is
returned as sessionContext for the renderer.
- Added agent identity and registry fields to session configuration and session creation, enabling relay routing metadata for agent sessions.
- Implemented non-persistent IRC relay emission to forward incoming and reply observations from non-main agents into the main session UI.
- Updated IRC UI rendering to support `irc:relay` messages with participant-aware arrow formatting and body display.
Before this change, every navigateTree → renderInitialMessages call path
performed two independent O(N) session-tree walks:
1. agent-session.ts:6586 buildDisplaySessionContext() [inside navigateTree]
2. ui-helpers.ts:402 sessionManager.buildSessionContext() [inside renderInitialMessages]
Changes:
- agent-session.ts: navigateTree() now calls sessionManager.buildSessionContext()
once, derives the display (deobfuscated) context from the raw result, and
returns the raw SessionContext in the result object.
- ui-helpers.ts: renderInitialMessages() accepts an optional prebuiltContext
parameter; reuses it when provided, falls back to buildSessionContext() otherwise.
- interactive-mode.ts: forwards prebuiltContext through the wrapper.
- modes/types.ts: updates InteractiveModeContext interface to match.
- selector-controller.ts: passes result.sessionContext from navigateTree into
renderInitialMessages(), closing the deduplication loop.
Bench (100-msg session, 200 iterations):
two walks [BEFORE]: 0.0702ms/op
one walk [AFTER]: 0.0298ms/op
Saved: 0.0404ms/navigation (57.5% reduction per navigate)
Tests: render-initial-messages-dedupe.test.ts asserts buildSessionContext is
called 0 times when a prebuilt context is passed, 1 time as fallback.
When plan-mode persists the plan file at the synthetic local://PLAN.md
URL, an Edit tool call would crash the entire session. The streaming-edit
pre-cache called resolveToCwd on the path unconditionally; that helper
asserts internal-scheme URLs cannot be resolved as filesystem paths and
threw synchronously inside the assistant-message-event interceptor. The
throw escaped as an Unhandled Rejection, killing the session.
Add an isInternalUrlPath() early-return guard at the top of:
- #getStreamingEditToolCall — returns undefined so the caller skips
pre-cache and the auto-generated guard for internal URLs entirely.
- #invalidateFileCacheForPath — early-returns; nothing to invalidate
for paths that were never cached.
Internal-scheme URLs don't have a stable filesystem path; the actual
Edit tool dispatches through its protocol handler (the same path Write
uses via resolvePlanPath), so the edit still applies — only the on-disk
pre-cache (Morph fast-apply optimization) is skipped.
Add a regression test in streaming-edit-abort.test.ts that drives a
streaming Edit toolcall with path: 'local://PLAN.md'. Without this fix
the test reproduces the original panic stack:
assertNotInternalUrl → resolveToCwd → #getStreamingEditToolCall
→ #preCacheStreamingEditFile → assistant message interceptor.
- Added an `irc_message` session event carrying custom IRC messages and emitted it when IRC records are created.
- Registered an IRC message handler in EventController that skips duplicate messages by role, custom type, and timestamp.
- The handler now appends IRC messages to chat, resets read grouping, and triggers a UI render.
- Updated `read` selector parsing for file and URL reads to accept optional leading `L` and `+` count-style ranges.
- Adjusted truncation notices and schema/help text to emit and suggest `sel` offsets without the `L` prefix, including continuation and suggestion messages.
- Updated hashline/output parsing and related tests to recognize the new `sel` formatting in truncation notices.
- Updated `formatMatchLine` to emit `*` for matched lines, a leading space for context, and a `|` anchor/content separator.
- Revised grep/hashline mismatch messages and prompts to describe the new marker and separator format.
- Aligned affected atom and hashline tests with the updated match-line prefixes and separators.
- dropSession: close persist writer before deletion to prevent EPERM
on Windows where an open file handle blocks unlink
- #runNewSessionFlow: guard UI reset on newSession return value;
session_before_switch hook cancellation now prevents chat state
being cleared and success banner being shown
- Added `AgentRegistry` singleton with session registration/unregistration and IRC routing metadata for peer lookups.
- Added IRC messaging prompts and tooling with `irc.enabled` setting, `list/send` tool paths, and peer roster rendering.
- Changed `/btw` to session-side `runEphemeralTurn`, added background IRC exchange flushing, and fixed empty-input checks.
- Added unit tests for IRC tool and BtwController ephemeral behavior, including disabled, busy, not-found, and abort cases.
/drop works like /new but permanently deletes the current session file
and artifacts instead of flushing (saving) it. Useful when the session
should not be kept.
- add drop?: boolean to NewSessionOptions
- branch in AgentSession.newSession(): skip flush, delete via
FileSessionStorage.deleteSessionWithArtifacts when drop=true;
deletion failure is non-fatal (logged, new session still starts)
- wire handleDropCommand() through InteractiveModeContext interface,
InteractiveMode delegation, and CommandController implementation
- guard: shows error if session has not been saved yet (no file to drop)
- register /drop in builtin-registry adjacent to /new
- Added `note` support to todo-write with `op: "note"` and required `text` input.
- Added optional `notes: string[]` to todo models and preserved notes in cloning and session task mapping.
- Implemented `op: "note"` append flow and markdown `>` block serialization/parsing for todo import/export.
- Updated HUD todo rendering to append superscript `+N` note markers and show in-progress note bodies.
- Documented note operation, required text field, and note rendering rules in todo-write docs and changelog.
- Session metadata now stores on-disk size bytes, populated from file stats when collecting sessions.
- The session selector metadata line now renders file size using formatBytes instead of message counts.
- Session test fixtures were updated with the new size property so they match the revised SessionInfo shape.
- Added tracking of the last successful non-error yield tool call when a yield execution ends.
- Cleared the tracked yield ID and skipped post-turn maintenance when appropriate, including when the last assistant message was a successful yield.
- Standardized unchanged-result error messages in atom and replace edit modes.