- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Added `HindsightSessionState` to `AgentSession` and bound hindsight lifecycle hooks to session state.
- Removed global hindsight state/queue handling and replaced it with per-session `HindsightRetainQueue` batching and scoped flushing.
- Reworked recall, reflect, and retain tools to use `session.getHindsightSessionState()` instead of sessionId-based lookup.
- Updated SDK/task/backend/controller flows to pass `session`/`parentHindsightSessionState`, scope `/memory` behavior, and document it in changelog.
- Added per-session retain queues with size/time auto-flush, recursive drain, and lifecycle flushes on end/clear/enqueue.
- Changed `hindsight-retain` to validate session state, enqueue writes, and return `Memory queued.` immediately.
- Added `notice` event support and handlers that route error/warning/status messages with source-aware formatting.
- Updated client and tests with shared request mapping, `RequestOptions`, `buildMemoryItem`, and expanded batch/list/doc APIs.
- Added inline hashline parse and apply support for `<` prepend and `+` append operations with prefix+suffix edits.
- Added fail-fast behavior to reject inline modify ops combined with delete or replace on same line.
- Renamed HASHLINE_* and mode symbols to HL_* in prompt tooling, read/search checks, and prompt templates.
- Standardized separators to `PI_HL_SEP`/`HL_EDIT_SEP` and fixed `HL_BODY_SEP='|'`, updating parser formatting behavior.
- Updated benchmark subtype constants and python cleanup test setup to use HL_* values and AgentRegistry mock failure injection.
- Added a new optional `beforeAgentStartPrompt` hook to `MemoryBackend` and implemented it in the Hindsight backend to recall long-term context for the first turn.
- Updated `AgentSession` startup flow to inject the recalled context into the turn-specific system prompt before the first response is generated.
- Preserved `<hindsight_memories>` tags in Hindsight developer instructions and added tests for first-turn injection and state caching.
- Added `memory.backend` and `hindsight.*` settings schema with migration from `memories.enabled` legacy mode.
- Added Hindsight memory backend runtime modules for resolved config, client creation, bank ID derivation, and state lifecycle.
- Added off/local/hindsight backends and resolver wiring across SDK, commands, and compaction context.
- Added `hindsight_recall`, `hindsight_reflect`, and `hindsight_retain` tools with schema validation and backend gating.
- Added Memory tab metadata and symbols to expose backend selection in the settings UI.
- Added package export barrels and tests for bank ID, content formatting, and hindsight config env precedence.
Manual handoff starts a fresh session and seeds it with a displayed custom handoff message, not an assistant message. Session persistence normally waits for an assistant message before creating the session file, which made the new handoff session exist only in memory until later activity.
Persist the seeded handoff session after injecting the handoff context, and record the previous session file as its parent so lineage remains discoverable.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
buildSystemPrompt injects today's date into the prompt body. The
applied-tool signature skips rebuilds when tools are byte-identical,
but did not cover the date — so a session spanning midnight with only
tool-stable MCP reconnects would keep yesterday's date indefinitely.
Append the current YYYY-MM-DD date as a suffix to the signature so any
reconnect after midnight triggers exactly one rebuild, then resumes
skipping normally for the rest of the new day.
Built-in tools whose prompt-rendered metadata depends on settings
(`TaskTool`, `SearchToolBm25Tool`, `EditTool`) expose `description`/
`label` via getters that re-evaluate on every access. The skip
optimization in `#applyActiveToolsByName` is correctness-safe for these
because `#computeAppliedToolSignature` reads `tool.description` live
each call, so a settings flip mutates the rendered string and differs
the signature on the next refresh.
This contract was implicit; a future refactor that caches per-tool
description strings would silently break it. Defending it explicitly:
- Added a regression test that wires a getter onto a CustomTool's
`description`, verifies `refreshMCPTools` skips while the underlying
state is unchanged, then mutates the state (without changing tool
object identity) and verifies the rebuild fires.
- Expanded the `#computeAppliedToolSignature` docstring to document
the getter-based coverage path and the SDK-init-time closure
constants in `sdk.ts` that genuinely cannot change at runtime
(`repeatToolDescriptions`, `eagerTasks`, `intentField`,
`mcpDiscoveryEnabled`, `secretsEnabled`).
Triggered by a review question on whether the skip breaks settings-
based prompt changes. It does not, but the property is non-obvious.
A tool's wire-visible name (`customWireName`) is rendered into the
system prompt body via `toolPromptNames`, but the applied-tool signature
only hashed name+label+description. A future tool whose wire name varied
without touching the other fields would silently produce a stale system
prompt that advertises the wrong callable name to the model — desyncing
prompt guidance from actual tool routing.
Today the only mutation path (edit-mode toggle) is also covered by a
description change and an explicit `refreshBaseSystemPrompt` from
`#syncEditToolModeAfterModelChange`, so this is a defensive fix rather
than a live bug. Including `customWireName` makes the signature a
self-consistent model of the prompt inputs.
- Extended `describeTool` in `#computeAppliedToolSignature` to include
`tool.customWireName ?? ""`. Applies to both the active-tool segment
and the (mcpDiscoveryEnabled) registry segment via the shared helper.
- Updated the docstring to call out wire-name coverage.
- Added a regression test that mutates `customWireName` between
identical-metadata refreshes and asserts the rebuild fires.
Per Codex review on #890.
Two cache-stability fixes for Anthropic prompt caching during MCP server
reconnects, which happen routinely (~5 min per server) in long sessions
due to SSE transport keepalive timeouts.
1) MCPManager: deterministic tool ordering
`#tools` is now sorted by name after every mutation. The previous
filter-out + push-to-end pattern in `#replaceServerTools` moved the
reconnecting server's tools to the end of the array, producing a new
byte order whenever the reconnect sequence differed from the initial
discovery sequence. With multiple healthy servers, each reconnect of
the non-last server flipped the order and invalidated the tools
cache breakpoint sent to Anthropic.
Sort applies in `discoverAndConnect` (initial population) and
`#replaceServerTools` (used by `reconnectServer` and
`refreshServerTools`). The comparator is character-code based,
locale-independent and deterministic. `sortMCPToolsByName` is
exported as a small generic helper and unit-tested.
2) AgentSession: skip system-prompt rebuild when inputs are unchanged
`#applyActiveToolsByName` (called from `refreshMCPTools` after every
reconnect) used to unconditionally call `rebuildSystemPrompt` and
`setSystemPrompt` even when the resulting prompt was byte-identical.
This wasted CPU on every flap and risked silent cache invalidation
if the rebuild path ever became non-deterministic.
Now `#applyActiveToolsByName` computes a stable signature of the
inputs `rebuildSystemPrompt` reads and skips the rebuild when the
signature matches the last successful one. The signature covers:
- active tool names in render order
- active tool labels and descriptions (rendered as `{{label}}:
\`{{name}}\`` in the prompt body)
- when MCP discovery is on, every registry tool's name + label +
description (the prompt summarizes discoverable-but-inactive
MCP tools)
- per-server MCP `instructions` text (embedded under "## MCP
Server Instructions" in the appended prompt; can change on
server upgrade while tool list stays identical)
Server instructions are read via a new optional
`getMcpServerInstructions` callback on `AgentSessionConfig`, wired
from the SDK as `() => mcpManager.getServerInstructions()`.
`refreshBaseSystemPrompt()` continues to rebuild unconditionally and
refreshes the cached signature, so explicit refreshes still pick up
ambient changes (edit-mode toggles, memory writes, etc.) that the
signature does not cover.
Signature inputs deliberately NOT covered: tool input schemas, memory
instructions read from disk, and other ambient state. Callers that
mutate those must call `refreshBaseSystemPrompt()` explicitly; existing
hooks (`#syncEditToolModeAfterModelChange`, memory hooks, `/clear`)
already do.
- Added `Process` class with pidfd (Linux), libproc (macOS), and handle (Windows) ownership for race-free signaling.
- Replaced `killTree`/`listDescendants` free functions with `Process.fromPid`, `fromPath`, `terminate`, and `waitForExit`.
- Added `TerminationTargets` for batching pgid+pid sets across pty and shell job teardown.
- Migrated `procmgr` and `ptree` to use the new native API, removing the `setNativeKillTree` injection pattern.
- Removed `id` fields from todo models/fixtures and switched session clones to content-based task identity.
- Replaced `/todo_write` `replace` with `init`, updated setup schemas to `list`/`phase`, and append content-only items.
- Updated `/todo` command flows to match phases and tasks by names/content (exact/prefix/substr, case-insensitive), with no ID targeting.
- Updated rendering/output labels to `# Todos`, `formatPhaseDisplayName`, and Roman-numeral phase headings across todo views.
- Aligned prompts, changelog, and todo tests/fixtures with the new init and content-based todo-write contract.
- Integrated `isUnexpectedSocketCloseMessage` into transient error detection so Bun socket-closure failures are treated as retryable.
- Added a retry fallback test that simulates a Bun socket close error and verifies the request is retried successfully with matching retry start/end events and recovered output.
- Fixed bash interceptor to check both raw and cwd-normalized commands, catching commands hidden behind leading `cd ... &&` wrappers.
- Fixed LSP client shutdown to await graceful shutdown with a 5s timeout before killing the process, and parallelized `shutdownAll` via `Promise.allSettled`.
- Fixed concurrent bash command tracking by replacing a single abort controller with a Set, preventing premature cancellation of parallel commands.
- Removed `./hooks` and `./hooks/*` export entries from the coding-agent package exports map.
- Updated pinned Rust nightly toolchain from `nightly-2026-03-27` to `nightly-2026-04-29` in `rust-toolchain.toml` and CI workflow.
- Replaced custom already-published detection in `ci-release-publish.ts` with `bun publish --tolerate-republish` flag.
Mirror anthropic.ts:disableThinkingIfToolChoiceForced for backends that 400
on combined reasoning + forced tool_choice. Kimi explicitly rejects this
combination ('tool_choice specified is incompatible with thinking enabled')
on its native API, OpenCode-Go, OpenRouter, etc. Anthropic itself enforces
the same constraint, so Claude reached through OpenAI-compat proxies
(LiteLLM, Vertex chat-completions, OpenRouter) needs the same handling.
Adds disableReasoningOnForcedToolChoice compat flag, defaulted on for any
Kimi (moonshotai/kimi*, kimi-* ids) or Anthropic (provider/baseUrl/claude*
ids) model. When tool_choice resolves to anything that forces a tool call
(required, named function), reasoning_effort and the OpenRouter-shaped
nested reasoning object are dropped for that turn. The forced tool_choice
itself stays so the agent still gets the tool call.
Replaces the previous (incorrect) approach of unconditionally dropping
tool_choice for kimi reasoning models, which broke explicit tool routing.
Fixes#827
Add optional defaultLevel to ThinkingConfig schema/type so models.yml can
declare a preferred starting thinking level per model. On model switch
the agent session adopts model.thinking.defaultLevel when present (with
explicit caller-supplied level still winning); otherwise current behavior
is preserved. SDK initial selection prefers the model's defaultLevel
before falling back to the global defaultThinkingLevel setting.
Fixes#775
The autoContinue post-compaction prompt echoed the summary's '## Next
Steps' heading, but that section is generated only from the compacted
tail; the kept ~20k recent tokens are not fed to the summarizer. When
the user pivoted within the kept window, 'Continue if you have next
steps.' anchored the model on the now-outdated plan instead of the
latest intent.
Move the prompt to prompts/system/auto-continue.md (per AGENTS.md
no-inline-prompts rule) and rewrite it to direct the model to re-read
the kept recent messages and follow the user's most recent request,
explicitly allowing it to stop when nothing remains.
Fixes#840
- Added a shared session path resolver that maps local:// URLs through local-protocol options, skips other internal schemes, and returns an absolute filesystem path for real files.
- Updated streaming-edit pre-cache and post-edit cache invalidation to use the shared resolver, preventing internal-scheme assertions while keeping filesystem-based flow for local plan files.
- Extended streaming-edit tests to confirm local:// plan edits complete without panicking and that auto-generated checks receive resolved absolute paths.
navigateTree() was unconditionally calling buildSessionContext() twice —
once to build stateContext for agent.replaceMessages, and again after
the session_tree emit to capture any hook-driven mutations. 6 of 7
callers discard result.sessionContext, so they paid an O(N) walk for
nothing.
Gate both the emit and the post-emit rebuild behind
extensionRunner.hasHandlers("session_tree"), mirroring the
session_before_tree guard at the top of the same function. When no
handlers are registered, stateContext is returned directly (the
intermediate ops don't mutate SessionManager).
rawContext was captured before the awaited session_tree hook emit,
so extension appendEntry/setLabel mutations during hook handling were
invisible to the UI until the next full rebuild.
Rename the pre-hook build to stateContext (used only for replaceMessages)
and add a second buildSessionContext() call post-hook whose result is
returned as sessionContext for the renderer.
- Added agent identity and registry fields to session configuration and session creation, enabling relay routing metadata for agent sessions.
- Implemented non-persistent IRC relay emission to forward incoming and reply observations from non-main agents into the main session UI.
- Updated IRC UI rendering to support `irc:relay` messages with participant-aware arrow formatting and body display.
Before this change, every navigateTree → renderInitialMessages call path
performed two independent O(N) session-tree walks:
1. agent-session.ts:6586 buildDisplaySessionContext() [inside navigateTree]
2. ui-helpers.ts:402 sessionManager.buildSessionContext() [inside renderInitialMessages]
Changes:
- agent-session.ts: navigateTree() now calls sessionManager.buildSessionContext()
once, derives the display (deobfuscated) context from the raw result, and
returns the raw SessionContext in the result object.
- ui-helpers.ts: renderInitialMessages() accepts an optional prebuiltContext
parameter; reuses it when provided, falls back to buildSessionContext() otherwise.
- interactive-mode.ts: forwards prebuiltContext through the wrapper.
- modes/types.ts: updates InteractiveModeContext interface to match.
- selector-controller.ts: passes result.sessionContext from navigateTree into
renderInitialMessages(), closing the deduplication loop.
Bench (100-msg session, 200 iterations):
two walks [BEFORE]: 0.0702ms/op
one walk [AFTER]: 0.0298ms/op
Saved: 0.0404ms/navigation (57.5% reduction per navigate)
Tests: render-initial-messages-dedupe.test.ts asserts buildSessionContext is
called 0 times when a prebuilt context is passed, 1 time as fallback.
When plan-mode persists the plan file at the synthetic local://PLAN.md
URL, an Edit tool call would crash the entire session. The streaming-edit
pre-cache called resolveToCwd on the path unconditionally; that helper
asserts internal-scheme URLs cannot be resolved as filesystem paths and
threw synchronously inside the assistant-message-event interceptor. The
throw escaped as an Unhandled Rejection, killing the session.
Add an isInternalUrlPath() early-return guard at the top of:
- #getStreamingEditToolCall — returns undefined so the caller skips
pre-cache and the auto-generated guard for internal URLs entirely.
- #invalidateFileCacheForPath — early-returns; nothing to invalidate
for paths that were never cached.
Internal-scheme URLs don't have a stable filesystem path; the actual
Edit tool dispatches through its protocol handler (the same path Write
uses via resolvePlanPath), so the edit still applies — only the on-disk
pre-cache (Morph fast-apply optimization) is skipped.
Add a regression test in streaming-edit-abort.test.ts that drives a
streaming Edit toolcall with path: 'local://PLAN.md'. Without this fix
the test reproduces the original panic stack:
assertNotInternalUrl → resolveToCwd → #getStreamingEditToolCall
→ #preCacheStreamingEditFile → assistant message interceptor.
- Added an `irc_message` session event carrying custom IRC messages and emitted it when IRC records are created.
- Registered an IRC message handler in EventController that skips duplicate messages by role, custom type, and timestamp.
- The handler now appends IRC messages to chat, resets read grouping, and triggers a UI render.
- Updated `formatMatchLine` to emit `*` for matched lines, a leading space for context, and a `|` anchor/content separator.
- Revised grep/hashline mismatch messages and prompts to describe the new marker and separator format.
- Aligned affected atom and hashline tests with the updated match-line prefixes and separators.
- Added `AgentRegistry` singleton with session registration/unregistration and IRC routing metadata for peer lookups.
- Added IRC messaging prompts and tooling with `irc.enabled` setting, `list/send` tool paths, and peer roster rendering.
- Changed `/btw` to session-side `runEphemeralTurn`, added background IRC exchange flushing, and fixed empty-input checks.
- Added unit tests for IRC tool and BtwController ephemeral behavior, including disabled, busy, not-found, and abort cases.
/drop works like /new but permanently deletes the current session file
and artifacts instead of flushing (saving) it. Useful when the session
should not be kept.
- add drop?: boolean to NewSessionOptions
- branch in AgentSession.newSession(): skip flush, delete via
FileSessionStorage.deleteSessionWithArtifacts when drop=true;
deletion failure is non-fatal (logged, new session still starts)
- wire handleDropCommand() through InteractiveModeContext interface,
InteractiveMode delegation, and CommandController implementation
- guard: shows error if session has not been saved yet (no file to drop)
- register /drop in builtin-registry adjacent to /new
- Added `note` support to todo-write with `op: "note"` and required `text` input.
- Added optional `notes: string[]` to todo models and preserved notes in cloning and session task mapping.
- Implemented `op: "note"` append flow and markdown `>` block serialization/parsing for todo import/export.
- Updated HUD todo rendering to append superscript `+N` note markers and show in-progress note bodies.
- Documented note operation, required text field, and note rendering rules in todo-write docs and changelog.
- Added tracking of the last successful non-error yield tool call when a yield execution ends.
- Cleared the tracked yield ID and skipped post-turn maintenance when appropriate, including when the last assistant message was a successful yield.
- Standardized unchanged-result error messages in atom and replace edit modes.
- Added command-marker metadata and lifecycle in process execution, including completion markers and exit-code writes.
- Refactored minimization to support `MarkedCommands` mode, token-based detection, and marker-aware stripping.
- Added `onMinimizedSave` and `saveBashOriginalArtifact` to persist full bash-original output artifacts.
- Expanded public exports and marker hooks so external command launch metadata can be controlled by consumers.
Addresses two P1 chatgpt-codex-connector findings on PR #623:
1. Block follow-up auto-continue during retry backoff
isStreaming only checks agent.state.isStreaming || #promptInFlightCount
and does NOT include #retryPromise. When a notification-driven follow-up
arrives during the deliberate sleep window in #handleRetryableError,
the idle gate fires and schedules agent.continue() immediately,
bypassing the configured retry delay and racing the retry timer.
Fix: include isRetrying in the gate.
2. Guard auto-continue against non-assistant resume state
agent.continue() only dequeues follow-ups from an assistant-ended
state; when the last message is user/toolResult (e.g. after an early
abort), continue() resumes prior context and runs an extra model
call on the stale prompt before draining the queue.
Fix: require messages.at(-1).role === 'assistant' before scheduling.
Extracted #canAutoContinueForFollowUp() so both the pre-schedule and
shouldContinue re-check use identical conditions.
Previously session.followUp() enqueued the message into agent.followUp()
but never scheduled a continue when the agent was idle. The message sat
in the queue until the next user turn triggered a prompt.
Mirrors the pattern already used by TTSR deferred injection: pair
agent.followUp() with #scheduleAgentContinue(). The shouldContinue guard
prevents spurious continues if multiple follow-ups arrive in a burst and
the queue drains before a scheduled task runs.
Only affects the idle path - during streaming the running loop's
getFollowUpMessages callback drains the queue at end of turn naturally.
Extract duplicate normalizeLocalScheme regex pattern into a shared function in path-utils.ts. Updated interactive-mode.ts, approved-plan.ts, agent-session.ts, bash-skill-urls.ts, and plan-mode-guard.ts to use the shared utility. Also fixed error message formatting (removed extra backslashes).
On Linux, Node's path.normalize() collapses the double slash in
local://PLAN.md to local:/PLAN.md, creating a directory called local:
in the project root instead of routing through the local:// protocol handler.
Defense-in-depth fixes across 5 layers:
1. resolveToCwd() now throws if a path starts with any internal URL
scheme prefix (local:, agent:, skill:, etc.), preventing all 59
call sites from treating URIs as relative filesystem paths.
2. resolvePlanPath() now matches on local: prefix (not just local://)
and normalizes local:/ to local:// before resolution, catching
all slash variants.
3. Bash URL expansion regex and early-exit checks now also match
local:/ (single slash), and normalize before resolution.
4. Edit preview/diff functions now gracefully skip internal URL paths
instead of crashing via the resolveToCwd guard.
5. All startsWith('local://') checks updated to startsWith('local:')
with normalization in agent-session, interactive-mode, and
approved-plan modules.
Also adds local: to .gitignore to prevent accidental commits of the
leaked directory.
The #emitSessionEvent refactoring changed the order: #emit was called before
#emitExtensionEvent, but for non-message_update events this introduced
microtask timing issues that broke auto-compaction and handoff tests.
Fix: keep the original order (extension first, then emit) for non-message_update
events. Only message_update events use fire-and-forget extension queueing with
emit-first ordering.