- Added GoalRuntime with wall-clock and token accounting, budget steering, and lifecycle operations (create, pause, resume, drop, complete).
- Exposed goal tool as a hidden agent tool, activated only when goal mode is enabled.
- Integrated goal continuation loop in InteractiveMode with auto-submit between turns.
- Added status line segment and theme icons for goal mode state.
- Removed ExitPlanModeTool and deleted exit-plan-mode docs/tests, dropping the old approval contract outputs.
- Replaced plan-mode approval flow from exit_plan_mode to resolve across session, SDK, controllers, and discovery.
- Added standing resolve handler accessors and updated resolve routing for queued or standing approval handlers.
- Added PlanApprovalDetails and enforced normalized, validated approval titles with readable plan-file requirements.
- Extended resolve schema and invocation signatures with optional extra metadata and reason trimming behavior updates.
- Updated plan and resolve prompts and changelog guidance to require resolve action, reason, and extra.title for apply/discard.
- Pinned final plan path before handleCompactCommand so queued messages use approved plan, not draft.
- Added regression coverage for setPlanReferencePath timing before compaction queue flush.
A fifth ExitPlanMode approval choice — sits between the existing
"Approve and execute" (purge session) and "Approve and keep context"
(full transcript). Runs `handleCompactCommand` against the plan-mode
transcript with a planning-specific custom instruction rendered from
`plan-mode-compact-instructions.md`, then dispatches the plan-approved
synthetic prompt so it lands as the first entry in the freshly-
summarized transcript — giving execution a fresh cache anchor with
the rationale carried over.
Cancel/fail contract:
- ok → bookkeeping runs, plan-approved synthetic prompt dispatched.
- cancelled → bookkeeping runs (tools restored, plan reference path
recorded), warning surfaced, dispatch skipped.
`markPlanReferenceSent` is intentionally deferred past
the cancel guard so `AgentSession.#buildPlanReferenceMessage`
re-injects the plan on the operator's next prompt() call.
If we marked it sent on cancel, the executor's first turn
would have no plan context.
- failed → bookkeeping runs, error already surfaced by executeCompaction,
dispatch proceeds best-effort. Approval intent stands.
Cancel vs. fail is discriminated via `instanceof CompactionCancelledError`
at the session/compaction error boundary (introduced in the previous
commit), so any abort source — operator Esc, extension hook, programmatic
abort — classifies uniformly without input-modality or message-string
coupling.
Op: extend
Introduce `CompactionCancelledError` and `CompactionOutcome` ("ok" |
"cancelled" | "failed") so callers can discriminate user-driven aborts
from generic failures via `instanceof`, instead of inspecting error
messages or `AbortError`-name strings.
`AgentSession.compact()`'s two abort-rejection sites now throw the
typed sentinel; the model-call wrapper normalizes AbortError-shaped
rejections to the sentinel only when the compaction's abort signal
is actually set, preserving every other exception unchanged so real
compaction bugs are not silently relabeled as cancellations.
`CommandController.executeCompaction` and `handleCompactCommand`
return `Promise<CompactionOutcome>`; the catch classifies via
`instanceof CompactionCancelledError`. Existing callers (`/compact`,
loop runner, auto-compact) ignore the return value — non-breaking.
Op: extend
- Updated addMessageToChat in UI helpers and InteractiveModeContext to return rendered components instead of void.
- EventController now tracks IRC message components and removes them after a 10-second TTL, avoiding duplicate expiry scheduling per message signature.
- EventController dispose now clears all pending IRC expiry timers and tests were added for immediate render, TTL removal, duplicate suppression, and timer cleanup.
- Trim logo from 14w to 12w (2-wide legs).
- Diagonal BL→TR gradient instead of per-line LTR.
- Truecolor: 3-stop magenta→violet→cyan path that skips the deep-blue valley.
- 256-color: same 6-stop ramp as fallback.
- Compute the colored logo once at module load instead of per render.
- Updated ExtensionUIContext, InteractiveModeContext, and InteractiveMode to require editor factories to return CustomEditor instances.
- Removed the runtime compatibility guard and warning for non-CustomEditor implementations in setEditorComponent.
- Removed the test that verified rejection of non-CustomEditor factories in interactive-mode editor-component tests.
- Render each exit_plan_mode submission as a fresh plan review entry so refined plans are emitted into terminal scrollback.
- Keep external-editor edits replacing the current preview instead of adding duplicate review entries.
- Updated the plan review regression test to preserve the first preview and append the second one after intervening chat content.
- Added optional `/loop` `count|duration` command arguments and wired `command.args` into loop handling.
- Implemented `loop-limit` parsing and runtime types/helpers for iteration and duration budget limits with validation.
- Updated interactive mode to enforce loop limits per iteration, check duration expiry, and clear budget state on disable.
- Added loop-limit parse/runtime tests and fixed `/loop` arg errors plus macOS `MallocStackLogging` environment leakage.
- Move planModeEnabled check before plan.enabled guard so /plan can
still exit an active session even if setting toggled off mid-session
- Clear stale plan/plan_paused mode_change entries when plan.enabled
is off during session restore, preventing unexpected re-entry when
the setting is later re-enabled
- Added session-draft persistence methods in SessionManager to write unsent editor text to an artifacts-sidecar draft file and delete it after single-shot consumption.
- Persisted editor text during interactive shutdown and restored that draft on resume when the editor was empty, enabling Ctrl+D draft recovery.
- Updated Ctrl+D handling in the editor/controller path and added tests covering draft round-trip, artifact cleanup, stale-draft eviction, and in-memory no-op behavior.
- Added `recordLocalSubmission` and `withLocalSubmission` to replace inline signature set mutations across input and compaction flush paths.
- Signatures are now cleaned up automatically on delivery failure, preventing stale entries from suppressing editor-draft protection on retries.
- Extended test fixtures and added cases covering signature lifecycle for idle, streaming, and fire-and-forget submission paths.
- Status-line rendering now treats `statusLine.sessionAccent` as disabled only when explicitly set to false.
- Updated status-line overflow tests to assert gap colors use session accent when enabled and theme border when disabled.
- Test teardown now restores `WSL_INTEROP` and `WSL_DISTRO_NAME` environment variables after mutation.
- Added an isSettingsInitialized helper to expose whether global settings were created.
- Updated interactive mode to use the session name when settings are not yet initialized.
- Guarded the status-line accent lookup behind settings initialization to avoid premature access.
- Added WSL detection in terminal initialization and skipped periodic OSC 11 polling when running under Windows Subsystem for Linux.
- Removed mode 2031 activity tracking and stopped OSC 11 polling immediately when DA1/Mode 2031 reports were received.
- Extended OSC 11 appearance tests to restore platform/env state and verify no polling is started under WSL.
Fixes#914
- Added a new boolean statusLine.sessionAccent setting to settings schema and propagated it through status-line preview, controller, and component setting updates.
- Updated status line rendering and interactive border coloring to disable session-based accent colors when the setting is false.
- Added a regression test verifying the status-line gap uses theme border color instead of session accent when session accents are disabled.
Fixes#918
- Removed title-source aware branching from session terminal-title and accent helpers, and updated callers to use session name plus cwd only.
- Dropped UUID-based recent-session naming by preferring explicit header titles or first user prompts and generating an "Untitled · <time>" fallback.
- Adjusted welcome session-row rendering for width-aware name truncation and disabled reasoning in title generation requests to keep terminal titles concise.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
- Parallelized startup by deferring plugin preload and running AGENTS.md scan plus context/template/command discovery in parallel.
- Added AgentsMdSearch exports and options so prebuilt search results were passed into system-prompt construction.
- Reworked logger timing to use AsyncLocalStorage-backed nested spans, initialize a root span, and emit hierarchical summaries.
- Added PI_TIMING-gated TS/TSX module-load timing via side-effect module-timer registration and wrapped key init/request paths with logger.time.
- Removed `id` fields from todo models/fixtures and switched session clones to content-based task identity.
- Replaced `/todo_write` `replace` with `init`, updated setup schemas to `list`/`phase`, and append content-only items.
- Updated `/todo` command flows to match phases and tasks by names/content (exact/prefix/substr, case-insensitive), with no ID targeting.
- Updated rendering/output labels to `# Todos`, `formatPhaseDisplayName`, and Roman-numeral phase headings across todo views.
- Aligned prompts, changelog, and todo tests/fixtures with the new init and content-based todo-write contract.
When entering plan mode while the session is streaming, #applyPlanModeModel
defers the switch into #pendingModelSwitch and snapshots the previous model.
On exit, the snapshot was restored but the deferred switch was left queued,
so the next agent_end flush landed the session on the plan-role model after
the user had already left plan mode.
Drop the pending switch in #exitPlanMode when its target matches the
plan-role resolution; leave any other queued switch alone.
Fixes#816
- Added a shared session path resolver that maps local:// URLs through local-protocol options, skips other internal schemes, and returns an absolute filesystem path for real files.
- Updated streaming-edit pre-cache and post-edit cache invalidation to use the shared resolver, preventing internal-scheme assertions while keeping filesystem-based flow for local plan files.
- Extended streaming-edit tests to confirm local:// plan edits complete without panicking and that auto-generated checks receive resolved absolute paths.
- Added support for batch PR operations by accepting `pr` as string or array and dropping `worktree` input.
- Updated `pr_view` and `pr_diff` to normalize PR IDs, process multiple PRs in parallel, and emit combined summaries.
- Refactored checkout into `checkoutPullRequest`, added repo-locking, fixed worktree paths, and summary metadata outputs.
- Updated `remote.add` handling with URL-aware idempotency and per-repo queueing for serialized git mutations.
- Added temp-home test scaffolding and expanded tests for batched PR flows and remote add conflict/no-op cases.
- Added a `/context` slash command flow from registry to interactive-mode command dispatch.
- Added `handleContextCommand()` to the mode context interface and command-controller wiring.
- Added context usage breakdown utilities, cell allocation, and 20x10 usage rendering for token categories.
- Reworked compaction token estimation to use tokenizer counts, role aggregation, image token estimates, and fallback handling.
- Exported `resolveThresholdTokens()` as a public compaction helper.
- Updated loop command handling to toggle `/loop` without a prompt argument and prompt for the next user input to start repeating.
- Stored each user-submitted prompt as the active loop prompt while loop mode is enabled so iterations auto-resubmit that input after each yield.
- Introduced `pauseLoop` behavior to clear the captured loop prompt and cancel pending auto-submit when Escape is pressed, while `handleLoopCommand` now no longer disables mode automatically with arguments.
- Tracked locally submitted user signatures for both optimistic and streamed-queue submissions in interactive-mode context state.
- Updated user message_start handling to avoid re-clearing the editor and re-adding chat for locally originated messages.
- Cleared consumed local signatures on queued-message restore and added tests for queued, external, and optimistic message_start editor behavior.
- Added a new `loop.mode` enum setting and UI option entries for prompt, compact, and reset behaviors.
- Updated interactive mode's loop auto-submit flow to execute selected compact or reset actions before re-submitting the prompt.
- Registered a new `just.enabled` setting in the tools settings schema.
- Added loop-mode command handling in interactive mode, including enable/disable toggling, repeated auto-submission scheduling, and Escape-based cancellation.
- Propagated loop mode state to the status line via a renamed `mode` segment that now renders plan or loop status, and updated presets/theme assets for the new loop icon.
- Migrated persisted status-line segment usage from `plan_mode` to `mode` in configuration defaults and normalization so legacy settings continue to load.
The L2 render cache used objectId(this.#theme) as part of the key.
After setTheme(), the same Markdown instance still holds the old
theme reference (same objectId) but the closures now read a different
palette — so the cache returned pre-switch styled lines (F3).
Export clearRenderCache() from markdown.ts and call it from the
onThemeChange handler in interactive-mode.ts so all L2 entries are
dropped on every theme switch, forcing a clean re-render.
Before this change, every navigateTree → renderInitialMessages call path
performed two independent O(N) session-tree walks:
1. agent-session.ts:6586 buildDisplaySessionContext() [inside navigateTree]
2. ui-helpers.ts:402 sessionManager.buildSessionContext() [inside renderInitialMessages]
Changes:
- agent-session.ts: navigateTree() now calls sessionManager.buildSessionContext()
once, derives the display (deobfuscated) context from the raw result, and
returns the raw SessionContext in the result object.
- ui-helpers.ts: renderInitialMessages() accepts an optional prebuiltContext
parameter; reuses it when provided, falls back to buildSessionContext() otherwise.
- interactive-mode.ts: forwards prebuiltContext through the wrapper.
- modes/types.ts: updates InteractiveModeContext interface to match.
- selector-controller.ts: passes result.sessionContext from navigateTree into
renderInitialMessages(), closing the deduplication loop.
Bench (100-msg session, 200 iterations):
two walks [BEFORE]: 0.0702ms/op
one walk [AFTER]: 0.0298ms/op
Saved: 0.0404ms/navigation (57.5% reduction per navigate)
Tests: render-initial-messages-dedupe.test.ts asserts buildSessionContext is
called 0 times when a prebuilt context is passed, 1 time as fallback.
/drop works like /new but permanently deletes the current session file
and artifacts instead of flushing (saving) it. Useful when the session
should not be kept.
- add drop?: boolean to NewSessionOptions
- branch in AgentSession.newSession(): skip flush, delete via
FileSessionStorage.deleteSessionWithArtifacts when drop=true;
deletion failure is non-fatal (logged, new session still starts)
- wire handleDropCommand() through InteractiveModeContext interface,
InteractiveMode delegation, and CommandController implementation
- guard: shows error if session has not been saved yet (no file to drop)
- register /drop in builtin-registry adjacent to /new
- Added `note` support to todo-write with `op: "note"` and required `text` input.
- Added optional `notes: string[]` to todo models and preserved notes in cloning and session task mapping.
- Implemented `op: "note"` append flow and markdown `>` block serialization/parsing for todo import/export.
- Updated HUD todo rendering to append superscript `+N` note markers and show in-progress note bodies.
- Documented note operation, required text field, and note rendering rules in todo-write docs and changelog.
Extract duplicate normalizeLocalScheme regex pattern into a shared function in path-utils.ts. Updated interactive-mode.ts, approved-plan.ts, agent-session.ts, bash-skill-urls.ts, and plan-mode-guard.ts to use the shared utility. Also fixed error message formatting (removed extra backslashes).
On Linux, Node's path.normalize() collapses the double slash in
local://PLAN.md to local:/PLAN.md, creating a directory called local:
in the project root instead of routing through the local:// protocol handler.
Defense-in-depth fixes across 5 layers:
1. resolveToCwd() now throws if a path starts with any internal URL
scheme prefix (local:, agent:, skill:, etc.), preventing all 59
call sites from treating URIs as relative filesystem paths.
2. resolvePlanPath() now matches on local: prefix (not just local://)
and normalizes local:/ to local:// before resolution, catching
all slash variants.
3. Bash URL expansion regex and early-exit checks now also match
local:/ (single slash), and normalize before resolution.
4. Edit preview/diff functions now gracefully skip internal URL paths
instead of crashing via the resolveToCwd guard.
5. All startsWith('local://') checks updated to startsWith('local:')
with normalization in agent-session, interactive-mode, and
approved-plan modules.
Also adds local: to .gitignore to prevent accidental commits of the
leaked directory.
- Propagated session title source through terminal-title update calls so auto and user names are handled consistently across controllers.
- Suppressed auto-generated names in status-line segments and completion messages by falling back to cwd-based titles.
- Updated title formatter tests to verify auto-generated and user-specified session titles produce different terminal labels.