- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
Hidden slider means the operator made no choice; a singleton cycle built around the active plan model must not be pinned as executionModel, otherwise approval re-applies the plan model after #exitPlanMode restored the pre-plan one.
Added regression coverage for the plan-only role configuration.
Refs #3554
Same-model role with an explicit thinking suffix that differs from the pre-plan thinking now passes through applyRoleModel instead of being treated as an implicit match.
Added regression coverage for the sonnet:off vs pre-plan thinking-high case.
Refs #3554
Compared the selected approval tier against the model restored after plan mode instead of the active plan-mode tier.
Added regression coverage for keeping the active planning model selected on approval.
Fixes#3554
- Added a `mode` property to `CompactOptions` to allow fine-grained control over compaction strategies.
- Implemented `soft`, `remote`, and `snapcompact` submode overrides for the `/compact` command.
- Integrated `parseCompactArgs` to enable robust subcommand routing and validation, including focus instruction rejection for specific modes.
- Established a `CompactMode` registry to manage compaction strategies and verify remote availability.
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
Addresses Codex review on #2520:
- Cancel path: `#approvePlan` returned on `compactOutcome === "cancelled"`
without restoring the deferred pre-plan model, stranding the next turn on
the plan model and leaking `#planModePreviousModelState`. The model
transition now runs for the cancelled outcome too (the operator aborted
only the compaction, not the approval) before the early return.
- Queue-flush ordering: `executeCompaction` flushes input queued during
compaction before returning, so the post-return model switch landed after
the queued turn began streaming (deferred one turn via #pendingModelSwitch).
Added a `beforeFlush(outcome)` hook to `executeCompaction`/`handleCompactCommand`;
`#approvePlan` runs the transition through it (and idempotently re-runs it
afterward to cover the message-count short-circuit).
Tests cover the cancel restore and the before-flush ordering.
Plan approval dispatched the executor's first synthetic prompt without
checking whether the agent was still streaming the post-resolve
continuation (or a turn started by the approve-time compaction/clear),
surfacing "Failed to finalize approved plan: ... Agent is already
processing". Loop auto-submit and goal continuations hit the same throw
via submitInteractiveInput, which always called prompt/promptCustomMessage
without a streamingBehavior.
#approvePlan now aborts any in-flight turn before the synthetic prompt,
and submitInteractiveInput routes submissions through the steer/follow-up
queue (streamingBehavior: "followUp") when the session is streaming.
Non-streaming call shapes are unchanged. Extends the manual-/goal fix
(#2454) to the continuation and plan-approval paths.
When a plan is approved via "Approve and compact context", #exitPlanMode
restored the pre-plan model before the compaction summarizer ran, so the
request cold-missed the plan model's warm prompt cache. The model switch is
now deferred: compaction runs on the plan model, and the switch to the
execution (slider) or pre-plan model happens only after a successful
compaction. A failed compaction stays on the plan model; cancellation is
unchanged. Also clears any queued plan-role model switch when deferring the
restore so it cannot later clobber the restored model.
- Added an `openPlanReview` flow that selected the newest `local://<slug>-plan.md`, resolved its title, and reopened approval.
- Registered a new `/plan-review` builtin slash command that invokes that flow and clears the editor text.
- Added tests covering latest-plan selection plus warnings when plan mode is inactive or when no local plan file exists.
- Added `isUserInterruptAbort` and `shouldRenderAbortReason` helpers so interrupt handling can distinguish Esc-based aborts from other abort reasons.
- Updated assistant transcript rendering to suppress the `Interrupted by user` line while continuing to show generic or other abort labels.
- Updated plan-review and transcript container tests to assert interrupted assistant messages no longer render the redundant interrupt line.
- Replaced approved-plan renaming with `resolveApprovedPlan` resolution and state/slug lookup.
- Updated ACP and interactive apply flows to propagate canonical `planFilePath` instead of renamed paths.
- Added local plan fallback lookup by mtime for unresolved slugs after plan approval.
- Restricted plan-mode writes to `local://` plan artifacts and simplified path handling.
- Added optional `reason` parameters to `Agent.abort` and `AgentSession.abort` APIs.
- Passed abort reasons through interrupt flows into underlying agent cancellation.
- Replaced hard-coded abort text with `resolveAbortLabel` for streaming and replayed messages.
- Fell back to generic `Request was aborted` text when no abort reason was provided.
- Added section-based plan parsing with per-section delete, undo, and annotate.
- Added Tab/Shift+Tab focus regions and per-line scroll navigation.
- Added Refine feedback loop emitting annotations back to the model.
- Added `OverlayOptions.fullscreen` borrowing the terminal's alt screen buffer.
- Replaced the plan review flow to open PlanReviewOverlay for approvals.
- Added scrollable markdown plan rendering with prompt, options, and footer in the overlay UI.
- Added disabled-option handling so cursor movement and confirmation skip unavailable rows.
- Added setPlanContent updates to refresh overlay text and reset scroll position on edits.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
- Aligned handoff, reminder, and system-prompt expectations with shortened copy.
- Added HTTP transport test for required initialize failures.
- Guarded SSE startup timeout against stale connection races.
- Updated interactive-mode plan review tests to capture shared fixtures, clear references, and run cleanup with explicit garbage collection before disposal.
- Increased the MCP HTTP transport test connection timeout from 200ms to 1,000ms.
- Adjusted the tool streaming command delay and tightened a start-pending-submission spy type in tests for better stability and type accuracy.
- Updated local module loading to force TS syntax stripping for .ts/.tsx/.mts modules and use the matching Bun transpiler loader.
- Extended TypeScript stripping to detect `import type`/`export type` syntax and rewired wrapping to strip TS syntax after final-expression extraction with TypeScript-aware parsing.
- Added tests verifying type-only imports are handled correctly in evaluator modules and rewritten code no longer contains type-only import declarations.
- Captured the operator's selected execution tier and passed it through plan approval options.
- Applied the stored model after exiting plan mode so #exitPlanMode's restore no longer reverts it before execution begins.
- Added a regression test that selects a different slider tier and verifies execution runs on the chosen model.
- Removed Ghostty-specific hardware-cursor forcing from TUI preference resolution and dropped the redundant terminal-cursor marker flag.
- Updated interactive mode editors to use `ui.getShowHardwareCursor()` so cursor mode now follows actual hardware-cursor visibility.
- Reworked terminal regressions tests to assert Ghostty respects the requested cursor preference and only emits cursor-show output when enabled.
- Rendered the plan review "keep context" selector label with the session context usage percentage when available.
- Kept the fallback label unchanged when context usage data was unavailable.
- Added coverage asserting both the percentage label and fallback label in plan review tests.
Fixes#1458
Patch axis: extend
Displacement: net-zero; reuses existing plan reference state instead of adding persistence or overwriting approved artifacts
Rule violations averted: no approved-plan overwrite, no transcript format migration, no public CLI/API expansion
PASS/FAIL: PASS after plan-mode focused tests and package check. Note: system-prompt-templates has an unrelated HOME=/tmp path-shortening expectation failure.
- Removed ExitPlanModeTool and deleted exit-plan-mode docs/tests, dropping the old approval contract outputs.
- Replaced plan-mode approval flow from exit_plan_mode to resolve across session, SDK, controllers, and discovery.
- Added standing resolve handler accessors and updated resolve routing for queued or standing approval handlers.
- Added PlanApprovalDetails and enforced normalized, validated approval titles with readable plan-file requirements.
- Extended resolve schema and invocation signatures with optional extra metadata and reason trimming behavior updates.
- Updated plan and resolve prompts and changelog guidance to require resolve action, reason, and extra.title for apply/discard.
- Updated conflict URI parsing to accept `path:conflict://N` and record the removed prefix in `recoveredPrefix`.
- Updated write conflict handling to resolve single or wildcard IDs through shared helpers and append a recovery note when a malformed prefix was stripped.
- Added regression tests for recovered prefixes and end-to-end write-path recovery and documented the change in the changelog.
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
- Pinned final plan path before handleCompactCommand so queued messages use approved plan, not draft.
- Added regression coverage for setPlanReferencePath timing before compaction queue flush.
- Extend the existing selector-list assertion to include the new
"Approve and compact context" entry in the expected five-choice array.
- Add three new cases that mock `handleCompactCommand` at the
CompactionOutcome boundary (the contract `#approvePlan` consumes):
- "ok" → asserts compactSpy called with the planning-specific
custom instruction rendered from `plan-mode-compact-instructions.md`
(matched by content: "Preparing to execute the approved plan" +
the final plan file path); plan-approved synthetic prompt dispatched
(discriminated by the `{ synthetic: true }` option flag);
`markPlanReferenceSent` called once.
- "cancelled" → no synthetic dispatch; `showWarning` surfaces the
deferred-dispatch message; `setPlanReferencePath` IS called with the
final path so the session knows the plan was approved; but
`markPlanReferenceSent` is NOT called so the next operator turn
re-injects the reference via #buildPlanReferenceMessage. This last
assertion is the load-bearing regression guard for the asymmetric
cancel/fail contract.
- "failed" → plan-approved synthetic prompt IS dispatched
(best-effort); markPlanReferenceSent fires.
Per AGENTS.md the tests use `vi.spyOn` + `vi.restoreAllMocks()` in
the existing afterEach; no `mock.module()`.
Op: extend
- Render each exit_plan_mode submission as a fresh plan review entry so refined plans are emitted into terminal scrollback.
- Keep external-editor edits replacing the current preview instead of adding duplicate review entries.
- Updated the plan review regression test to preserve the first preview and append the second one after intervening chat content.
- Rendered diff content before the truncation notice instead of after.
- Updated hidden lines label to show count with "+" prefix.
- Added fallback label when no lines are hidden.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Fixed plan review previews to re-append at the chat tail on refresh, keeping them adjacent to the active selector instead of updating off-screen.
- Added comprehensive test coverage for plan review rendering behavior across multiple refresh cycles.