- Added `tui.scrollbackRebuild` configuration with interactive startup/controller wiring to apply `setScrollbackRebuild`.
- Exposed prewalk session state in `SegmentContext` and rendered a dedicated prewalk segment/icon in the status line.
- Added divergence-aware TUI full-paint logic that enables scrollback erase-and-replay rebuilds for non-multiplexer divergence cases.
- Updated rendering and streaming tests to verify rebuild behavior (`3J`) and eliminate stale marker expectations under drift scenarios.
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Removed the Google Interactions transport and deleted interaction-specific request options from the shared AI stream typing/API surface.
- Simplified Google provider routing to eliminate interactions auto-selection logic and keep `streamGoogle` on the `:streamGenerateContent` path.
- Updated Vertex request handling to use resolved stream hosts without `/interactions`/`Api-Revision` and removed related interaction constants.
- Deleted obsolete Interactions tests and updated remaining Google stream tests to no longer reference `useInteractionsApi`/`storeInteraction`/`previousInteractionId`.
- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.
- Replaced legacy reasoning-slide functionality with new downshift and plan-yolo capabilities.
- Updated CLI arguments, slash commands, and configuration schemas to support the new model-switching and execution behaviors.
- Refactored agent session logic to handle downshift arming, plan-yolo headless execution, and context scrubbing.
- Renamed and added system prompts to align with the updated downshift and plan-yolo workflows.
- Clear inherited todo list state and persist empty edit at fork creation to prevent parent task reminders from affecting the tangential session.
- Re-inject the fork notice after each auto-compaction event to ensure the boundary between the parent and child session survives history summarization.
- Align the provider cache key with the parent's actual pinned key to correctly mirror cached session context.
- Update `AgentSession` to perform a full entry rewrite during tool result pruning to ensure session files match pruned state for reliable resuming and branching.
- Added test coverage for auto-compaction entry warning attribution during dead-end scenarios.
- Verified automatic continuation and warning suppression when image-drop recovery successfully clears context pressure.
- Mocked session context usage and image-drop functionality to simulate high-usage state recovery.
- Added `StrippedToolCallsMarker` to identify assistant messages with dangling tool calls that were removed during context building.
- Updated `buildSessionContext` to preserve turn entries in transcripts even when all content is stripped, marking them with the count of elided calls.
- Updated `UiHelpers` to render a dim, italicized placeholder in the TUI when an assistant message has elided tool calls, preventing silent activity gaps.
- Added `warning` field to `CompactionEntry` and `CompactionSummaryMessage` to persist dead-end status in session history.
- Introduced a multi-tier rescue mechanism in `AgentSession` that automatically performs `elide` and `dropImages` passes when maintenance fails to recover sufficient headroom.
- Integrated visual indicators for dead-ends into the `CompactionSummaryMessageComponent`, surfacing warnings directly on the compaction divider and detail block.
- Updated `AgentSession` logic to re-evaluate progress after each rescue tier and emit recovery notices, ensuring transparent reporting of automated history rewrites.
- Update in-flight context snapshots after historical messages are modified by compaction or elision.
- Prevent stale run-start token counts from triggering false-positive dead-end "no progress" warnings during auto-continuation.
- Add regression test to verify that prompt headroom measurements correctly account for post-compaction context sizes.
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
- Compat tests assert the exported SCHEMA_VERSION instead of a stale
hardcoded 5; the perf-sample change bumped agent.db to v6.
- Backfill cap fixture inserts its 300 rows in one transaction; per-row
implicit transactions fsynced 300 times and timed out on CI runners.
- Permit model fallback even if the retry budget is exhausted when the current provider is locked by credential rotation or usage limits.
- Reset the retry budget when successfully switching to a fallback model to ensure the new model has a full allowance of retries.
- Added a regression test to verify that credential rotation failure triggers model fallback.
- Added `model_perf` table to store persistent, recency-weighted TPS and TTFT statistics per model.
- Introduced `AgentStorage.recordModelPerf` to asynchronously capture and aggregate turn timing metrics.
- Implemented a background backfill mechanism to migrate historical request data from `stats.db` into new performance aggregates.
- Updated `AgentSession` to record performance metrics upon turn completion.
- Bumped `SCHEMA_VERSION` to 6 and cleared existing performance data to account for improved duration-based TPS calculation.
- Relocated `AsyncDrain` class from `coding-agent` to `utils` package.
- Updated `HistoryStorage` to import `AsyncDrain` from shared utilities.
- Centralized the utility to allow reuse across the codebase.
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.
- Updated task synchronization to include all todo phases regardless of completion status.
- Removed logic that filtered out completed or abandoned tasks during session synchronization.
- Forwarded the session provider state map to ephemeral side-channel turns so Codex can allocate websocket transport state for the side request session id.
- Extended the /btw side-channel regression to assert providerSessionState reaches the stream options with the websocket preference.
Fixes#5213
- Preserved the session websocket preference for ephemeral /btw side-channel turns so Codex websocket-only models do not fall back to SSE.
- Prioritized active /btw and /omfg panels in Esc handling before loop, maintenance, and main-turn interrupts.
- Added regression coverage for side-channel websocket options and /btw Escape priority.
Fixes#5213
- Enabled persistence of dangling tool calls during transcript rebuilds by adding `keepDanglingToolCalls` configuration.
- Integrated `seal()` logic for pending tool blocks to properly finalize history during idle session states.
- Enhanced UI helpers and session context to preserve assistant turns during streaming or mid-turn rebuilds.
- Added comprehensive unit and integration tests to verify correct tool call tracking and transcript integrity.
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
- Transitioned vibe tools to an ephemeral registration model where they are installed only when entering `/vibe` mode and removed upon exit.
- Added `activateVibeTools` and `deactivateVibeTools` methods to `AgentSession` to manage these transient tool registrations.
- Removed vibe tools from the default global tool registry, preventing unnecessary background exposure.
- Implemented Vibe mode to enable worker session management and director-role context injection.
- Added a `/vibe` slash command and integrated status line UI to display mode activity.
- Configured restricted toolsets and guards to prevent concurrent conflicts with existing Goal or Plan modes.
- Provided system prompts and tool templates to support specialized agent communication and task orchestration.
Emitted a failed auto-retry event when empty assistant stop responses exhaust their retry cap so the TUI can show an error instead of silently settling.
Fixes#5128
- Introduced `stringifyJson` helper to preserve bigint precision by serializing them as decimal strings.
- Replaced native `JSON.stringify` across compaction and session management modules to prevent serialization errors when handling bigint values in tool arguments.
- Added regression tests in `agent` and `coding-agent` packages to ensure bigint tool arguments remain intact through compaction and persistence flows.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
Separated advisor provider session identity from local advisor labels so Codex requests carry stable UUIDv7 values while transcripts keep their advisor-specific names.
Fixes#5040
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.
- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.
- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.
Fixes#5035