- Exported ensureChromiumExecutable so the test can probe launchability.
- CI runner holds the downloaded Chrome but lacks libnspr4 & co., so the
binary fails at dynamic-link time; probe --version and skipIf instead
of failing the release run.
- Added `tui.scrollbackRebuild` configuration with interactive startup/controller wiring to apply `setScrollbackRebuild`.
- Exposed prewalk session state in `SegmentContext` and rendered a dedicated prewalk segment/icon in the status line.
- Added divergence-aware TUI full-paint logic that enables scrollback erase-and-replay rebuilds for non-multiplexer divergence cases.
- Updated rendering and streaming tests to verify rebuild behavior (`3J`) and eliminate stale marker expectations under drift scenarios.
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Changed daemon log reads to return both sanitized display text and a raw `terminalText` slice, and included it on log RPC responses for PTY runs when grep was not used.
- Extended the logs result contract and launch tool rendering to consume `terminalText`, reconstruct terminal output, and display it in framed, preview-capped sections.
- Kept terminal row layout stable by writing space characters for empty cells when reading rows, preserving spacing during output reconstruction.
- Bounded expanded partial edit diff rendering in `formatStreamingDiff` to `previewWindowRows()` instead of an unbounded budget, preventing runaway preview growth during live updates.
- Updated streaming diff tests to simulate terminal height and verify expanded previews stay full only within the viewport, then switch to a truncated tail with the `more lines above` marker when too tall.
- Reinitialized in-memory `Settings` before each initial-messages test since the test suite reads global display configuration and needs isolation.
- Added `isAcceptedElicitation` in the ACP agent to narrow accepted elicitations before accessing response content.
- Added `isFormElicitation` in ACP tests and used it to narrow form-mode requests before assertions.
- Initialized in-memory settings in the repro issue test by resetting settings before theme setup.
- Reduced the UI/TUI CI test bucket chunk size from 10 to 5 to avoid cumulative Bun GC heap aborts.
- Removed the Google Interactions transport and deleted interaction-specific request options from the shared AI stream typing/API surface.
- Simplified Google provider routing to eliminate interactions auto-selection logic and keep `streamGoogle` on the `:streamGenerateContent` path.
- Updated Vertex request handling to use resolved stream hosts without `/interactions`/`Api-Revision` and removed related interaction constants.
- Deleted obsolete Interactions tests and updated remaining Google stream tests to no longer reference `useInteractionsApi`/`storeInteraction`/`previousInteractionId`.
- Verify terminal screen row replay correctly handles cursor rewrites.
- Ensure final text color and weight are preserved during terminal output rendering.
- Assert that superseded text is excluded from the rendered output.
- Updated `hashline` test expectations to match new recovery message strings.
- Initialized settings in `interactive-mode-status` tests to prevent state leakage.
- Added ANSI color code verification to `launch` tool tests to align with updated output.
- Cleaned up whitespace in `tan-context-switch.md` system prompt.
- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.
- Added comprehensive tests for downshifting model behavior, including plan nudge injection and completion safety mechanisms.
- Verified session persistence accuracy by validating that from-disk rebuilds match the live agent state.
- Confirmed correct cache key propagation during tan commands to ensure provider caching is preserved.
- Removed obsolete reasoning slide tests.
- Clear inherited todo list state and persist empty edit at fork creation to prevent parent task reminders from affecting the tangential session.
- Re-inject the fork notice after each auto-compaction event to ensure the boundary between the parent and child session survives history summarization.
- Align the provider cache key with the parent's actual pinned key to correctly mirror cached session context.
- Update `AgentSession` to perform a full entry rewrite during tool result pruning to ensure session files match pruned state for reliable resuming and branching.
- Implemented a consolidated TUI renderer for launch tool operations to unify status headers and metadata.
- Added tracking for unfulfilled readiness conditions to distinguish between timeout causes and daemon states.
- Enhanced timeout reporting to explicitly surface specific unmet log or port conditions.
- Included comprehensive unit and regression tests for tool rendering and start-timeout diagnostics.
- Added a configuration setting `launch.enabled` to control the availability of the project-scoped launch tool.
- Integrated the setting into the bash tool and tool registry to conditionally enable or disable the launch functionality.
- Updated documentation and settings schema to include the new toggle.
- Introduced a project-scoped `launch` tool to orchestrate long-running services, debuggers, and watchers with persistent execution capabilities.
- Implemented a robust daemon broker with Unix/Windows IPC transport that manages process lifecycles, readiness monitoring, and automatic log rotation.
- Added detached process support to ensure services persist independently of the main application lifecycle, including recovery and cleanup mechanisms.
- Updated the `bash` interceptor to prioritize the new `launch` tool for background processes and provided comprehensive documentation for lifecycle and signal handling.
- Added support for role switching via `@` search prefix in the model picker.
- Implemented `quickRoles` configuration and logic to handle role-specific selections via the session API.
- Updated the model browser to support preserving query order and custom label coloring for improved navigation.
- Integrated role cycle tracking and updated UI hints to support the new role-browsing workflow.
- Redesigned agent hub entries as two-line cards, separating identity and status from task descriptions.
- Added explicit model and thinking level badges to agent status information.
- Updated agent hub rendering to support multi-line entries and adaptive vertical scrolling based on row height.
- Simplified status display by using glyphs instead of redundant status labels.
- Adjusted rendering logic to prioritize displaying agent task summaries on an independent indented line.
- Added test coverage for auto-compaction entry warning attribution during dead-end scenarios.
- Verified automatic continuation and warning suppression when image-drop recovery successfully clears context pressure.
- Mocked session context usage and image-drop functionality to simulate high-usage state recovery.
- Update in-flight context snapshots after historical messages are modified by compaction or elision.
- Prevent stale run-start token counts from triggering false-positive dead-end "no progress" warnings during auto-continuation.
- Add regression test to verify that prompt headroom measurements correctly account for post-compaction context sizes.
- Initialized session metadata including system prompt, task, and toolset within the controller.
- Added session initialization tracking to the clone creation flow to ensure session state visibility.
- Updated unit tests to verify that session initialization data is correctly appended when a tan is created.
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
- Stop propagating real-time updates for backgrounded Bash jobs to avoid UI flickering once a job enters the background.
- Refine background task tracking in `EventController` to distinguish between persistent background tasks and transient backgrounded Bash commands.
- Update UI rendering to display cleaner background job metadata in the footer instead of inline text notices.
- Removed unreliable Bing and Yahoo HTML-scraping search providers.
- Deleted `src/web/search/providers/bing.ts` and `src/web/search/providers/yahoo.ts` implementation files.
- Updated `provider.ts`, `types.ts`, and `public.ts` to prune provider registration and configuration.
- Adjusted `web-search-public.test.ts` to exclude removed engines from test coverage.
- Introduced ModelPickerComponent to enable temporary, floating model selection overlays.
- Centralized role management logic using a new resolveRoleAssignments function in the browser.
- Streamlined model hub by removing legacy pick-mode logic and defaulting to fullscreen views.
- Synchronized model picker state with the registry to ensure accurate offline refreshes.
- Implemented cell budget clamping for timeouts to prevent stalled browser operations from exceeding execution limits.
- Added `recover` capabilities for tab workers to clear blocking dialogs and safely terminate hung navigations.
- Introduced `//!world=main` support for `tab.evaluate` via Puppeteer patch to allow execution in main execution contexts.
- Improved failure attribution for tab terminations by tracking dialogs, stalled operations, and specific termination reasons.
- Enhanced `wait()` in browser tools to accept a predicate function in addition to milliseconds.
- Implemented automatic polling with configurable timeout and interval, resolving with the first truthy value.
- Throws a descriptive `ToolError` on timeout instead of waiting for the full execution deadline.
- Added comprehensive unit tests to verify polling behavior, cancellation, and error handling.
- pickWeightedTip now takes the tip list and a uniform sample, exported for tests.
- Weighted-selection test sweeps a synthetic tip list, so shipping zero [NEW] tips no longer fails the suite.
- Updated import rewriting to identify and publish `var` and `function` declarations to the global scope when a cell contains top-level `await`.
- Prevented these declarations from being trapped within the async wrapper's function scope, allowing them to remain accessible to subsequent evaluation cells.
- Added support for home, end, and page navigation in the model browser.
- Restricted the hover background band to mouse interactions, using cursor glyphs and text accents for keyboard selection instead.
- Synchronized focus states between the model hub sidebar and the browser pane.
- Removed redundant "login" labels from locked provider entries.
- Prevented navigation keys from triggering the input-priority grace period, ensuring responsive movement after idle states.
- Limited the input queue-drain delay to Ctrl+C and Escape double-press gestures.
Keep ordinary 401 retries bounded while replay-safe quota failures
walk every distinct eligible credential. Anchor blocks to the failed
credential and stop on cycles, aborts, or 64 attempts.
- Compat tests assert the exported SCHEMA_VERSION instead of a stale
hardcoded 5; the perf-sample change bumped agent.db to v6.
- Backfill cap fixture inserts its 300 rows in one transaction; per-row
implicit transactions fsynced 300 times and timed out on CI runners.
- Permit model fallback even if the retry budget is exhausted when the current provider is locked by credential rotation or usage limits.
- Reset the retry budget when successfully switching to a fallback model to ensure the new model has a full allowance of retries.
- Added a regression test to verify that credential rotation failure triggers model fallback.
- Verified throughput (TPS) and latency (TTFT) aggregation logic for multiple samples.
- Confirmed data sanitization by dropping invalid samples like NaN, zero duration, or missing values.
- Ensured backfilling correctly filters stale, errored, and empty turns from the stats database.
- Validated that backfill operations respect the maximum sample limit per model.
- Validated the async deferral behavior in recording perf samples.
- Integrated performance statistics (TPS/TTFT) into ModelBrowser and ModelHub components with adaptive visibility.
- Decoupled mouse wheel behavior from selection logic to enable smoother viewport panning.
- Improved UI navigation by preventing unwanted wrapping in role indexes and adding viewport snapping logic.
- Added comprehensive test suites to verify performance display rendering and corrected input interaction behavior.
- Prioritized fuzzy match quality over the most-recently-used model order when filtering the model browser.
- Added bucketing to match scores to ensure stable sorting when match quality is identical.
- Added unit tests to verify that exact query matches take precedence over the MRU model.
- Implemented `/queue` command and `->`/`=>` shorthands to support deferred, sequential message processing.
- Added a robust parsing utility to handle various list-based queue inputs and automate yield management.
- Integrated visual decorations and state tracking to provide real-time feedback on queueing status.
- Enabled non-cursor line text decoration in the TUI to support dynamic queue header rendering and list numbering.
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.