- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.
- Added comprehensive tests for downshifting model behavior, including plan nudge injection and completion safety mechanisms.
- Verified session persistence accuracy by validating that from-disk rebuilds match the live agent state.
- Confirmed correct cache key propagation during tan commands to ensure provider caching is preserved.
- Removed obsolete reasoning slide tests.
- Clear inherited todo list state and persist empty edit at fork creation to prevent parent task reminders from affecting the tangential session.
- Re-inject the fork notice after each auto-compaction event to ensure the boundary between the parent and child session survives history summarization.
- Align the provider cache key with the parent's actual pinned key to correctly mirror cached session context.
- Update `AgentSession` to perform a full entry rewrite during tool result pruning to ensure session files match pruned state for reliable resuming and branching.
- Implemented a consolidated TUI renderer for launch tool operations to unify status headers and metadata.
- Added tracking for unfulfilled readiness conditions to distinguish between timeout causes and daemon states.
- Enhanced timeout reporting to explicitly surface specific unmet log or port conditions.
- Included comprehensive unit and regression tests for tool rendering and start-timeout diagnostics.
- Added a configuration setting `launch.enabled` to control the availability of the project-scoped launch tool.
- Integrated the setting into the bash tool and tool registry to conditionally enable or disable the launch functionality.
- Updated documentation and settings schema to include the new toggle.
- Introduced a project-scoped `launch` tool to orchestrate long-running services, debuggers, and watchers with persistent execution capabilities.
- Implemented a robust daemon broker with Unix/Windows IPC transport that manages process lifecycles, readiness monitoring, and automatic log rotation.
- Added detached process support to ensure services persist independently of the main application lifecycle, including recovery and cleanup mechanisms.
- Updated the `bash` interceptor to prioritize the new `launch` tool for background processes and provided comprehensive documentation for lifecycle and signal handling.
- Added support for role switching via `@` search prefix in the model picker.
- Implemented `quickRoles` configuration and logic to handle role-specific selections via the session API.
- Updated the model browser to support preserving query order and custom label coloring for improved navigation.
- Integrated role cycle tracking and updated UI hints to support the new role-browsing workflow.
- Redesigned agent hub entries as two-line cards, separating identity and status from task descriptions.
- Added explicit model and thinking level badges to agent status information.
- Updated agent hub rendering to support multi-line entries and adaptive vertical scrolling based on row height.
- Simplified status display by using glyphs instead of redundant status labels.
- Adjusted rendering logic to prioritize displaying agent task summaries on an independent indented line.
- Added test coverage for auto-compaction entry warning attribution during dead-end scenarios.
- Verified automatic continuation and warning suppression when image-drop recovery successfully clears context pressure.
- Mocked session context usage and image-drop functionality to simulate high-usage state recovery.
- Update in-flight context snapshots after historical messages are modified by compaction or elision.
- Prevent stale run-start token counts from triggering false-positive dead-end "no progress" warnings during auto-continuation.
- Add regression test to verify that prompt headroom measurements correctly account for post-compaction context sizes.
- Initialized session metadata including system prompt, task, and toolset within the controller.
- Added session initialization tracking to the clone creation flow to ensure session state visibility.
- Updated unit tests to verify that session initialization data is correctly appended when a tan is created.
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
- Stop propagating real-time updates for backgrounded Bash jobs to avoid UI flickering once a job enters the background.
- Refine background task tracking in `EventController` to distinguish between persistent background tasks and transient backgrounded Bash commands.
- Update UI rendering to display cleaner background job metadata in the footer instead of inline text notices.
- Removed unreliable Bing and Yahoo HTML-scraping search providers.
- Deleted `src/web/search/providers/bing.ts` and `src/web/search/providers/yahoo.ts` implementation files.
- Updated `provider.ts`, `types.ts`, and `public.ts` to prune provider registration and configuration.
- Adjusted `web-search-public.test.ts` to exclude removed engines from test coverage.
- Introduced ModelPickerComponent to enable temporary, floating model selection overlays.
- Centralized role management logic using a new resolveRoleAssignments function in the browser.
- Streamlined model hub by removing legacy pick-mode logic and defaulting to fullscreen views.
- Synchronized model picker state with the registry to ensure accurate offline refreshes.
- Implemented cell budget clamping for timeouts to prevent stalled browser operations from exceeding execution limits.
- Added `recover` capabilities for tab workers to clear blocking dialogs and safely terminate hung navigations.
- Introduced `//!world=main` support for `tab.evaluate` via Puppeteer patch to allow execution in main execution contexts.
- Improved failure attribution for tab terminations by tracking dialogs, stalled operations, and specific termination reasons.
- Enhanced `wait()` in browser tools to accept a predicate function in addition to milliseconds.
- Implemented automatic polling with configurable timeout and interval, resolving with the first truthy value.
- Throws a descriptive `ToolError` on timeout instead of waiting for the full execution deadline.
- Added comprehensive unit tests to verify polling behavior, cancellation, and error handling.
- pickWeightedTip now takes the tip list and a uniform sample, exported for tests.
- Weighted-selection test sweeps a synthetic tip list, so shipping zero [NEW] tips no longer fails the suite.
- Updated import rewriting to identify and publish `var` and `function` declarations to the global scope when a cell contains top-level `await`.
- Prevented these declarations from being trapped within the async wrapper's function scope, allowing them to remain accessible to subsequent evaluation cells.
- Added support for home, end, and page navigation in the model browser.
- Restricted the hover background band to mouse interactions, using cursor glyphs and text accents for keyboard selection instead.
- Synchronized focus states between the model hub sidebar and the browser pane.
- Removed redundant "login" labels from locked provider entries.
- Prevented navigation keys from triggering the input-priority grace period, ensuring responsive movement after idle states.
- Limited the input queue-drain delay to Ctrl+C and Escape double-press gestures.
- Compat tests assert the exported SCHEMA_VERSION instead of a stale
hardcoded 5; the perf-sample change bumped agent.db to v6.
- Backfill cap fixture inserts its 300 rows in one transaction; per-row
implicit transactions fsynced 300 times and timed out on CI runners.
- Permit model fallback even if the retry budget is exhausted when the current provider is locked by credential rotation or usage limits.
- Reset the retry budget when successfully switching to a fallback model to ensure the new model has a full allowance of retries.
- Added a regression test to verify that credential rotation failure triggers model fallback.
- Verified throughput (TPS) and latency (TTFT) aggregation logic for multiple samples.
- Confirmed data sanitization by dropping invalid samples like NaN, zero duration, or missing values.
- Ensured backfilling correctly filters stale, errored, and empty turns from the stats database.
- Validated that backfill operations respect the maximum sample limit per model.
- Validated the async deferral behavior in recording perf samples.
- Integrated performance statistics (TPS/TTFT) into ModelBrowser and ModelHub components with adaptive visibility.
- Decoupled mouse wheel behavior from selection logic to enable smoother viewport panning.
- Improved UI navigation by preventing unwanted wrapping in role indexes and adding viewport snapping logic.
- Added comprehensive test suites to verify performance display rendering and corrected input interaction behavior.
- Prioritized fuzzy match quality over the most-recently-used model order when filtering the model browser.
- Added bucketing to match scores to ensure stable sorting when match quality is identical.
- Added unit tests to verify that exact query matches take precedence over the MRU model.
- Implemented `/queue` command and `->`/`=>` shorthands to support deferred, sequential message processing.
- Added a robust parsing utility to handle various list-based queue inputs and automate yield management.
- Integrated visual decorations and state tracking to provide real-time feedback on queueing status.
- Enabled non-cursor line text decoration in the TUI to support dynamic queue header rendering and list numbering.
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.
- Prevent selected chips from being hidden when the role strip row overflows.
- Implement horizontal scrolling with a leading ellipsis to maintain visibility of the selection and lookahead context.
- Forwarded the session provider state map to ephemeral side-channel turns so Codex can allocate websocket transport state for the side request session id.
- Extended the /btw side-channel regression to assert providerSessionState reaches the stream options with the websocket preference.
Fixes#5213
- Preserved the session websocket preference for ephemeral /btw side-channel turns so Codex websocket-only models do not fall back to SSE.
- Prioritized active /btw and /omfg panels in Esc handling before loop, maintenance, and main-turn interrupts.
- Added regression coverage for side-channel websocket options and /btw Escape priority.
Fixes#5213
- Added `conflict://*` support to the `write` tool, allowing resolution of multiple conflicts in a single call using per-id directives.
- Implemented `parseBulkDirectives` to interpret `ID: @side` mappings from the raw input content.
- Enabled partial bulk resolution where unlisted conflict IDs remain registered for subsequent operations.
- Updated conflict documentation and tool summaries to reflect the new bulk resolution capability.
- Enabled persistence of dangling tool calls during transcript rebuilds by adding `keepDanglingToolCalls` configuration.
- Integrated `seal()` logic for pending tool blocks to properly finalize history during idle session states.
- Enhanced UI helpers and session context to preserve assistant turns during streaming or mid-turn rebuilds.
- Added comprehensive unit and integration tests to verify correct tool call tracking and transcript integrity.
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
- Update `IrcBus.wait` to accept a `liveness` configuration, automatically aborting waits if the specified sender or relevant peers transition to an idle state.
- Enhance `IrcTool` to utilize liveness monitoring during `wait` operations, ensuring the agent does not hang if target peers become unreachable.
- Refactor `IrcBus.wait` to unify internal settlement logic and improve cleanup robustness.
- Implemented logic to automatically detect and trim redundant lines duplicating adjacent file content within conflict markers.
- Added delimiter balancing and boundary tracking to ensure accurate removal of echoed text while preserving EOL formatting.
- Updated user feedback to report the number of trimmed echo lines during write and conflict resolution operations.
- Expanded test coverage to include multi-line echo scenarios and integration validation of the repair process.
- Integrated a horizontal separator in the model browser list to distinguish between frequently used or role-assigned models and other available models.
- Removed the "Recent" category from the Model Hub sidebar to streamline navigation.
- Updated related tests to reflect the removal of the Recent category and the new visual list layout.
Validated custom tool factory results before registering plugin-provided tools so one malformed feature entry is reported and skipped instead of crashing startup.
Added regression coverage for null entries and mixed valid/null factory arrays.
Fixes#5189
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.