- Added an environment configuration option to choose between docker and apple-container backends.
- Updated the server to default to apple-container when the container CLI is detected.
- Removed unused auto_compaction_end event handling in the coding-agent print mode.
- Sanitized CLI JSON output by stripping sensitive payload data and implementing incremental delta extraction.
- Optimized log parsing in the harbor-manager by using tail-scanning and streaming file handles to prevent OOM errors.
- Enhanced process management to support persistent runs across manager restarts and improved orphan process handling.
- Added incremental cost probing for live trials and verified status transitions for runner lifecycle events.
- Remove the `gen:docs` step from the Docker build process and the developer documentation.
- Update documentation in `docs-index.ts` to reflect that the index is now injected via environment variable instead of a generated file.
- Update error messaging to point to binary or bundle rebuilds rather than the deprecated generation script.
- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.
- Refactored experiment tests to reflect the migration from slide to downshift configurations.
- Updated test expectations in manager and runner to validate correct downstream command argument generation.
- Added a legacy compatibility check to ensure existing slide configurations are still handled correctly in summarization.
- Added comprehensive tests for downshifting model behavior, including plan nudge injection and completion safety mechanisms.
- Verified session persistence accuracy by validating that from-disk rebuilds match the live agent state.
- Confirmed correct cache key propagation during tan commands to ensure provider caching is preserved.
- Removed obsolete reasoning slide tests.
- Replaced legacy reasoning-slide functionality with new downshift and plan-yolo capabilities.
- Updated CLI arguments, slash commands, and configuration schemas to support the new model-switching and execution behaviors.
- Refactored agent session logic to handle downshift arming, plan-yolo headless execution, and context scrubbing.
- Renamed and added system prompts to align with the updated downshift and plan-yolo workflows.
- Clear inherited todo list state and persist empty edit at fork creation to prevent parent task reminders from affecting the tangential session.
- Re-inject the fork notice after each auto-compaction event to ensure the boundary between the parent and child session survives history summarization.
- Align the provider cache key with the parent's actual pinned key to correctly mirror cached session context.
- Update `AgentSession` to perform a full entry rewrite during tool result pruning to ensure session files match pruned state for reliable resuming and branching.
- Implemented a consolidated TUI renderer for launch tool operations to unify status headers and metadata.
- Added tracking for unfulfilled readiness conditions to distinguish between timeout causes and daemon states.
- Enhanced timeout reporting to explicitly surface specific unmet log or port conditions.
- Included comprehensive unit and regression tests for tool rendering and start-timeout diagnostics.
- Implemented bsd-to-gnu flag translation for `date`, `find`, `mktemp`, `sed`, and `tail` utilities to support native bsd command usage.
- Adjusted argument parsing logic in `ln` and `base32` to accommodate bsd-compatible short-flag aliases and behavior.
- Added comprehensive regression tests and bsd-specific compatibility suites across modified utilities to ensure consistent behavior.
- Updated project documentation to reflect the expanded cross-platform utility support.
- Added `rewrite_bsd_invocation` to detect macOS-style `stat -f FORMAT` syntax, preventing misinterpretation as GNU filesystem mode.
- Implemented `translate_bsd_format` to map BSD-specific format directives to GNU equivalents, including support for time forms, user/group IDs, and file metadata.
- Enabled BSD flag cluster parsing (`-LnfqF`) to allow common macOS command patterns to function correctly within the in-process `stat` utility.
- Added comprehensive unit tests to validate BSD syntax translation, sub-field handling, and error reporting for unsupported directives.
- Added a system prompt to define boundaries for tangential agent forks.
- Injected the context-switch notice into new clones to prevent task leakage from parent sessions.
- Added a configuration setting `launch.enabled` to control the availability of the project-scoped launch tool.
- Integrated the setting into the bash tool and tool registry to conditionally enable or disable the launch functionality.
- Updated documentation and settings schema to include the new toggle.
- Introduced a project-scoped `launch` tool to orchestrate long-running services, debuggers, and watchers with persistent execution capabilities.
- Implemented a robust daemon broker with Unix/Windows IPC transport that manages process lifecycles, readiness monitoring, and automatic log rotation.
- Added detached process support to ensure services persist independently of the main application lifecycle, including recovery and cleanup mechanisms.
- Updated the `bash` interceptor to prioritize the new `launch` tool for background processes and provided comprehensive documentation for lifecycle and signal handling.
- Added support for role switching via `@` search prefix in the model picker.
- Implemented `quickRoles` configuration and logic to handle role-specific selections via the session API.
- Updated the model browser to support preserving query order and custom label coloring for improved navigation.
- Integrated role cycle tracking and updated UI hints to support the new role-browsing workflow.
- Redesigned agent hub entries as two-line cards, separating identity and status from task descriptions.
- Added explicit model and thinking level badges to agent status information.
- Updated agent hub rendering to support multi-line entries and adaptive vertical scrolling based on row height.
- Simplified status display by using glyphs instead of redundant status labels.
- Adjusted rendering logic to prioritize displaying agent task summaries on an independent indented line.
- Added test coverage for auto-compaction entry warning attribution during dead-end scenarios.
- Verified automatic continuation and warning suppression when image-drop recovery successfully clears context pressure.
- Mocked session context usage and image-drop functionality to simulate high-usage state recovery.
- Added `StrippedToolCallsMarker` to identify assistant messages with dangling tool calls that were removed during context building.
- Updated `buildSessionContext` to preserve turn entries in transcripts even when all content is stripped, marking them with the count of elided calls.
- Updated `UiHelpers` to render a dim, italicized placeholder in the TUI when an assistant message has elided tool calls, preventing silent activity gaps.
- Added `warning` field to `CompactionEntry` and `CompactionSummaryMessage` to persist dead-end status in session history.
- Introduced a multi-tier rescue mechanism in `AgentSession` that automatically performs `elide` and `dropImages` passes when maintenance fails to recover sufficient headroom.
- Integrated visual indicators for dead-ends into the `CompactionSummaryMessageComponent`, surfacing warnings directly on the compaction divider and detail block.
- Updated `AgentSession` logic to re-evaluate progress after each rescue tier and emit recovery notices, ensuring transparent reporting of automated history rewrites.
- Replaced markdown-style headings with concise `¶user:`, `¶think:`, `¶ai:`, and `¶call:` scope markers.
- Updated the serializer to merge consecutive messages or blocks of the same type under a shared prefix.
- Updated the documentation prompt to reflect the new compact formatting and scope rules.
- Added regression tests verifying scope merging and correct formatting of tool calls and intents.
- Updated transcript session context to use display.collapseCompacted setting instead of a hardcoded boolean.
- Enabled user-defined control over whether compacted history is collapsed during live rendering sessions.
- Introduced conditional scrollback clearing during UI renders when transcript compaction is enabled.
- Updated `CommandController` and `EventController` to respect the `display.collapseCompacted` setting.
- Configured `SelectorController` to trigger a chat rebuild and UI reset when the compaction setting changes.
- Updated `InteractiveMode` to dynamically toggle between collapsed and full inline history based on user settings.
- Added display.collapseCompacted boolean setting to the appearance tab.
- Enabled default collapse behavior for pre-compaction history in the transcript view.
- Update in-flight context snapshots after historical messages are modified by compaction or elision.
- Prevent stale run-start token counts from triggering false-positive dead-end "no progress" warnings during auto-continuation.
- Add regression test to verify that prompt headroom measurements correctly account for post-compaction context sizes.
- Initialized session metadata including system prompt, task, and toolset within the controller.
- Added session initialization tracking to the clone creation flow to ensure session state visibility.
- Updated unit tests to verify that session initialization data is correctly appended when a tan is created.
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
- Stop propagating real-time updates for backgrounded Bash jobs to avoid UI flickering once a job enters the background.
- Refine background task tracking in `EventController` to distinguish between persistent background tasks and transient backgrounded Bash commands.
- Update UI rendering to display cleaner background job metadata in the footer instead of inline text notices.
- Removed unreliable Bing and Yahoo HTML-scraping search providers.
- Deleted `src/web/search/providers/bing.ts` and `src/web/search/providers/yahoo.ts` implementation files.
- Updated `provider.ts`, `types.ts`, and `public.ts` to prune provider registration and configuration.
- Adjusted `web-search-public.test.ts` to exclude removed engines from test coverage.
- Introduced ModelPickerComponent to enable temporary, floating model selection overlays.
- Centralized role management logic using a new resolveRoleAssignments function in the browser.
- Streamlined model hub by removing legacy pick-mode logic and defaulting to fullscreen views.
- Synchronized model picker state with the registry to ensure accurate offline refreshes.
- Refactor status event display to show the most recent events as the live edge.
- Replace the bottom-trailing "more" indicator with a leading "earlier" marker when events exceed the display limit.
- Dynamically adjust the expanded display size based on the available preview window rows.
- Implemented cell budget clamping for timeouts to prevent stalled browser operations from exceeding execution limits.
- Added `recover` capabilities for tab workers to clear blocking dialogs and safely terminate hung navigations.
- Introduced `//!world=main` support for `tab.evaluate` via Puppeteer patch to allow execution in main execution contexts.
- Improved failure attribution for tab terminations by tracking dialogs, stalled operations, and specific termination reasons.
- Shifted testing expectations from mandatory inclusion to proof-based verification depending on the task type.
- Clarified that experiments, UI changes, and investigations do not require tests unless specific conditions are met.
- Reorganized the verification workflow to prioritize behavioral smoke testing over automated test execution for non-permanent changes.
- Enhanced `wait()` in browser tools to accept a predicate function in addition to milliseconds.
- Implemented automatic polling with configurable timeout and interval, resolving with the first truthy value.
- Throws a descriptive `ToolError` on timeout instead of waiting for the full execution deadline.
- Added comprehensive unit tests to verify polling behavior, cancellation, and error handling.
- pickWeightedTip now takes the tip list and a uniform sample, exported for tests.
- Weighted-selection test sweeps a synthetic tip list, so shipping zero [NEW] tips no longer fails the suite.
- Updated import rewriting to identify and publish `var` and `function` declarations to the global scope when a cell contains top-level `await`.
- Prevented these declarations from being trapped within the async wrapper's function scope, allowing them to remain accessible to subsequent evaluation cells.
- Added support for home, end, and page navigation in the model browser.
- Restricted the hover background band to mouse interactions, using cursor glyphs and text accents for keyboard selection instead.
- Synchronized focus states between the model hub sidebar and the browser pane.
- Removed redundant "login" labels from locked provider entries.
- Prevented navigation keys from triggering the input-priority grace period, ensuring responsive movement after idle states.
- Limited the input queue-drain delay to Ctrl+C and Escape double-press gestures.
- Compat tests assert the exported SCHEMA_VERSION instead of a stale
hardcoded 5; the perf-sample change bumped agent.db to v6.
- Backfill cap fixture inserts its 300 rows in one transaction; per-row
implicit transactions fsynced 300 times and timed out on CI runners.