Commit Graph
4835 Commits
Author SHA1 Message Date
DarkPhilosophy 4db2d9617b Merge remote-tracking branch 'can1357/main' into feat/advisor-per-agent-toggle 2026-07-13 23:16:26 +03:00
can1357 f405525bf4 feat: removed boomerang feature and associated validation workflows
- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
2026-07-13 14:15:54 +02:00
can1357 590270ca26 feat(coding-agent): gated downshift trigger on todo initialization
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
2026-07-13 10:37:20 +02:00
can1357 95ecc61bc9 feat(coding-agent): implemented downshift boomerang flow for context handoff
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
2026-07-13 10:37:20 +02:00
can1357 4d019e5617 feat(coding-agent): trigger downshift on post-plan todo initialization
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.
2026-07-13 10:36:39 +02:00
can1357 e42589d43d test(coding-agent): validated session persistence and downshift logic
- Added comprehensive tests for downshifting model behavior, including plan nudge injection and completion safety mechanisms.
- Verified session persistence accuracy by validating that from-disk rebuilds match the live agent state.
- Confirmed correct cache key propagation during tan commands to ensure provider caching is preserved.
- Removed obsolete reasoning slide tests.
2026-07-13 06:03:48 +02:00
can1357 0a98aa252b feat(coding-agent/modes): hardened tan fork isolation and session sync
- Clear inherited todo list state and persist empty edit at fork creation to prevent parent task reminders from affecting the tangential session.
- Re-inject the fork notice after each auto-compaction event to ensure the boundary between the parent and child session survives history summarization.
- Align the provider cache key with the parent's actual pinned key to correctly mirror cached session context.
- Update `AgentSession` to perform a full entry rewrite during tool result pruning to ensure session files match pruned state for reliable resuming and branching.
2026-07-13 06:00:30 +02:00
can1357 e770cdc4d9 fix(coding-agent): unified tool renderer and clarified launch diagnostics
- Implemented a consolidated TUI renderer for launch tool operations to unify status headers and metadata.
- Added tracking for unfulfilled readiness conditions to distinguish between timeout causes and daemon states.
- Enhanced timeout reporting to explicitly surface specific unmet log or port conditions.
- Included comprehensive unit and regression tests for tool rendering and start-timeout diagnostics.
2026-07-13 05:16:36 +02:00
can1357 f83e409218 feat(coding-agent): added toggle for launch tool
- Added a configuration setting `launch.enabled` to control the availability of the project-scoped launch tool.
- Integrated the setting into the bash tool and tool registry to conditionally enable or disable the launch functionality.
- Updated documentation and settings schema to include the new toggle.
2026-07-13 04:22:28 +02:00
can1357 4cfec93458 feat(coding-agent/launch): introduced persistent project service tool
- Introduced a project-scoped `launch` tool to orchestrate long-running services, debuggers, and watchers with persistent execution capabilities.
- Implemented a robust daemon broker with Unix/Windows IPC transport that manages process lifecycles, readiness monitoring, and automatic log rotation.
- Added detached process support to ensure services persist independently of the main application lifecycle, including recovery and cleanup mechanisms.
- Updated the `bash` interceptor to prioritize the new `launch` tool for background processes and provided comprehensive documentation for lifecycle and signal handling.
2026-07-13 04:21:55 +02:00
can1357 af7345e87e feat(coding-agent): implemented role switching and filtering in model picker
- Added support for role switching via `@` search prefix in the model picker.
- Implemented `quickRoles` configuration and logic to handle role-specific selections via the session API.
- Updated the model browser to support preserving query order and custom label coloring for improved navigation.
- Integrated role cycle tracking and updated UI hints to support the new role-browsing workflow.
2026-07-13 04:21:54 +02:00
can1357 ac16253613 wip: rslide (experimental) 2026-07-13 04:21:54 +02:00
can1357 42d81f189f feat(coding-agent): redesigned agent hub layout for task clarity
- Redesigned agent hub entries as two-line cards, separating identity and status from task descriptions.
- Added explicit model and thinking level badges to agent status information.
- Updated agent hub rendering to support multi-line entries and adaptive vertical scrolling based on row height.
- Simplified status display by using glyphs instead of redundant status labels.
- Adjusted rendering logic to prioritize displaying agent task summaries on an independent indented line.
2026-07-13 03:03:26 +02:00
can1357 22f2c1947b test(session): validated auto-compaction guard and recovery behavior
- Added test coverage for auto-compaction entry warning attribution during dead-end scenarios.
- Verified automatic continuation and warning suppression when image-drop recovery successfully clears context pressure.
- Mocked session context usage and image-drop functionality to simulate high-usage state recovery.
2026-07-13 01:41:09 +02:00
can1357 bf5eb3769f feat(coding-agent): rebased pending context snapshot after compaction
- Update in-flight context snapshots after historical messages are modified by compaction or elision.
- Prevent stale run-start token counts from triggering false-positive dead-end "no progress" warnings during auto-continuation.
- Add regression test to verify that prompt headroom measurements correctly account for post-compaction context sizes.
2026-07-13 01:28:18 +02:00
can1357 46ed33f27b feat(coding-agent): initialized session metadata during tan creation
- Initialized session metadata including system prompt, task, and toolset within the controller.
- Added session initialization tracking to the clone creation flow to ensure session state visibility.
- Updated unit tests to verify that session initialization data is correctly appended when a tan is created.
2026-07-13 01:09:42 +02:00
can1357 58d6130b50 feat(coding-agent): enabled model fallback for hard errors
- Extended `AgentSession` to consult `retry.fallbackChains` on non-retryable (hard) model errors.
- Implemented `#isHardErrorFallbackEligible` to validate eligibility for model switching before surfacing terminal errors.
- Updated `#handleRetryableError` to orchestrate immediate model switching for hard errors, bypassing backoff-retries for the failing model.
- Ensured hard errors propagate to the user if no fallback candidates are available or if no credential can be resolved for the fallback model.
2026-07-13 00:59:43 +02:00
can1357 87a64b2f6a feat(coding-agent): improved background job lifecycle and display
- Stop propagating real-time updates for backgrounded Bash jobs to avoid UI flickering once a job enters the background.
- Refine background task tracking in `EventController` to distinguish between persistent background tasks and transient backgrounded Bash commands.
- Update UI rendering to display cleaner background job metadata in the footer instead of inline text notices.
2026-07-13 00:54:51 +02:00
can1357 d0f90f35ae refactor(coding-agent): removed unreliable web search providers
- Removed unreliable Bing and Yahoo HTML-scraping search providers.
- Deleted `src/web/search/providers/bing.ts` and `src/web/search/providers/yahoo.ts` implementation files.
- Updated `provider.ts`, `types.ts`, and `public.ts` to prune provider registration and configuration.
- Adjusted `web-search-public.test.ts` to exclude removed engines from test coverage.
2026-07-13 00:34:27 +02:00
can1357 8dbc43b6e8 feat(coding-agent-modes): implemented floating model selection system
- Introduced ModelPickerComponent to enable temporary, floating model selection overlays.
- Centralized role management logic using a new resolveRoleAssignments function in the browser.
- Streamlined model hub by removing legacy pick-mode logic and defaulting to fullscreen views.
- Synchronized model picker state with the registry to ensure accurate offline refreshes.
2026-07-13 00:13:04 +02:00
can1357 8c8afaf477 fix(browser): made tab evaluation use the main world 2026-07-13 00:05:43 +02:00
can1357 0d07da529a feat(coding-agent-tools): implemented browser execution safety controls
- Implemented cell budget clamping for timeouts to prevent stalled browser operations from exceeding execution limits.
- Added `recover` capabilities for tab workers to clear blocking dialogs and safely terminate hung navigations.
- Introduced `//!world=main` support for `tab.evaluate` via Puppeteer patch to allow execution in main execution contexts.
- Improved failure attribution for tab terminations by tracking dialogs, stalled operations, and specific termination reasons.
2026-07-13 00:05:43 +02:00
can1357 bd7d395222 feat(coding-agent/tools): added predicate polling support to wait()
- Enhanced `wait()` in browser tools to accept a predicate function in addition to milliseconds.
- Implemented automatic polling with configurable timeout and interval, resolving with the first truthy value.
- Throws a descriptive `ToolError` on timeout instead of waiting for the full execution deadline.
- Added comprehensive unit tests to verify polling behavior, cancellation, and error handling.
2026-07-13 00:05:42 +02:00
can1357 6bd51d4ad3 refactor(coding-agent): decoupled tip weight test from tips.txt data
- pickWeightedTip now takes the tip list and a uniform sample, exported for tests.
- Weighted-selection test sweeps a synthetic tip list, so shipping zero [NEW] tips no longer fails the suite.
2026-07-12 20:46:29 +02:00
can1357 b6f83021c9 fix(coding-agent): ensured top-level declarations persist in async cells
- Updated import rewriting to identify and publish `var` and `function` declarations to the global scope when a cell contains top-level `await`.
- Prevented these declarations from being trapped within the async wrapper's function scope, allowing them to remain accessible to subsequent evaluation cells.
2026-07-12 13:21:50 +02:00
can1357 1822603b2d feat(coding-agent): improved model browser keyboard navigation and focus visuals
- Added support for home, end, and page navigation in the model browser.
- Restricted the hover background band to mouse interactions, using cursor glyphs and text accents for keyboard selection instead.
- Synchronized focus states between the model hub sidebar and the browser pane.
- Removed redundant "login" labels from locked provider entries.
- Prevented navigation keys from triggering the input-priority grace period, ensuring responsive movement after idle states.
- Limited the input queue-drain delay to Ctrl+C and Escape double-press gestures.
2026-07-12 13:05:01 +02:00
Hayden Evanandcan1357 4fa5b61b05 Add plan review copy hotkey 2026-07-12 12:43:10 +02:00
can1357 20c0a2e410 test(coding-agent): fixed storage tests for schema v6 and slow CI disks
- Compat tests assert the exported SCHEMA_VERSION instead of a stale
  hardcoded 5; the perf-sample change bumped agent.db to v6.
- Backfill cap fixture inserts its 300 rows in one transaction; per-row
  implicit transactions fsynced 300 times and timed out on CI runners.
2026-07-12 03:13:15 +02:00
can1357 d54dcc2224 feat(coding-agent): allowed model fallback after retry budget exhaustion
- Permit model fallback even if the retry budget is exhausted when the current provider is locked by credential rotation or usage limits.
- Reset the retry budget when successfully switching to a fallback model to ensure the new model has a full allowance of retries.
- Added a regression test to verify that credential rotation failure triggers model fallback.
2026-07-12 01:35:50 +02:00
can1357 41317cc231 test(perf): validated performance sample aggregation and backfill logic
- Verified throughput (TPS) and latency (TTFT) aggregation logic for multiple samples.
- Confirmed data sanitization by dropping invalid samples like NaN, zero duration, or missing values.
- Ensured backfilling correctly filters stale, errored, and empty turns from the stats database.
- Validated that backfill operations respect the maximum sample limit per model.
- Validated the async deferral behavior in recording perf samples.
2026-07-12 01:06:27 +02:00
can1357 a0dcb8ae20 feat(ui): enabled performance monitoring and viewport navigation
- Integrated performance statistics (TPS/TTFT) into ModelBrowser and ModelHub components with adaptive visibility.
- Decoupled mouse wheel behavior from selection logic to enable smoother viewport panning.
- Improved UI navigation by preventing unwanted wrapping in role indexes and adding viewport snapping logic.
- Added comprehensive test suites to verify performance display rendering and corrected input interaction behavior.
2026-07-12 01:06:22 +02:00
can1357 666327608c feat(coding-agent): improved model search ranking by match quality
- Prioritized fuzzy match quality over the most-recently-used model order when filtering the model browser.
- Added bucketing to match scores to ensure stable sorting when match quality is identical.
- Added unit tests to verify that exact query matches take precedence over the MRU model.
2026-07-12 00:00:35 +02:00
can1357 e7955ddf3c feat(coding-agent): introduced sequential message queueing and commands
- Implemented `/queue` command and `->`/`=>` shorthands to support deferred, sequential message processing.
- Added a robust parsing utility to handle various list-based queue inputs and automate yield management.
- Integrated visual decorations and state tracking to provide real-time feedback on queueing status.
- Enabled non-cursor line text decoration in the TUI to support dynamic queue header rendering and list numbering.
2026-07-11 22:07:51 +02:00
can1357 54bafa1cce feat(coding-agent): implemented interactive fallback chain configuration
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.
2026-07-11 22:03:35 +02:00
can1357 2ab2da2b1d fix(coding-agent): improved role-assignment strip visibility on overflow
- Prevent selected chips from being hidden when the role strip row overflows.
- Implement horizontal scrolling with a leading ellipsis to maintain visibility of the selection and lookahead context.
2026-07-11 21:17:35 +02:00
DarkPhilosophy ca8bf8dda1 Merge remote-tracking branch 'can1357/main' into feat/advisor-per-agent-toggle
# Conflicts:
#	packages/ai/test/pi-native-client.test.ts
#	packages/coding-agent/src/modes/components/advisor-config.ts
#	packages/coding-agent/src/session/agent-session.ts
2026-07-11 22:13:08 +03:00
can1357 7957d0b88c Merge remote-tracking branch 'origin/farm/3b0815b6/fix-btw-gpt-luna-esc' 2026-07-11 20:38:37 +02:00
can1357 a643e94466 chore: remove redundant tests 2026-07-11 19:50:23 +02:00
roboomp 9269823d16 fix(coding-agent): shared codex state for btw websockets
- Forwarded the session provider state map to ephemeral side-channel turns so Codex can allocate websocket transport state for the side request session id.
- Extended the /btw side-channel regression to assert providerSessionState reaches the stream options with the websocket preference.

Fixes #5213
2026-07-11 17:46:28 +00:00
roboomp 2c161d2a8f fix(coding-agent): preserved btw codex websocket routing
- Preserved the session websocket preference for ephemeral /btw side-channel turns so Codex websocket-only models do not fall back to SSE.
- Prioritized active /btw and /omfg panels in Esc handling before loop, maintenance, and main-turn interrupts.
- Added regression coverage for side-channel websocket options and /btw Escape priority.

Fixes #5213
2026-07-11 17:31:37 +00:00
can1357 6f47df88b4 Merge remote-tracking branch 'origin/farm/84fe7c9e/harden-plugin-feature-load' 2026-07-11 19:18:56 +02:00
can1357 d39a3ed453 chore: fix stale tests 2026-07-11 19:10:40 +02:00
can1357 7a0ae70313 feat(coding-agent/tools): implemented bulk conflict resolution via conflict://*
- Added `conflict://*` support to the `write` tool, allowing resolution of multiple conflicts in a single call using per-id directives.
- Implemented `parseBulkDirectives` to interpret `ID: @side` mappings from the raw input content.
- Enabled partial bulk resolution where unlisted conflict IDs remain registered for subsequent operations.
- Updated conflict documentation and tool summaries to reflect the new bulk resolution capability.
2026-07-11 18:23:57 +02:00
can1357 5b20a7dea4 feat(coding-agent): implemented persistence for tool calls during rebuilds
- Enabled persistence of dangling tool calls during transcript rebuilds by adding `keepDanglingToolCalls` configuration.
- Integrated `seal()` logic for pending tool blocks to properly finalize history during idle session states.
- Enhanced UI helpers and session context to preserve assistant turns during streaming or mid-turn rebuilds.
- Added comprehensive unit and integration tests to verify correct tool call tracking and transcript integrity.
2026-07-11 18:19:48 +02:00
can1357 33b6774aa1 feat(coding-agent): implemented resumable subagent yielding for tasks
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
2026-07-11 17:27:35 +02:00
can1357 0828c53cab feat(coding-agent): integrated liveness monitoring into irc wait operations
- Update `IrcBus.wait` to accept a `liveness` configuration, automatically aborting waits if the specified sender or relevant peers transition to an idle state.
- Enhance `IrcTool` to utilize liveness monitoring during `wait` operations, ensuring the agent does not hang if target peers become unreachable.
- Refactor `IrcBus.wait` to unify internal settlement logic and improve cleanup robustness.
2026-07-11 17:14:03 +02:00
can1357 7e5e7e864d feat(coding-agent-tools): implemented auto-trimming for echo lines
- Implemented logic to automatically detect and trim redundant lines duplicating adjacent file content within conflict markers.
- Added delimiter balancing and boundary tracking to ensure accurate removal of echoed text while preserving EOL formatting.
- Updated user feedback to report the number of trimmed echo lines during write and conflict resolution operations.
- Expanded test coverage to include multi-line echo scenarios and integration validation of the repair process.
2026-07-11 17:04:23 +02:00
can1357 c1e9a84112 feat(coding-agent): added visual separator to model browser list
- Integrated a horizontal separator in the model browser list to distinguish between frequently used or role-assigned models and other available models.
- Removed the "Recent" category from the Model Hub sidebar to streamline navigation.
- Updated related tests to reflect the removal of the Recent category and the new visual list layout.
2026-07-11 16:43:13 +02:00
roboomp 459682cc63 fix(plugins): skipped invalid custom tool entries
Validated custom tool factory results before registering plugin-provided tools so one malformed feature entry is reported and skipped instead of crashing startup.

Added regression coverage for null entries and mixed valid/null factory arrays.

Fixes #5189
2026-07-11 14:31:07 +00:00
can1357 408a92d91a feat(coding-agent): enabled asynchronous background task execution
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
2026-07-11 16:17:03 +02:00