- pickWeightedTip now takes the tip list and a uniform sample, exported for tests.
- Weighted-selection test sweeps a synthetic tip list, so shipping zero [NEW] tips no longer fails the suite.
- Added support for home, end, and page navigation in the model browser.
- Restricted the hover background band to mouse interactions, using cursor glyphs and text accents for keyboard selection instead.
- Synchronized focus states between the model hub sidebar and the browser pane.
- Removed redundant "login" labels from locked provider entries.
- Prevented navigation keys from triggering the input-priority grace period, ensuring responsive movement after idle states.
- Limited the input queue-drain delay to Ctrl+C and Escape double-press gestures.
- Removed the `computeTokensPerSecond` helper function in favor of a direct calculation.
- Standardized throughput to use total request duration instead of post-TTFT decode time.
- Updated UI usage display to calculate tokens per second against total duration.
- Prevented over-inflation of throughput metrics caused by hidden reasoning tokens.
- Integrated performance statistics (TPS/TTFT) into ModelBrowser and ModelHub components with adaptive visibility.
- Decoupled mouse wheel behavior from selection logic to enable smoother viewport panning.
- Improved UI navigation by preventing unwanted wrapping in role indexes and adding viewport snapping logic.
- Added comprehensive test suites to verify performance display rendering and corrected input interaction behavior.
AgentSession defers and coalesces the wire-level agent_end while a
prompt is in flight (#emitSessionEvent), so a multi-attempt retry saga
often surfaces only ONE agent_end to EventController — which can be
the final settle, not an intermediate attempt. Consuming #retryPending
against whichever agent_end arrived first (previous commit) could
therefore discard the real final failure notification.
Switch to gating purely on the retry lifecycle: #retryPending is set
by auto_retry_start and cleared only by auto_retry_end (both
outcomes), never consumed by sendErrorNotification itself. Those
lifecycle events are never deferred, so they reliably bracket the
window a retry is actually outstanding regardless of how agent_end
coalescing lands.
Close the residual gap this creates: #handleRetryableError's
classifier-refusal and Fireworks-fallback-ineligible branches could
short-circuit a saga that already announced auto_retry_start without
ever emitting auto_retry_end, latching #retryPending open forever.
Both branches now emit a final auto_retry_end(false) when a prior
attempt already started the saga. #handleAgentStart also clears
#retryPending defensively so a saga that still somehow never resolves
cannot suppress a later, unrelated turn's notification.
A retryable error's agent_end fires with the failed assistant message
(stopReason === 'error') the instant #handleRetryableError schedules a
retry (auto_retry_start), before the retry has a chance to recover.
sendErrorNotification read that transient agent_end the same as a real
final settle, so error.notify=on raised a 'Stopped with error' toast
even for turns that went on to succeed on retry.
Track a #retryPending flag set on auto_retry_start and consumed
(check-then-clear) by sendErrorNotification, so exactly one mid-retry
agent_end is suppressed per attempt. #handleAutoRetryEnd also clears it
directly on both success and failure so a recovered retry never leaves
it stuck; consuming it on read (rather than only via auto_retry_end)
keeps it self-healing for the classifier-refusal short-circuit path,
which can return from #handleRetryableError without ever emitting a
fresh auto_retry_end.
- Prioritized fuzzy match quality over the most-recently-used model order when filtering the model browser.
- Added bucketing to match scores to ensure stable sorting when match quality is identical.
- Added unit tests to verify that exact query matches take precedence over the MRU model.
- Implemented `/queue` command and `->`/`=>` shorthands to support deferred, sequential message processing.
- Added a robust parsing utility to handle various list-based queue inputs and automate yield management.
- Integrated visual decorations and state tracking to provide real-time feedback on queueing status.
- Enabled non-cursor line text decoration in the TUI to support dynamic queue header rendering and list numbering.
- Implemented model-specific keys and provider wildcards for `retry.fallbackChains` with updated resolution logic.
- Added interactive fallback chain management in the model roles UI, including support for reordering and editing.
- Improved fallback chain specificity rules and added comprehensive validation with startup warnings.
- Fixed mouse interaction alignment and hover state coordinate mapping in the roles view.
- Prevent selected chips from being hidden when the role strip row overflows.
- Implement horizontal scrolling with a leading ellipsis to maintain visibility of the selection and lookahead context.
- Updated task synchronization to include all todo phases regardless of completion status.
- Removed logic that filtered out completed or abandoned tasks during session synchronization.
- Preserved the session websocket preference for ephemeral /btw side-channel turns so Codex websocket-only models do not fall back to SSE.
- Prioritized active /btw and /omfg panels in Esc handling before loop, maintenance, and main-turn interrupts.
- Added regression coverage for side-channel websocket options and /btw Escape priority.
Fixes#5213
- Enabled persistence of dangling tool calls during transcript rebuilds by adding `keepDanglingToolCalls` configuration.
- Integrated `seal()` logic for pending tool blocks to properly finalize history during idle session states.
- Enhanced UI helpers and session context to preserve assistant turns during streaming or mid-turn rebuilds.
- Added comprehensive unit and integration tests to verify correct tool call tracking and transcript integrity.
- Included the current session name in the fullscreen pause UI.
- Updated `renderPauseScreen` and `InteractiveModeContext` to propagate and display the session title.
- Added test coverage for both full and compact pause screen layouts.
- Introduced an `AgentPauseGate` mechanism to suspend and resume agent loops and tool executions safely.
- Added a `/pause` slash command to trigger a new fullscreen UI that manages agent suspension and lifecycle.
- Integrated pause checks into the core agent loop and tool execution pipeline to ensure responsive state handling.
- Provided a new pause screen component to facilitate user interaction and resume control during suspension.
- Integrated a horizontal separator in the model browser list to distinguish between frequently used or role-assigned models and other available models.
- Removed the "Recent" category from the Model Hub sidebar to streamline navigation.
- Updated related tests to reflect the removal of the Recent category and the new visual list layout.
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
- Implemented tiered search strategy using synchronous literal matching for immediate results and debounced asynchronous fuzzy matching to prevent input blocking.
- Cached session search data and introduced an indexed `FuzzyText` structure to reduce redundant string processing.
- Added comprehensive test suite verifying search convergence, stale task orphaning, and selection stability.
- Enhanced `tui` fuzzy matching utilities to support prebuilt search indexes for improved multi-match performance.
Addresses the second review round (internal re-review + Codex on c38840482):
- Active-account matching (logout preselection, /usage in-use marker) is
org-decisive when EITHER side carries an org: a legacy bare-email active
row no longer flags org-scoped siblings via the shared email (reverse of
the previous fix). Both-org-less keeps the email/account fallback, so
providers without orgs are unaffected.
- Status-line usage context key includes orgId, so rotating between two
same-email subscriptions invalidates the cached quota immediately
instead of showing the previous org's numbers for the cache TTL.
- CredentialHealthResult carries orgId/orgName and auth-gateway check
labels rows with the org, so a failing row names the subscription.
- getOAuthAccountIdentity preserves org-only identities; the login
success message renders them.
- ACP /usage account-id fallback labels get the org suffix too.
- Regression tests for both matching directions (marker + logout).
One Anthropic account email can hold multiple organizations (a Team seat
plus a personal Max plan), each with its own org-scoped OAuth token and
independent 5h/7d limit pools. Credentials were deduped by bare email, so
logging in with the second subscription silently replaced the first, and
usage reports from the two pools merged into one row with mixed numbers.
- capture organization uuid/name at login (token exchange response, with
a claude_cli/bootstrap fallback); token refreshes never rewrite it
- key anthropic credential identity as email + org; a legacy email-keyed
row is claimed in place by the first org-scoped login with the same
email, and org-less credentials never clobber org-scoped rows
- partition usage-report dedupe and the per-credential usage cache by
org so the two subscriptions' limit pools stay distinct for rotation
- show the organization in omp usage (redaction-safe) and name the
stored account/org in the login success message
- Implemented custom role creation within the Model Hub, including a virtual row for direct initiation and name stripping.
- Added quick-switch cycle editing functionality with persistent ordering and live preview of role membership.
- Optimized sidebar scope navigation to mute empty entries and improve keyboard focus stability during search.
- Updated the Model Hub sidebar UI to prioritize Roles and included comprehensive tests for new navigation and management flows.
- Implemented dynamic reordering of model providers in the sidebar to prioritize matches during active searches.
- Refactored model hub registry synchronization to maintain separate state for fixed, unlocked, and locked provider entries.
- Added a dedicated composition step to reassemble sidebar entries based on current search result counts.
- Added test coverage to verify that provider order reverts to alphabetical upon clearing the search query.
- Enable horizontal navigation between the scope sidebar and model list using left and right arrow keys.
- Update UI help strings to reflect the new spatial navigation controls.
- Replaced the legacy model selector with a full-screen Model Hub, introducing mouse support and a fuzzy-searchable browser.
- Integrated comprehensive model management, including role assignment, thinking-level visualization, and manual provider discovery.
- Implemented a cancellable OAuth login flow and integrated it directly into the Model Hub for provider authentication.
- Centralized model logic and migrated existing tests to support the new component architecture.
- Refined dialog layout with stable height clamping and improved preview visibility logic.
- Simplified interaction flow by replacing the "Next" row with "Submit" tab confirmation.
- Enabled accessible option toggling using Enter and Space keys.
- Removed legacy chat integration and streamlined internal dialog state management.
- Split mixed assistant text into per-tool segments instead of one post-tool tail.
- Insert each segment immediately after its preceding tool component across live and rebuilt transcripts.
- Extended the regression test to cover two tool calls with middle and final assistant text.
Fixes#4871
- Closed RPC host tool bridge before draining queued commands after stdin EOF.
- Rejected active host tool calls and future queued host tool calls with the disconnect error instead of emitting new host_tool_call frames.
- Added regression coverage for dispatcher drain with active and queued host tool requests.
Fixes#5153
- Split mixed assistant messages so pre-tool text stays before tool panels while trailing text renders after the tool timeline.
- Applied the split to live streaming, transcript rebuilds, and file-backed transcript rendering.
- Added a focused EventController regression test for text/toolCall/text Cursor-shaped turns.
Fixes#4871
- Added a serialized RPC input dispatcher so control-plane frames can resolve dialogs while ordinary commands remain ordered.
- Made pending extension UI requests fail closed on RPC disconnect so EOF drains active and queued commands.
- Covered extension UI response overtaking, queue ordering, queue recovery, and EOF dialog rejection in rpc input tests.
Fixes#5153
- Transitioned vibe tools to an ephemeral registration model where they are installed only when entering `/vibe` mode and removed upon exit.
- Added `activateVibeTools` and `deactivateVibeTools` methods to `AgentSession` to manage these transient tool registrations.
- Removed vibe tools from the default global tool registry, preventing unnecessary background exposure.
- Improved real-time screen rendering by implementing dynamic builder functions for vibe mode components.
- Optimized cursor and spinner rendering to re-derive state from shared mutable options on every paint.
- Added session snapshotting and improved wait logic to accurately detect settled turns and running sessions.
- Ensured reliable toolset restoration during mode transitions through updated logic and new contract tests.
- Implemented Vibe mode to enable worker session management and director-role context injection.
- Added a `/vibe` slash command and integrated status line UI to display mode activity.
- Configured restricted toolsets and guards to prevent concurrent conflicts with existing Goal or Plan modes.
- Provided system prompts and tool templates to support specialized agent communication and task orchestration.
- Added a fallback to emit error messages during `agent_end` if no error was previously streamed during the turn.
- Added tracking to prevent duplicate error messages when a provider error is successfully delivered during streaming.
- Added a stderr hint for interactive users launching the ACP server directly from a terminal.
- Treated missing or invalid changelog markers as first install and persisted the current version without replaying historical notes.
- Shared bounded changelog rendering between startup and recent changelog views, with a 64 KiB startup cap and full-history hint on truncation.
- Added marker, truncation, recent/full rendering, and PTY startup regression coverage.
Fixes#5135
Rendered the /move directory picker using the overlay layout width instead of the legacy fixed 68-column frame.
Added regression coverage for non-68-column overlay frames.
Fixes#5067
- Transitioned DeepSeek models to an explicit `[High, Max]` effort ladder, removing stale alias maps.
- Updated the OpenAI compatibility layer to enforce authoritative `supportsReasoningEffort` and `omitReasoningEffort` flags.
- Synchronized `models.json` definitions to reflect accurate reasoning effort capabilities across the catalog.