- Integrated a horizontal separator in the model browser list to distinguish between frequently used or role-assigned models and other available models.
- Removed the "Recent" category from the Model Hub sidebar to streamline navigation.
- Updated related tests to reflect the removal of the Recent category and the new visual list layout.
Validated custom tool factory results before registering plugin-provided tools so one malformed feature entry is reported and skipped instead of crashing startup.
Added regression coverage for null entries and mixed valid/null factory arrays.
Fixes#5189
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
- Implemented tiered search strategy using synchronous literal matching for immediate results and debounced asynchronous fuzzy matching to prevent input blocking.
- Cached session search data and introduced an indexed `FuzzyText` structure to reduce redundant string processing.
- Added comprehensive test suite verifying search convergence, stale task orphaning, and selection stability.
- Enhanced `tui` fuzzy matching utilities to support prebuilt search indexes for improved multi-match performance.
Codex review round 8 on 679b45e5f:
- matchesReplacementCredential: an org-scoped incoming credential now
matches an existing row keyed by ANY of its base identities
(email/account/project), bare or org-qualified with the same org, plus
the bare org-only key — so an account-keyed row stored while email
recovery failed is claimed once a later same-subscription login
recovers the email, instead of duplicating. Org-less incoming
credentials keep exact-key matching only.
- auth-broker migrate: index org-only broker snapshot rows by their org
id so reruns skip an already-migrated row instead of re-uploading a
stale refresh token over the broker's newer one.
Codex review round 7 on e90a72bdd flagged that broker usage-report
matching, header-overlay keying, and omp-usage coverage all treated a
matching organization as a sufficient match. Two Team members share the
org id while drawing on per-user pools, so the first same-org report
(or a lone sibling report) was handed to the wrong member. The org is
now a gate: within the same-org subset the member's own base identity
(account/email/project) must still match, with org-only entities (no
base identifiers) matching on the org alone when unambiguous. The
overlay merger (findMatchingReportIndex) had the identical same-org
flaw and receives the symmetric fix. Org-presence-mismatch semantics
are unchanged: org-scoped vs org-less stays fall-through/unreported,
and both-org-less keeps the legacy base-identity fallback.
Addresses the Codex review round on c3fcb4fa7:
- Active-account matching (/logout preselection, /usage 'in use by this
session' marker) treated a shared org as sufficient: two Anthropic Team
seats in one org (same orgId, per-user pools, distinct email/account)
matched each other's rows and reports. The org is now a gate that
qualifies the base identity — mismatched org presence or different orgs
still never match (round-3 semantics unchanged), a shared org falls
through to the account/email/project checks, and only an org-only
active identity (no base identifiers recovered) matches on the org
alone.
- An org-only credential row keyed org:<id> (stored when login recovered
neither email nor account) was never claimed by a later same-org login
that does recover the identity, duplicating one subscription across two
rows. matchesReplacementCredential now upgrades and re-keys such a row
in place; the upgrade stays one-way — an org-only incoming key still
only claims rows via exact key equality.
Regression tests: same-org different-member rows/reports stay unmatched
while the member's own row matches; org-only identities match same-org
rows on the org alone; org-only row upgraded in place on identity
recovery with the one-way direction preserved.
Addresses the fourth review round (Codex no-email finding on d37e3992c,
confirmed and scoped by internal review):
- resolveProviderCredentialIdentityKey: the anthropic org qualifier now
rides on whichever base identity exists (email > account > project),
not only email. The account UUID is identical across the orgs of one
login account, so the bare account fallback let a second subscription
replace the first whenever the email could not be recovered (token
response omits it AND bootstrap fails). Org-only credentials key on
the org alone instead of losing identity entirely.
- matchesReplacementCredential: the one-way legacy claim strips a
trailing |org: from ANY anthropic base key (account/project included);
only anthropic keys carry the qualifier, so other providers are
unaffected.
- Usage-report dedupe falls back to the org-qualified account for
no-email anthropic reports instead of returning no identifiers.
- Broker report/overlay routing (matchUsageReport/findMatchingReportIndex)
is org-decisive on EITHER side: an org-less legacy credential no longer
receives an org-attributed sibling's pool via the lone-candidate or
email/account fallback, and an org-less overlay only merges into
org-less reports.
- omp usage unreported-account attribution follows the same either-side
rule, so a legacy row whose fetch failed surfaces as 'no usage data'
instead of being hidden by a sibling's report.
- Regression tests: no-email identity coexistence/replace/claim, no-email
report dedupe, org-less broker routing, either-side unreported
attribution.
Addresses the second review round (internal re-review + Codex on c38840482):
- Active-account matching (logout preselection, /usage in-use marker) is
org-decisive when EITHER side carries an org: a legacy bare-email active
row no longer flags org-scoped siblings via the shared email (reverse of
the previous fix). Both-org-less keeps the email/account fallback, so
providers without orgs are unaffected.
- Status-line usage context key includes orgId, so rotating between two
same-email subscriptions invalidates the cached quota immediately
instead of showing the previous org's numbers for the cache TTL.
- CredentialHealthResult carries orgId/orgName and auth-gateway check
labels rows with the org, so a failing row names the subscription.
- getOAuthAccountIdentity preserves org-only identities; the login
success message renders them.
- ACP /usage account-id fallback labels get the org suffix too.
- Regression tests for both matching directions (marker + logout).
Addresses Codex review on #5170:
- OAuthAccess/OAuthAccessFailure and every resolution site now carry
orgId/orgName; dry-balance bench keys and labels are org-qualified so
two same-email subscriptions stay two benchmark targets.
- When the active identity is org-scoped, logout active-marking and the
/usage in-use marker match ONLY the same org — an org-less legacy row
or pre-upgrade report can no longer be flagged active via the shared
email, so the logout preselection cannot land on the wrong row.
- Regression tests: org-scoped active vs sibling org and vs legacy
bare-email row (logout + in-use marker), org-suffixed logout labels.
One Anthropic account email can hold multiple organizations (a Team seat
plus a personal Max plan), each with its own org-scoped OAuth token and
independent 5h/7d limit pools. Credentials were deduped by bare email, so
logging in with the second subscription silently replaced the first, and
usage reports from the two pools merged into one row with mixed numbers.
- capture organization uuid/name at login (token exchange response, with
a claude_cli/bootstrap fallback); token refreshes never rewrite it
- key anthropic credential identity as email + org; a legacy email-keyed
row is claimed in place by the first org-scoped login with the same
email, and org-less credentials never clobber org-scoped rows
- partition usage-report dedupe and the per-credential usage cache by
org so the two subscriptions' limit pools stay distinct for rotation
- show the organization in omp usage (redaction-safe) and name the
stored account/org in the login success message
- Implemented custom role creation within the Model Hub, including a virtual row for direct initiation and name stripping.
- Added quick-switch cycle editing functionality with persistent ordering and live preview of role membership.
- Optimized sidebar scope navigation to mute empty entries and improve keyboard focus stability during search.
- Updated the Model Hub sidebar UI to prioritize Roles and included comprehensive tests for new navigation and management flows.
- Implemented dynamic reordering of model providers in the sidebar to prioritize matches during active searches.
- Refactored model hub registry synchronization to maintain separate state for fixed, unlocked, and locked provider entries.
- Added a dedicated composition step to reassemble sidebar entries based on current search result counts.
- Added test coverage to verify that provider order reverts to alphabetical upon clearing the search query.
Included advisor tool-result text in the quarantine source check so legitimate findings from granted read/grep tools are not treated as model-generated contamination.
Kept assistant text out of the source set to avoid laundering prior advisor hallucinations.
Fixes#5181
- Enable horizontal navigation between the scope sidebar and model list using left and right arrow keys.
- Update UI help strings to reflect the new spatial navigation controls.
- Replaced the legacy model selector with a full-screen Model Hub, introducing mouse support and a fuzzy-searchable browser.
- Integrated comprehensive model management, including role assignment, thinking-level visualization, and manual provider discovery.
- Implemented a cancellable OAuth login flow and integrated it directly into the Model Hub for provider authentication.
- Centralized model logic and migrated existing tests to support the new component architecture.
Scanned allowed advise tool notes for output-only destructive directives before the tool can route them to the primary agent.
Kept provenance checks against the watched session update so legitimate warnings about user-provided dangerous text still pass.
Fixes#5181
Prevented late interrupting advisor findings from waking the primary after a terminal text answer when no queued work remains.
Added regression coverage for the advisor-confirmation path so duplicate primary turns are caught.
Fixes#4840
Cleared provider-native replay payloads and stop details when Advisor output is quarantined so persisted transcripts only contain the sanitized error.
Added regression coverage for OpenAI Responses-style providerPayload leakage.
Fixes#5181
Quarantined Advisor assistant turns that request tools outside the granted tool pool before they can enter the Advisor context.
Reset the Advisor runtime after quarantine so the next update re-primes from the primary transcript instead of replaying contaminated private context.
Fixes#5181
- Refined dialog layout with stable height clamping and improved preview visibility logic.
- Simplified interaction flow by replacing the "Next" row with "Submit" tab confirmation.
- Enabled accessible option toggling using Enter and Space keys.
- Removed legacy chat integration and streamlined internal dialog state management.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
Lazy-initialized the header-generator dependency so compiled runtimes without its fs-loaded data_files fall back to the bundled Chrome header profile instead of failing extension imports.
Added a regression test that hides header-generator data_files in a fresh Bun subprocess and verifies fallback headers are returned.
Fixes#5178
- Clarified prompt instructions to emphasize using specific agent roles over the default worker.
- Updated the task prompt template to provide clearer guidance on agent selection policies.
- Removed unused template logic related to the default agent identification.
- Configured the default `task` subagent to use `auto` thinking.
- Enabled `auto` as a valid thinking-level value in agent frontmatter.
- Adjusted thinking-level precedence to ensure that explicit `:level` suffixes in resolved model patterns override agent-defined defaults.
- Recognized preformatted chat context in message-preproc and bypassed paired-tag stripping that consumed the entire envelope.
- Integrated scaffolding-tag removal in low-signal title prefilter.
- Added corresponding tests for formatTitleUserMessage and isLowSignalTitleInput.
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
- Made GenerateImage omit chatgpt-account-id when Codex bearer tokens do not expose an account id.
- Added Codex hosted image coverage for opaque proxy keys and JWT-derived account headers.
Fixes#5174
- Bun.build-API compiled Windows executables report import.meta.main === false
(standalone loader keys the entry module with backslashes but registers the
main path with forward slashes), so cli.ts never dispatched: the binary
exited silently with code 0 and omp update rolled back after failing to
verify the new version.
- Entry dispatch and worker-host declaration now also honor the define-folded
PI_COMPILED marker, which is only true inside compiled binaries where the
entry module is by definition the process entry.
- Verified on a Windows VM: --version prints omp/16.4.3 and --smoke-test
passes (stats sync worker + tiny-model subprocess) on a cross-compiled
binary; v16.4.3 release binary reproduces the silent exit.
- Sanitized task subagent progress, fallback output, retry/error text, and yield previews before rendering in the parent TUI.
- Added regression coverage for carriage-return and CSI bytes in expanded subagent output.
Fixes#5159
- Split mixed assistant text into per-tool segments instead of one post-tool tail.
- Insert each segment immediately after its preceding tool component across live and rebuilt transcripts.
- Extended the regression test to cover two tool calls with middle and final assistant text.
Fixes#4871
Added MiMo to the hashline edit-mode fallback list so it uses replace mode by default unless explicitly configured otherwise.
Kept DeepSeek Flash on hashline after maintainer could not reproduce provider-agnostic failure locally.
Added regression coverage for MiMo fallback and DeepSeek Flash staying on hashline.
Fixes#3772
- Closed RPC host tool bridge before draining queued commands after stdin EOF.
- Rejected active host tool calls and future queued host tool calls with the disconnect error instead of emitting new host_tool_call frames.
- Added regression coverage for dispatcher drain with active and queued host tool requests.
Fixes#5153
- Split mixed assistant messages so pre-tool text stays before tool panels while trailing text renders after the tool timeline.
- Applied the split to live streaming, transcript rebuilds, and file-backed transcript rendering.
- Added a focused EventController regression test for text/toolCall/text Cursor-shaped turns.
Fixes#4871
- Added a serialized RPC input dispatcher so control-plane frames can resolve dialogs while ordinary commands remain ordered.
- Made pending extension UI requests fail closed on RPC disconnect so EOF drains active and queued commands.
- Covered extension UI response overtaking, queue ordering, queue recovery, and EOF dialog rejection in rpc input tests.
Fixes#5153