Commit Graph

418 Commits

Author SHA1 Message Date
can1357 9f17927ee7 Merge PR #5467: fix(cli): flush config JSON output before exit (@roboomp) 2026-07-14 22:58:46 +02:00
roboomp df7aa6d01f fix(cli): awaited config JSON stdout flush
- Waited for the stdout write callback before the config command exits.
- Covered JSON output larger than the 64 KiB pipe buffer.

Fixes #5309
2026-07-14 18:06:40 +00:00
can1357 946ba4bcea Merge PR #5044: fix(cli): parse max-time duration suffixes (@roboomp)
# Conflicts:
#	packages/coding-agent/src/cli/flag-tables.ts
2026-07-14 18:46:14 +02:00
can1357 e31a508140 Merge PR #5057: fix(cli): route npm self-updates through npm (@roboomp) 2026-07-14 18:41:16 +02:00
can1357 7f9a2777f9 Merge PR #5170: fix(ai): scope Anthropic credential identity by organization, not just email (@chan1103) 2026-07-14 18:29:14 +02:00
can1357 4c1c5f40d8 feat(launch): rendered launch logs from daemon terminal byte streams
- Changed daemon log reads to return both sanitized display text and a raw `terminalText` slice, and included it on log RPC responses for PTY runs when grep was not used.
- Extended the logs result contract and launch tool rendering to consume `terminalText`, reconstruct terminal output, and display it in framed, preview-capped sections.
- Kept terminal row layout stable by writing space characters for empty cells when reading rows, preserving spacing during output reconstruction.
2026-07-13 19:21:45 +02:00
can1357 4df6f6683d chore: finalizing the new /prewalk 2026-07-13 15:18:31 +02:00
can1357 f405525bf4 feat: removed boomerang feature and associated validation workflows
- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
2026-07-13 14:15:54 +02:00
can1357 95ecc61bc9 feat(coding-agent): implemented downshift boomerang flow for context handoff
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
2026-07-13 10:37:20 +02:00
can1357 9f1ff90a39 feat(coding-agent): implemented downshift and plan-yolo agent workflows
- Replaced legacy reasoning-slide functionality with new downshift and plan-yolo capabilities.
- Updated CLI arguments, slash commands, and configuration schemas to support the new model-switching and execution behaviors.
- Refactored agent session logic to handle downshift arming, plan-yolo headless execution, and context scrubbing.
- Renamed and added system prompts to align with the updated downshift and plan-yolo workflows.
2026-07-13 06:03:46 +02:00
can1357 e770cdc4d9 fix(coding-agent): unified tool renderer and clarified launch diagnostics
- Implemented a consolidated TUI renderer for launch tool operations to unify status headers and metadata.
- Added tracking for unfulfilled readiness conditions to distinguish between timeout causes and daemon states.
- Enhanced timeout reporting to explicitly surface specific unmet log or port conditions.
- Included comprehensive unit and regression tests for tool rendering and start-timeout diagnostics.
2026-07-13 05:16:36 +02:00
can1357 ac16253613 wip: rslide (experimental) 2026-07-13 04:21:54 +02:00
can1357 b6559861d0 feat(cli): standardized throughput calculation to total duration
- Removed the `computeTokensPerSecond` helper function in favor of a direct calculation.
- Standardized throughput to use total request duration instead of post-TTFT decode time.
- Updated UI usage display to calculate tokens per second against total duration.
- Prevented over-inflation of throughput metrics caused by hidden reasoning tokens.
2026-07-12 01:06:24 +02:00
can1357 6bb0878b63 feat(ai): added cache invalidation command for usage reports
- Added an `invalidate` action to the usage CLI to clear cached usage reports for specific or all providers.
- Updated `invalidateUsageCache` in `AuthStorage` to be asynchronous and notify the underlying store of invalidations.
2026-07-11 20:55:19 +02:00
chan1103 1e8f183f49 fix(ai): claim same-subscription rows across base identities; skip migrated org-only rows
Codex review round 8 on 679b45e5f:

- matchesReplacementCredential: an org-scoped incoming credential now
  matches an existing row keyed by ANY of its base identities
  (email/account/project), bare or org-qualified with the same org, plus
  the bare org-only key — so an account-keyed row stored while email
  recovery failed is claimed once a later same-subscription login
  recovers the email, instead of duplicating. Org-less incoming
  credentials keep exact-key matching only.
- auth-broker migrate: index org-only broker snapshot rows by their org
  id so reruns skip an already-migrated row instead of re-uploading a
  stale refresh token over the broker's newer one.
2026-07-11 22:49:12 +09:00
chan1103 e095af3be0 fix(ai): require member identity within a shared org for report routing, overlays, and coverage
Codex review round 7 on e90a72bdd flagged that broker usage-report
matching, header-overlay keying, and omp-usage coverage all treated a
matching organization as a sufficient match. Two Team members share the
org id while drawing on per-user pools, so the first same-org report
(or a lone sibling report) was handed to the wrong member. The org is
now a gate: within the same-org subset the member's own base identity
(account/email/project) must still match, with org-only entities (no
base identifiers) matching on the org alone when unambiguous. The
overlay merger (findMatchingReportIndex) had the identical same-org
flaw and receives the symmetric fix. Org-presence-mismatch semantics
are unchanged: org-scoped vs org-less stays fall-through/unreported,
and both-org-less keeps the legacy base-identity fallback.
2026-07-11 22:49:12 +09:00
chan1103 c2456d882f fix(ai): org-qualify account/project fallback identities; org-decisive routing on either side
Addresses the fourth review round (Codex no-email finding on d37e3992c,
confirmed and scoped by internal review):

- resolveProviderCredentialIdentityKey: the anthropic org qualifier now
  rides on whichever base identity exists (email > account > project),
  not only email. The account UUID is identical across the orgs of one
  login account, so the bare account fallback let a second subscription
  replace the first whenever the email could not be recovered (token
  response omits it AND bootstrap fails). Org-only credentials key on
  the org alone instead of losing identity entirely.
- matchesReplacementCredential: the one-way legacy claim strips a
  trailing |org: from ANY anthropic base key (account/project included);
  only anthropic keys carry the qualifier, so other providers are
  unaffected.
- Usage-report dedupe falls back to the org-qualified account for
  no-email anthropic reports instead of returning no identifiers.
- Broker report/overlay routing (matchUsageReport/findMatchingReportIndex)
  is org-decisive on EITHER side: an org-less legacy credential no longer
  receives an org-attributed sibling's pool via the lone-candidate or
  email/account fallback, and an org-less overlay only merges into
  org-less reports.
- omp usage unreported-account attribution follows the same either-side
  rule, so a legacy row whose fetch failed surfaces as 'no usage data'
  instead of being hidden by a sibling's report.
- Regression tests: no-email identity coexistence/replace/claim, no-email
  report dedupe, org-less broker routing, either-side unreported
  attribution.
2026-07-11 22:49:12 +09:00
chan1103 a47c3c90ec fix(ai): org-decisive active matching on either side; org in status-line cache key and health results
Addresses the second review round (internal re-review + Codex on c38840482):

- Active-account matching (logout preselection, /usage in-use marker) is
  org-decisive when EITHER side carries an org: a legacy bare-email active
  row no longer flags org-scoped siblings via the shared email (reverse of
  the previous fix). Both-org-less keeps the email/account fallback, so
  providers without orgs are unaffected.
- Status-line usage context key includes orgId, so rotating between two
  same-email subscriptions invalidates the cached quota immediately
  instead of showing the previous org's numbers for the cache TTL.
- CredentialHealthResult carries orgId/orgName and auth-gateway check
  labels rows with the org, so a failing row names the subscription.
- getOAuthAccountIdentity preserves org-only identities; the login
  success message renders them.
- ACP /usage account-id fallback labels get the org suffix too.
- Regression tests for both matching directions (marker + logout).
2026-07-11 22:49:12 +09:00
chan1103 bee01bfc4a fix(ai): carry org through OAuth access results; org-scoped active never matches org-less rows
Addresses Codex review on #5170:

- OAuthAccess/OAuthAccessFailure and every resolution site now carry
  orgId/orgName; dry-balance bench keys and labels are org-qualified so
  two same-email subscriptions stay two benchmark targets.
- When the active identity is org-scoped, logout active-marking and the
  /usage in-use marker match ONLY the same org — an org-less legacy row
  or pre-upgrade report can no longer be flagged active via the shared
  email, so the logout preselection cannot land on the wrong row.
- Regression tests: org-scoped active vs sibling org and vs legacy
  bare-email row (logout + in-use marker), org-suffixed logout labels.
2026-07-11 22:49:12 +09:00
chan1103 044d722a36 fix(ai): scope Anthropic credential identity by organization
One Anthropic account email can hold multiple organizations (a Team seat
plus a personal Max plan), each with its own org-scoped OAuth token and
independent 5h/7d limit pools. Credentials were deduped by bare email, so
logging in with the second subscription silently replaced the first, and
usage reports from the two pools merged into one row with mixed numbers.

- capture organization uuid/name at login (token exchange response, with
  a claude_cli/bootstrap fallback); token refreshes never rewrite it
- key anthropic credential identity as email + org; a legacy email-keyed
  row is claimed in place by the first org-scoped login with the same
  email, and org-less credentials never clobber org-scoped rows
- partition usage-report dedupe and the per-credential usage cache by
  org so the two subscriptions' limit pools stay distinct for rotation
- show the organization in omp usage (redaction-safe) and name the
  stored account/org in the login success message
2026-07-11 22:49:12 +09:00
roboomp 3876f60d65 fix(cli): routed npm self updates through npm
- Detected npm-owned global bin paths and Windows npm launcher shims before falling back to binary replacement.
- Added npm install argv construction that pins the registry and native package versions.
- Covered npm shim resolution and npm update arguments in update-cli tests.

Fixes #5053
2026-07-10 08:53:20 +00:00
roboomp d670dd5d94 fix(cli): reported max-time parse errors cleanly
- Added a CLI usage-error type for argument validation failures.
- Reported invalid --max-time values from launch and ACP as clean usage errors with exit code 2.
- Covered the no-stack-trace CLI error path for invalid max-time input.

Fixes #5041
2026-07-10 07:45:27 +00:00
roboomp 7fa2c3f42d fix(coding-agent): preserved fork prompt cache affinity
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.

- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.

- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.

Fixes #5035
2026-07-10 07:22:28 +00:00
roboomp 6fab752eb4 fix(cli): parsed max-time duration suffixes
- Parsed --max-time values with s, m, and h suffixes into seconds instead of dropping the deadline.
- Raised a visible parse error for invalid or non-positive max-time values.
- Updated help text and regression coverage for the CLI timeout parser.

Fixes #5041
2026-07-10 07:18:53 +00:00
can1357 3dec7ac12c Merge remote-tracking branch 'origin/farm/99fd01f3/fix-linux-tiny-cuda-runtime' 2026-07-04 05:15:18 +02:00
roboomp f257fd533f fix(tiny): repaired cuda side-runtime install
Downloaded missing ONNX Runtime CUDA provider sidecars when the compiled tiny-model side runtime is used with PI_TINY_DEVICE=cuda.

Preserved actionable CUDA worker diagnostics in tiny-models text output and added focused regression coverage.

Fixes #4475
2026-07-03 18:21:19 +00:00
roboomp 721f6d4a08 fix(mcp): make full URL the primary OAuth copy target so SSH sessions work
Codex review flagged that advertising `launchUrl`
(http://localhost:<omp-port>/launch) as the visible `Copy URL:` breaks
SSH/WSL/headless users: their local browser resolves the URL against
the local machine (no OMP listening) and fails before ever hitting the
provider. On terminals without OSC 8 support, they lose the manual
`/login <redirect>` path entirely.

Every OAuth-facing surface now shows the full authorization URL as the
primary copy target and offers `launchUrl` as an additional "Local
shortcut (this machine only)" line for wide-terminal local users who
want the truncation-safe convenience:

- MCPAuthorizationLinkPrompt renders `Copy URL:` with the full URL and
  appends the local-shortcut row only when `launchUrl` differs. OSC 52
  clipboard staging in the MCP onAuth handler switches to the full URL
  (OSC 52 is a wire-level protocol — the terminal writes to the
  caller's LOCAL clipboard even when OMP is on a remote SSH box).
- LoginDialogComponent.showAuth, selector-controller onAuth,
  setup-wizard sign-in, and the auth-broker CLI mirror the pattern:
  full URL first, launchUrl as an optional local shortcut.
- Setup wizard uses `wrapTextWithAnsi`, not truncation, so the RFC
  7636 §4.3 downgrade bug that motivated launchUrl is unreachable
  through it; still surfaces launchUrl for wide-terminal convenience.

Regression tests in
`packages/coding-agent/test/modes/controllers/mcp-authorization-link.test.ts`
now assert:
- Full URL is the primary `Copy URL:` line so SSH sessions can complete.
- launchUrl still appears beneath as `Local shortcut (this machine only): …`
  when it differs from the full URL.
- No shortcut row when launchUrl is absent OR equals the full URL.
2026-07-03 08:57:55 +00:00
roboomp 97c1d08cce fix(mcp): surface a short launch URL and log Windows opener failures for OAuth
Two independent defects broke /mcp reauth against S256-only providers on
Windows boxes whose PATH no longer references System32:

1. openPath spawned bare rundll32 and swallowed the
   `Executable not found in $PATH` throw with a bare `catch {}`, so the MCP
   controller's outer try/catch was dead and the transcript unconditionally
   claimed "Opening browser automatically...".
2. TUI#prepareLine silently truncates any composed row wider than the
   viewport. MCPAuthorizationLinkPrompt rendered `Copy URL: <full URL>` as a
   single ~271-column line whose trailing parameter is
   code_challenge_method=S256. On the reporter's 270-col terminal the cut
   landed inside that parameter, dropping the method while keeping
   code_challenge — which RFC 7636 §4.3 treats as plain PKCE, which Linear
   correctly rejects with "The plain PKCE method is not allowed. Use S256
   instead."

OAuthCallbackFlow now hosts a `GET /launch` route on the same loopback
callback server it already runs; the route 302-redirects to the pending
authorization URL and is advertised as `OAuthAuthInfo.launchUrl` — a
~30-char copy target no viewport can meaningfully truncate. The MCP OAuth
fallback, /login, setup wizard, auth-broker CLI, and login-dialog all
prefer the launch URL for the visible copy target, keep the full URL in
the OSC 8 hyperlink for click-through, and the MCP flow additionally
stages the copy target on the clipboard via OSC 52 (same pattern the
setup wizard uses).

openPath now resolves rundll32.exe through %SystemRoot%\System32 (with a
C:\Windows fallback when SystemRoot is unset) and logs both synchronous
spawn throws and non-zero exits via the shared logger, so silent
misconfigurations show up in ~/.omp/logs/omp.*.log. The dead try/catch
around openPath in the MCP controller is removed.

Fixes #4418
2026-07-03 08:19:14 +00:00
can1357 8c648e4d02 fix(cli): enabled consistent reporting for sibling provider limits
- Added collectProviderLimitTemplates to aggregate limit IDs across all reports.
- Updated formatUsageBreakdown to render placeholders for limits missing from specific providers.
- Standardized label widths across sibling providers to align status bars.
2026-07-03 04:53:28 +02:00
can1357 8b3d0a7190 Merge remote-tracking branch 'origin/farm/1f41837c/timeout-bare-fetches' 2026-07-02 23:43:11 +02:00
can1357 41f06129b3 Merge remote-tracking branch 'origin/farm/2b21222d/fix-update-plugin-list-flag' 2026-07-02 23:43:09 +02:00
roboomp a86bf4f41b fix(coding-agent): surfaced invalid models config errors
Converted ArkType validation failures into ConfigError results and surfaced models.yml validation failures in CLI model listing and noninteractive startup.

Fixes #4305
2026-07-02 11:58:56 +00:00
roboomp 7cb413988f fix(cli): routed update plugin shorthand
Added the documented update -l path for marketplace plugin upgrades and covered the dispatch contract with update command tests.

Fixes #4304
2026-07-02 11:48:42 +00:00
roboomp 76c480646e fix(cli): added fetch timeouts
Added timeout-backed AbortSignals to update, Hindsight, and Smithery fetch calls so stalled endpoints abort instead of hanging indefinitely.

Added regression coverage for the timeout signals on the exposed command/client paths.

Fixes #4229
2026-07-02 08:39:21 +00:00
can1357 95b91c7f73 feat(coding-agent/tools)!: replaced paths arrays with path strings
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
2026-07-02 08:30:33 +02:00
can1357 51684b4b1d refactor(coding-agent): streamlined codebase by deduplicating helper logic and shims
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
2026-07-02 02:40:08 +02:00
roboomp a4ae4c130c fix(coding-agent): preserved explicit :auto suffix in modelRoles
The model selector's persistence path dropped the `:auto` selector when parsing role values, producing a warning ('Invalid thinking level "auto"') and rendering the badge as `inherit` instead of `auto`. Reload of the default role also lost the auto state whenever the role value carried an explicit `:auto` suffix instead of relying on `defaultThinkingLevel`.

Widen the resolver chain (`parseThinkingSuffix`, `splitThinkingSuffix`, `parseModelString`, `parseModelPattern*`, `ResolvedModelRoleValue`, `ResolvedRoleModel`, `ResolveCliModelResult`) to carry the `AUTO_THINKING` sentinel end to end, and coerce it back to `undefined` at concrete-only boundaries (glob scope patterns, retry fallback, advisor, commit pipeline, guided-goal, bench).

Regression tests cover:

- `resolveModelRoleValue("provider/model:auto")` returns explicit auto without a warning.

- `ModelSelector` renders `DEFAULT (auto)` and `SMOL (auto)` when the role value has `:auto`.

- `cycleRoleModels` activates auto thinking on entering a `:auto` role.

- Startup resume activates auto thinking when `modelRoles.default` carries `:auto`.

Fixes #4128
2026-07-01 08:18:42 +00:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 31d9fb1145 refactor: consolidated filesystem traversal into a dedicated library
- Replaced custom fast-walk and fs-cache implementations in pi-natives with a new dedicated pi-walker library.
- Integrated the thread-safe, parallel pi-walker library across pi-natives, pi-shell, pi-uu-grep, and uu-find.
- Rewrote file search, fuzzy finding, and glob-matching logic to leverage pi-walker configurations and visitor traits.
- Optimized shell process tracking in pi-shell by replacing global descendant-diff logic with an isolated, per-run SpawnRegistry.
- Added parallel rayon-based walking and optimized fast paths for directory scanning and entry classification.
2026-06-30 23:25:50 +02:00
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
roboomp f5be8953c4 fix(cli): surfaced tiny model download errors
Preserved worker-side download errors through TinyTitleClient and included them in tiny-models text and JSON failures.

Fixes #3839
2026-06-29 22:56:40 +00:00
can1357 64be11e289 Merge PR #3806: fix(cli): preserve Windows extension paths (@roboomp) 2026-06-29 16:45:47 +02:00
can1357 17c4aa0391 Merge PR #3795: fix(cli): honor web search provider settings in omp search (@roboomp) 2026-06-29 16:45:45 +02:00
can1357 1eb32a5df1 feat(coding-agent/bench-cli): tracked and close provider session states
- Added management of provider session states during benchmark execution.
- Implemented a teardown process to close and clear session states after request completion.
2026-06-29 16:31:21 +02:00
roboomp 777b609e63 fix(cli): preserved windows extension paths
Rejoined split Windows extension module paths before launch parsing finishes and stripped extended-length Win32 prefixes before Bun import and worker spawn APIs see them.

Fixes #3804
2026-06-29 11:50:36 +00:00
roboomp 3560d108db fix(cli): applied search provider settings
Initialized standalone search commands with configured web-search provider globals before resolving the implicit provider chain.

Fixes #3793
2026-06-29 08:55:53 +00:00
can1357 312b71b0a5 feat(sdk): enabled websocket transport configuration for sdk and cli
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
2026-06-29 08:05:02 +02:00
can1357 c8c24b0888 feat(bench): enabled concurrent execution and service tier selection
- Added --par flag to execute benchmark runs concurrently with a default degree of 4.
- Added --service-tier flag to allow overriding the provider service tier per benchmark.
- Increased default benchmark run count from 1 to 10 to provide more robust averaging.
- Updated benchmarking logic to process requests in a concurrency-limited pool while preserving output order.
- Implemented pre-flight credential checks to prevent unnecessary worker spawning when authentication is missing.
2026-06-29 06:51:13 +02:00
roboomp 58c0a305d6 fix(agent): shortened isolated task paths
Used compact hashed isolation directory segments and the short m mount dir so long task ids are not copied into subagent working paths.

Kept worktree cleanup compatible with legacy merged task-isolation directories.

Fixes #3756
2026-06-28 21:25:20 +00:00
can1357 21ed419aeb fix(coding-agent): keep registry-qualified bun cache markers 2026-06-28 18:45:12 +02:00