Commit Graph

445 Commits

Author SHA1 Message Date
usr-bin-roygbiv b9504f65e7 feat: add native Codex computer use 2026-07-24 01:40:04 +00:00
can1357 5ded27b0b1 fix(cli): stopped $-pattern expansion when injecting cache prefix 2026-07-24 02:25:25 +02:00
can1357 3f2ef2cea1 Merge PR #6413: feat(cli): benchmark prompt cache reuse (@riverpilot)
# Conflicts:
#	packages/ai/src/providers/pi-native-server.ts
#	packages/ai/src/stream.ts
#	packages/ai/src/types.ts
2026-07-24 02:25:24 +02:00
Alex TYRODE 60920e1c00 Merge upstream main into feat/auth-broker-account-pools 2026-07-23 21:37:51 +00:00
Alex TYRODE e6ab870d37 fix(ai): close auth broker account pool gaps 2026-07-23 21:28:17 +00:00
Alexander Kirilin 4274100bbe fix(cli): preserve cache prefix whitespace 2026-07-23 16:52:20 -04:00
Alexander Kirilin f25cc6cf09 fix(cli): harden prompt cache benchmark diagnostics 2026-07-23 16:37:39 -04:00
roboomp 5f47afba16 fix(session): validate blob refs before path join
parseBlobRef sliced the blob:sha256: suffix and returned it unvalidated;
get/getSync then fed it into path.join(this.dir, hash), so a crafted ref
like blob:sha256:../../../etc/passwd escaped the blob directory and read
arbitrary files into resolved image history (base64, raw UTF-8, and the ACP
sync path).

Reject any suffix that is not a canonical 64-char lowercase hex hash in
parseBlobRef, the single choke point for every resolution entry point. Reuse
the shared BLOB_HASH_RE in gc-cli instead of its duplicate HASH_RE.

Fixes #4088
2026-07-23 19:51:53 +00:00
Alexander Kirilin 88296f9c55 fix(cli): preserve cache benchmark affinity diagnostics 2026-07-23 15:40:07 -04:00
Alexander Kirilin 9c471a5088 feat(cli): benchmark prompt cache reuse 2026-07-23 15:18:37 -04:00
can1357 26726fdcb9 Merge PR #6255: add dynamic multi-root workspace context (@maatheusgois-dd) 2026-07-23 17:30:33 +02:00
can1357 da5da41169 Merge PR #6382: feat(ai): report Anthropic extra usage in omp usage (@LunarECL) 2026-07-23 16:35:47 +02:00
LunarECL e7fdc99afd feat(ai): report Anthropic extra usage 2026-07-23 21:25:15 +09:00
can1357 0f54c0df70 fix(tools): autoqa consent handling from default off to opt-in 2026-07-23 13:24:45 +02:00
Gareth-Rouse e2839d1288 feat(ai): added Synthetic usage provider
/usage and omp usage now report Synthetic (synthetic.new) quota state
via GET /v2/quotas (free, does not consume quota): the rolling 5-hour
request limit with per-tick regeneration rate, and the weekly credit
quota in USD. Applies to API-key credentials for the synthetic
provider; every payload section is parsed defensively since only
subscription is documented.
2026-07-23 08:46:18 +01:00
can1357 59877a01bd fix(update): fetch musl release asset on musl hosts and align musl release test
- omp update's getBinaryName now detects a musl Linux host (Alpine release file or /lib/ld-musl-* loader, mirroring scripts/install.sh) and downloads omp-linux-musl-<arch>, so self-update no longer replaces a musl install with the glibc build.
- Updated musl-release dry-run assertions to the Bun.build output format the binaries script now emits; promoted the musl changelog entry to [Unreleased].
2026-07-22 23:12:24 +02:00
roboomp 29f773ece2 fix(cli): disposed model-listing extensions
Emitted session_shutdown after model rendering and centralized managed timer cleanup across one-shot listings and agent sessions.

Added regression coverage for the extension shutdown lifecycle.

Fixes #6297
2026-07-22 15:51:47 +00:00
maatheusgois-dd 32d8b84e26 Add dynamic multi-root workspace context (#2569)
A session now carries an ordered list of workspace directories beyond cwd,
managed live from the terminal. New /add-dir, /remove-dir, and /dirs slash
commands let you add and remove folders mid-session; the repeatable --add-dir
CLI flag seeds them at launch, and the workspace.additionalDirectories
setting persists defaults per project. Additional roots are persisted in the
session header, survive reopen/fork/move, and are surfaced to the agent in the
system prompt so it knows they exist and can read/grep/glob them by absolute
path. Design aligns with the endorsed community implementation on
feature/session-workspace.

Co-authored-by: oh-my-pi <https://omp.sh>
2026-07-21 23:38:32 -03:00
Christian Stewart 670304eafa fix(coding-agent): resolve bare model role aliases
Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-18 03:51:09 -07:00
roboomp 22a8524058 fix(tools): recognized zip-family archives and pruned unconvertible extensions
- Treated ZIP-based .jar/.war/.ear/.apk as zip archives in archiveFormatFromPath and parseArchivePathCandidates so read/write member access works.
- Shared one archive-extension alternation between format detection and path splitting to stop them drifting.
- Derived the markit convertible-extension set from a single source of truth (utils/markit) matching the registered converters (pdf/docx/pptx/xlsx/epub), dropping legacy .doc/.ppt/.xls/.rtf that had no converter and only produced Unsupported format errors.
- Updated read/write tool prompts to document the zip-family extensions.

Fixes #5808
2026-07-17 08:04:10 +00:00
can1357 93cc1fed1c merged PR #5417: fix(openai): render native response images
# Conflicts:
#	packages/coding-agent/src/modes/components/chat-transcript-builder.ts
#	packages/coding-agent/test/agent-session-skill-keywords.test.ts
2026-07-16 03:48:21 +02:00
roboomp e291d77cb0 fix(tool): routed omp grep CLI path through expandPath
The `omp grep` subcommand resolved its path argument with a bare
`path.resolve` in `runGrepCommand`, bypassing `expandPath`. The
leading-colon strip from #5529 never fired, so `:/abs/path` was
mangled into `<cwd>/:/abs/path` and failed to resolve.

Route the path arg through `expandPath` so it inherits the leading-`:`
strip plus `@`-prefix, tilde, and unicode-space normalizations, matching
`read`/`edit`/in-agent `grep`.

Fixes #5624
2026-07-15 22:17:35 +00:00
can1357 9afedb591e feat(coding-agent): added opt-in task prewalk and tightened --tools and xdev behavior
- Added a `task.prewalk` option (default `false`), removed default task `prewalk` flags, and updated prewalk resolution so bunded generic task execution only prewalks when explicitly enabled.
- Enforced strict `--tools` validation in CLI parsing, making unknown tool names fail fast with `CliUsageError` instead of being silently filtered.
- Migrated legacy discovery settings (`tools.discoveryMode`, `tools.essentialOverride`, MCP discovery keys) into updated `tools.xdev` handling with preserved explicit override behavior.
- Hardened xdev/ACP execution flow by capping `docsAll` payloads with overflow listing and remapping `xd://` dispatches/approval gating for correct execute/read behavior and reduced duplicate prompts.
2026-07-15 18:39:36 +02:00
can1357 5ff277349c refactor(coding-agent): consolidated tool surface onto xd:// devices and hub
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
2026-07-15 15:16:29 +02:00
can1357 404ebb0fb0 Merge remote-tracking branch 'origin/farm/ae7a5593/ttsr-inline-regex-flags-and-scope-quoting' 2026-07-15 09:54:46 +02:00
can1357 9f17927ee7 Merge PR #5467: fix(cli): flush config JSON output before exit (@roboomp) 2026-07-14 22:58:46 +02:00
roboomp 86824c94e9 fix(ttsr): registered rules with inline regex flags and malformed scope
Rules whose condition led with a PCRE-style inline flag group (e.g.
`(?i)`) never registered: `new RegExp("(?i)...")` throws in Bun/JS, so
the condition failed to compile and `TtsrManager.addRule` dropped the
rule as having zero usable conditions.

- Add `compileRuleCondition` in capability/rule.ts translating a leading
  `(?i)`/`(?m)`/`(?s)` group into native RegExp flags; wire it into the
  TtsrManager and both ttsr-cli compile sites.
- Strip surrounding quotes from scope tokens so a malformed
  `scope: "text","thinking"` recovers to canonical `text`/`thinking`.
- Reparse each value in parseFrontmatter's YAML fallback so one bad line
  can't leave sibling values wrapped in literal quotes.

Fixes #4796
2026-07-14 19:01:47 +00:00
roboomp df7aa6d01f fix(cli): awaited config JSON stdout flush
- Waited for the stdout write callback before the config command exits.
- Covered JSON output larger than the 64 KiB pipe buffer.

Fixes #5309
2026-07-14 18:06:40 +00:00
can1357 946ba4bcea Merge PR #5044: fix(cli): parse max-time duration suffixes (@roboomp)
# Conflicts:
#	packages/coding-agent/src/cli/flag-tables.ts
2026-07-14 18:46:14 +02:00
can1357 e31a508140 Merge PR #5057: fix(cli): route npm self-updates through npm (@roboomp) 2026-07-14 18:41:16 +02:00
can1357 7f9a2777f9 Merge PR #5170: fix(ai): scope Anthropic credential identity by organization, not just email (@chan1103) 2026-07-14 18:29:14 +02:00
roboomp b6b947bdbe fix(openai): rendered native response images
- Normalized completed image_generation_call results into assistant image blocks.
- Persisted image bytes through the session blob store and rendered them in live, replay, ACP, proxy, telemetry, and HTML paths.
- Added response normalization, persistence, and TUI rendering regressions.

Fixes #4768
2026-07-14 16:15:53 +00:00
can1357 4c1c5f40d8 feat(launch): rendered launch logs from daemon terminal byte streams
- Changed daemon log reads to return both sanitized display text and a raw `terminalText` slice, and included it on log RPC responses for PTY runs when grep was not used.
- Extended the logs result contract and launch tool rendering to consume `terminalText`, reconstruct terminal output, and display it in framed, preview-capped sections.
- Kept terminal row layout stable by writing space characters for empty cells when reading rows, preserving spacing during output reconstruction.
2026-07-13 19:21:45 +02:00
can1357 4df6f6683d chore: finalizing the new /prewalk 2026-07-13 15:18:31 +02:00
can1357 f405525bf4 feat: removed boomerang feature and associated validation workflows
- Removed the `downshift.boomerang` feature, including CLI flags, configuration schema settings, and internal session state logic.
- Deleted the associated prompt definition file and all unit tests related to the boomerang validation flow.
- Cleaned up frontend forms and server request handling in the harbor-manager package to reflect the feature removal.
- Updated the downshift planning prompt instructions to maintain continuity in task creation.
2026-07-13 14:15:54 +02:00
can1357 95ecc61bc9 feat(coding-agent): implemented downshift boomerang flow for context handoff
- Introduced `--downshift-boomerang` CLI flag and `downshift.boomerang` configuration setting.
- Integrated boomerang validation hand-back logic and context cutting within the agent session.
- Added system prompt templates and runtime checks to facilitate message handling for boomerang transitions.
- Included unit tests to verify context management and validation pass requirements during the boomerang process.
2026-07-13 10:37:20 +02:00
can1357 9f1ff90a39 feat(coding-agent): implemented downshift and plan-yolo agent workflows
- Replaced legacy reasoning-slide functionality with new downshift and plan-yolo capabilities.
- Updated CLI arguments, slash commands, and configuration schemas to support the new model-switching and execution behaviors.
- Refactored agent session logic to handle downshift arming, plan-yolo headless execution, and context scrubbing.
- Renamed and added system prompts to align with the updated downshift and plan-yolo workflows.
2026-07-13 06:03:46 +02:00
can1357 e770cdc4d9 fix(coding-agent): unified tool renderer and clarified launch diagnostics
- Implemented a consolidated TUI renderer for launch tool operations to unify status headers and metadata.
- Added tracking for unfulfilled readiness conditions to distinguish between timeout causes and daemon states.
- Enhanced timeout reporting to explicitly surface specific unmet log or port conditions.
- Included comprehensive unit and regression tests for tool rendering and start-timeout diagnostics.
2026-07-13 05:16:36 +02:00
can1357 ac16253613 wip: rslide (experimental) 2026-07-13 04:21:54 +02:00
can1357 b6559861d0 feat(cli): standardized throughput calculation to total duration
- Removed the `computeTokensPerSecond` helper function in favor of a direct calculation.
- Standardized throughput to use total request duration instead of post-TTFT decode time.
- Updated UI usage display to calculate tokens per second against total duration.
- Prevented over-inflation of throughput metrics caused by hidden reasoning tokens.
2026-07-12 01:06:24 +02:00
can1357 6bb0878b63 feat(ai): added cache invalidation command for usage reports
- Added an `invalidate` action to the usage CLI to clear cached usage reports for specific or all providers.
- Updated `invalidateUsageCache` in `AuthStorage` to be asynchronous and notify the underlying store of invalidations.
2026-07-11 20:55:19 +02:00
chan1103 1e8f183f49 fix(ai): claim same-subscription rows across base identities; skip migrated org-only rows
Codex review round 8 on 679b45e5f:

- matchesReplacementCredential: an org-scoped incoming credential now
  matches an existing row keyed by ANY of its base identities
  (email/account/project), bare or org-qualified with the same org, plus
  the bare org-only key — so an account-keyed row stored while email
  recovery failed is claimed once a later same-subscription login
  recovers the email, instead of duplicating. Org-less incoming
  credentials keep exact-key matching only.
- auth-broker migrate: index org-only broker snapshot rows by their org
  id so reruns skip an already-migrated row instead of re-uploading a
  stale refresh token over the broker's newer one.
2026-07-11 22:49:12 +09:00
chan1103 e095af3be0 fix(ai): require member identity within a shared org for report routing, overlays, and coverage
Codex review round 7 on e90a72bdd flagged that broker usage-report
matching, header-overlay keying, and omp-usage coverage all treated a
matching organization as a sufficient match. Two Team members share the
org id while drawing on per-user pools, so the first same-org report
(or a lone sibling report) was handed to the wrong member. The org is
now a gate: within the same-org subset the member's own base identity
(account/email/project) must still match, with org-only entities (no
base identifiers) matching on the org alone when unambiguous. The
overlay merger (findMatchingReportIndex) had the identical same-org
flaw and receives the symmetric fix. Org-presence-mismatch semantics
are unchanged: org-scoped vs org-less stays fall-through/unreported,
and both-org-less keeps the legacy base-identity fallback.
2026-07-11 22:49:12 +09:00
chan1103 c2456d882f fix(ai): org-qualify account/project fallback identities; org-decisive routing on either side
Addresses the fourth review round (Codex no-email finding on d37e3992c,
confirmed and scoped by internal review):

- resolveProviderCredentialIdentityKey: the anthropic org qualifier now
  rides on whichever base identity exists (email > account > project),
  not only email. The account UUID is identical across the orgs of one
  login account, so the bare account fallback let a second subscription
  replace the first whenever the email could not be recovered (token
  response omits it AND bootstrap fails). Org-only credentials key on
  the org alone instead of losing identity entirely.
- matchesReplacementCredential: the one-way legacy claim strips a
  trailing |org: from ANY anthropic base key (account/project included);
  only anthropic keys carry the qualifier, so other providers are
  unaffected.
- Usage-report dedupe falls back to the org-qualified account for
  no-email anthropic reports instead of returning no identifiers.
- Broker report/overlay routing (matchUsageReport/findMatchingReportIndex)
  is org-decisive on EITHER side: an org-less legacy credential no longer
  receives an org-attributed sibling's pool via the lone-candidate or
  email/account fallback, and an org-less overlay only merges into
  org-less reports.
- omp usage unreported-account attribution follows the same either-side
  rule, so a legacy row whose fetch failed surfaces as 'no usage data'
  instead of being hidden by a sibling's report.
- Regression tests: no-email identity coexistence/replace/claim, no-email
  report dedupe, org-less broker routing, either-side unreported
  attribution.
2026-07-11 22:49:12 +09:00
chan1103 a47c3c90ec fix(ai): org-decisive active matching on either side; org in status-line cache key and health results
Addresses the second review round (internal re-review + Codex on c38840482):

- Active-account matching (logout preselection, /usage in-use marker) is
  org-decisive when EITHER side carries an org: a legacy bare-email active
  row no longer flags org-scoped siblings via the shared email (reverse of
  the previous fix). Both-org-less keeps the email/account fallback, so
  providers without orgs are unaffected.
- Status-line usage context key includes orgId, so rotating between two
  same-email subscriptions invalidates the cached quota immediately
  instead of showing the previous org's numbers for the cache TTL.
- CredentialHealthResult carries orgId/orgName and auth-gateway check
  labels rows with the org, so a failing row names the subscription.
- getOAuthAccountIdentity preserves org-only identities; the login
  success message renders them.
- ACP /usage account-id fallback labels get the org suffix too.
- Regression tests for both matching directions (marker + logout).
2026-07-11 22:49:12 +09:00
chan1103 bee01bfc4a fix(ai): carry org through OAuth access results; org-scoped active never matches org-less rows
Addresses Codex review on #5170:

- OAuthAccess/OAuthAccessFailure and every resolution site now carry
  orgId/orgName; dry-balance bench keys and labels are org-qualified so
  two same-email subscriptions stay two benchmark targets.
- When the active identity is org-scoped, logout active-marking and the
  /usage in-use marker match ONLY the same org — an org-less legacy row
  or pre-upgrade report can no longer be flagged active via the shared
  email, so the logout preselection cannot land on the wrong row.
- Regression tests: org-scoped active vs sibling org and vs legacy
  bare-email row (logout + in-use marker), org-suffixed logout labels.
2026-07-11 22:49:12 +09:00
chan1103 044d722a36 fix(ai): scope Anthropic credential identity by organization
One Anthropic account email can hold multiple organizations (a Team seat
plus a personal Max plan), each with its own org-scoped OAuth token and
independent 5h/7d limit pools. Credentials were deduped by bare email, so
logging in with the second subscription silently replaced the first, and
usage reports from the two pools merged into one row with mixed numbers.

- capture organization uuid/name at login (token exchange response, with
  a claude_cli/bootstrap fallback); token refreshes never rewrite it
- key anthropic credential identity as email + org; a legacy email-keyed
  row is claimed in place by the first org-scoped login with the same
  email, and org-less credentials never clobber org-scoped rows
- partition usage-report dedupe and the per-credential usage cache by
  org so the two subscriptions' limit pools stay distinct for rotation
- show the organization in omp usage (redaction-safe) and name the
  stored account/org in the login success message
2026-07-11 22:49:12 +09:00
roboomp 3876f60d65 fix(cli): routed npm self updates through npm
- Detected npm-owned global bin paths and Windows npm launcher shims before falling back to binary replacement.
- Added npm install argv construction that pins the registry and native package versions.
- Covered npm shim resolution and npm update arguments in update-cli tests.

Fixes #5053
2026-07-10 08:53:20 +00:00
roboomp d670dd5d94 fix(cli): reported max-time parse errors cleanly
- Added a CLI usage-error type for argument validation failures.
- Reported invalid --max-time values from launch and ACP as clean usage errors with exit code 2.
- Covered the no-stack-trace CLI error path for invalid max-time input.

Fixes #5041
2026-07-10 07:45:27 +00:00
roboomp 7fa2c3f42d fix(coding-agent): preserved fork prompt cache affinity
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.

- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.

- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.

Fixes #5035
2026-07-10 07:22:28 +00:00