Commit Graph

7348 Commits

Author SHA1 Message Date
can1357 fde55bf927 fix(model): added bracket-affix stripping and string-keyed resolution cache
- Replaced WeakMap model cache with provider/id string keys for stable reuse.
- Returned official model ids directly when matched, before heuristics.
- Collapsed non-message token path to system prompt and tool schema totals.
2026-06-06 22:21:58 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 5721034739 fix(ui): forced full replay on tool output expand toggle
- Replaced viewport-only repaint with resetDisplay so committed scrollback reflects new heights.
- Added per-server rust-analyzer workspace-ready timing overrides as a test seam.
- Added clearSuppressedSelectors to reset retry-fallback cooldown state.
- Removed obsolete shared eval executors test.
2026-06-06 21:56:00 +02:00
can1357 5fc443f4af fix(ai): adjusted usage ranking comparator for stable metric ordering
- Added a tolerance-aware `compareUsageRankingMetric` helper with finite-value handling.
- Replaced usage provider sorting comparisons with the new comparator for secondary and primary usage metrics.
- Kept existing tie-breaker fields while stabilizing ordering for nearly equal metric values.
2026-06-06 21:36:39 +02:00
can1357 3b5b182553 Merge remote-tracking branch 'origin/farm/10b5f6a2/surface-subagent-abort-reason' 2026-06-06 21:34:33 +02:00
can1357 6e88643dfa Merge remote-tracking branch 'origin/farm/2750a3c2/fix-todo-renderer-and-anthropic-reasoning-flag' 2026-06-06 21:33:39 +02:00
can1357 1b73329b51 Merge remote-tracking branch 'origin/farm/e0b4f6c6/update-kagi-v1-docs' 2026-06-06 21:33:16 +02:00
can1357 133137c9a6 fix(eval): surfaced subagent abort reasons and disabled runtime cap
- Used `||` so empty stderr falls through to abortReason in agent bridge.
- Preferred assistant errorMessage over "Cancelled by caller" on internal aborts.
- Forced `maxRuntimeMs: 0` for eval subagents via ExecutorOptions override.
2026-06-06 21:33:07 +02:00
can1357 0bac7012cd feat(coding-agent/web): added GitHub Actions run/job scraping
- Parsed /actions/runs URLs into run and job render handlers.
- Rendered run metadata with per-job breakdown, showing steps for failed jobs.
- Fetched job logs via API token, stripping ISO timestamp prefixes.
2026-06-06 21:31:49 +02:00
can1357 bdbbfa9778 fix(eval): surfaced subagent abort reason
- Used `||` so empty stderr no longer masks the real abort reason.
2026-06-06 20:54:00 +02:00
roboomp 2620d3970c docs(web-search): updated kagi description to v1 endpoint
The runtime cutover to Kagi's V1 search API landed in #1272 but
docs/tools/web_search.md still described the sunset V0 endpoint
(GET /api/v0/search with 'Authorization: Bot ...'). Realigned the
Querying and Output bullets with the actual implementation in
packages/coding-agent/src/web/kagi.ts:

- POST https://kagi.com/api/v1/search with Bearer auth and JSON body.
- recency maps to filters.after as a UTC YYYY-MM-DD string.
- Output now includes the categorized bucket merge (search/video/news/
  infobox with title tags), adjacent_question + related_search related
  questions, direct_answer-derived answer, and meta.trace requestId.

Fixes #2009
2026-06-06 18:50:37 +00:00
Can Bölük 4ae58e1abc Merge pull request #2004 from basedcorp99/fix/antigravity-usage-display
fix(usage): sanitize Antigravity /usage display — dedupe by tier, fix account count, merge window data
2026-06-06 20:48:52 +02:00
Can Bölük fa7271c2c2 Merge branch 'main' into fix/antigravity-usage-display 2026-06-06 20:48:25 +02:00
Can Bölük 4fb0742da7 Merge pull request #2002 from basedcorp99/fix/cca-mixed-type-collapse-strips-type-keys
fix(ai): strip type-specific keys when CCA mixed-type collapse picks non-matching type
2026-06-06 20:48:12 +02:00
basedcorp99 f854596f83 fix: format antigravity usage tests 2026-06-06 20:43:19 +02:00
basedcorp99 f8ef2cf8e4 fix: format CCA schema normalization 2026-06-06 20:42:20 +02:00
basedcorp99 7c8fb4d8f6 fix(ai): also strip sibling type-specific keys during CCA mixed-type collapse
Address review feedback:
- Replace `as string` assertion with typed `chosenType` local
- Strip sibling keys from nextSchema that were copied via
  copySchemaWithout but belong to a type other than the chosen one
  (e.g. sibling `items` on a now-string-typed schema)
- Export ALL_CCA_TYPE_SPECIFIC_KEYS from fields.ts for sibling filtering
- Add regression test for the sibling-key edge case
2026-06-06 20:42:20 +02:00
basedcorp99 dee4db1602 chore(ai): update changelog with PR number 2026-06-06 20:42:20 +02:00
basedcorp99 807df56ba7 docs(ai): add unreleased changelog entry for CCA mixed-type combiner collapse fix 2026-06-06 20:42:20 +02:00
basedcorp99 2623bd75a2 test(ai): add stripResidualCombiners regression for mixed-type string|array collapse 2026-06-06 20:42:20 +02:00
basedcorp99 f3210ab862 fix(ai): strip type-specific keys when CCA mixed-type collapse picks non-matching type
When collapseMixedTypeCombinerVariants collapses an anyOf with mixed
types (e.g. string | array), it previously picked the first non-null
type but indiscriminately copied ALL mergedVariantFields — including
type-specific keys like "items" that only belong to array. This
produced schemas like {type: "string", items: {...}} which Google
Cloud Code Assist API rejects with 400.

Fix: filter mergedVariantFields against the chosen types allowed keys
(CLOUD_CODE_ASSIST_TYPE_SPECIFIC_KEYS) before copying, so array-only
keys are dropped when the winner is string (and vice versa).

Fixes 400 error on github tools "pr" parameter (anyOf string/array).
2026-06-06 20:42:20 +02:00
basedcorp99 60f79199ca fix(usage): only add project: dedup identifier when no email; preserve raw windowId
- #getUsageReportIdentifiers now only pushes project: when no email was
  found, preventing two users with different emails on the same GCP
  project from being merged at the usage-report level.

- Antigravity dedup key now uses quotaInfo.windowId directly before
  falling back to parseWindow's id. When resetTime is absent but
  windowId is set, separate windows no longer collapse to 'default'.
2026-06-06 20:42:15 +02:00
basedcorp99 794a64aae1 fix(usage): address review — order email before project, add tests, nits
- BLOCKING: reorder resolveProviderCredentialIdentityKey so email
  identity takes priority over project — two users with different
  emails on the same GCP project no longer get merged/hard-deleted.

- Added #getUsageReportScopeProjectId helper so Gemini CLI reports
  (which set projectId on limit.scope but not metadata) still get
  dedup coverage. Both metadata and scope projectId paths checked.

- formatAggregateAmount now falls back to limits.length when no
  scope.accountId values are present, preserving pre-existing
  behaviour for providers that don't set accountId on limits.

- Added 9 contract tests for the antigravity usage merge logic:
  tier dedup, worst-fraction-wins, mixed-case collapsing,
  reset-time-from-other-entry, windowId separation, metadata,
  sort order, and null-on-no-project.

- Nits: label='Usage' (so formatLimitTitle renders 'Usage (Default)'
  not bare 'Default'), id uses params.provider instead of hardcoded
  string, tier field drops redundant ?? undefined.
2026-06-06 20:42:15 +02:00
basedcorp99 15c0dff28e fix(usage): fall back to limit.scope.projectId when metadata.projectId is absent
Gemini CLI provider stores projectId on limit.scope but not in report
metadata, so the metadata-only projectId fallback added earlier missed
that case. Now all three lookup sites (dedup identifiers, TUI account
label, ACP account label) also check limit.scope.projectId.
2026-06-06 20:42:15 +02:00
basedcorp99 ca24f34044 fix(usage): merge antigravity dedup entries — keep bar data + reset time
When models within the same (tier, windowId) group have complementary
data — some carry remainingFraction but no resetTime, others carry
resetTime but no remainingFraction — merge them so the displayed entry
has both a real bar and the 'resets in…' line.

Also lowercases tier names for dedup keys so 'Default' and 'default'
are recognized as the same tier.

Adds projectId to OAuth credential identity extraction and to the
usage-report dedup identifiers so duplicate credential rows (same
Google Cloud project, separate login sessions) are pruned and merged
at both the store and usage-report levels.
2026-06-06 20:42:15 +02:00
basedcorp99 c933d34398 fix(usage): antigravity /usage display — dedupe by tier, fix account count, add projectId identity
- Antigravity usage provider now deduplicates model quota entries by tier
  instead of emitting one bar per model (15+ redundant bars for one account).
  The upstream API groups quota by tier — models within the same tier share
  the same quota bucket, so per-model bars were misleading noise.

- Reports now carry credential email and accountId in metadata so the
  /usage display and deduplicator can show meaningful account identities
  instead of 'account 1'.

- formatAggregateAmount no longer uses limits.length as account count.
  Instead counts unique accountId values from limit scopes — a single
  account's N incomplete limits no longer display as 'N accts'.

- Usage report dedup now considers metadata.projectId for Google Cloud
  providers so duplicate credential rows with the same project merge.

- account labels in both TUI and ACP markdown paths now fall back to
  metadata.projectId before the generic 'account N' placeholder.
2026-06-06 20:42:15 +02:00
roboomp 5046f0f47c fix(ai): treat missing anthropic baseUrl as official in thinking replay
resolveAnthropicBaseUrl defaults to https://api.anthropic.com when
model.baseUrl is absent (e.g. same-id custom overrides that only tweak
metadata), and the existing isAnthropicApiBaseUrl helper already treats
an empty/undefined baseUrl as official. The new isOfficialAnthropicEndpoint
helper classified the same model as non-official, so
shouldReplayUnsignedThinking would replay unsigned thinking as
type: thinking against the first-party API — which rejects it.

Drop the redundant helper and use isAnthropicApiBaseUrl as the single
source of truth. Add a regression test that pins the missing-baseUrl
case to the text fallback.
2026-06-06 18:34:17 +00:00
roboomp d08e4ebb1f fix(ai): generalize anthropic unsigned thinking replay
Anthropic-compatible reasoning providers commonly emit thinking blocks
without first-party Anthropic signatures while still expecting those
blocks back as native thinking on continuation. The previous follow-up
for #2005 fixed Xiaomi by provider/host allowlist, but the protocol
contract is broader and matches the behavior described in #1996.

Replace the Xiaomi-specific branch with a protocol-level rule:
- official api.anthropic.com keeps demoting unsigned thinking to text
- non-official anthropic-messages reasoning models replay unsigned
  thinking as type: thinking with an empty signature
- existing known non-signing DeepSeek/Z.AI compatibility remains

The regression test now uses a generic Anthropic-compatible reasoning
endpoint, includes the Xiaomi MiMo reporter configuration only as a
fixture, and guards non-reasoning unknown endpoints plus official
Anthropic behavior.

Fixes #2005
2026-06-06 18:29:43 +00:00
roboomp ae3477abfd style: bun run fix 2026-06-06 18:20:57 +00:00
roboomp 6dcbb07793 fix(ai): replay xiaomi mimo anthropic-compat thinking blocks unsigned
The Anthropic-compat endpoints hosted under *.xiaomimimo.com (every
Xiaomi MiMo Token Plan region plus api.xiaomimimo.com) emit thinking
blocks without a signature. convertAnthropicMessages defaulted to
"signing capable" for any endpoint not explicitly allowlisted as
non-signing, so MiMo's unsigned thinking blocks were demoted to text on
every continuation request. Without its prior reasoning replayed, MiMo
destabilized tool-call argument serialization — the root cause behind
the args?.ops?.map crash already mitigated at the renderer in #2005.

Extend isNonSigningAnthropicEndpoint to cover the xiaomi catalog
provider, every xiaomi-token-plan-* provider id, and any baseUrl on
xiaomimimo.com so the existing non-signing replay branch fires for MiMo
the same way it does for DeepSeek and Z.AI.

Fixes #2005
2026-06-06 18:20:27 +00:00
roboomp e83dbc177b style: bun run fix 2026-06-06 18:15:09 +00:00
roboomp cab465cba4 fix(eval): surfaced subagent abort reason through python agent() bridge
Python eval agent() collapsed every subagent runtime-limit abort into a
generic 'RuntimeError: bridge call __agent__ failed' instead of the real
reason. runEvalAgent built its failure message with:

  result.error ?? result.stderr ?? result.abortReason ?? <default>

? is nullish-coalescing, so result.stderr = "" (the executor's value for
a runtime-limit abort) short-circuited the chain and never reached
abortReason. The host bridge then shipped {ok: false, error: ""}, and
prelude.py's '<msg> or <fallback>' picked the named-bridge fallback.

Extracted buildSubagentFailureMessage(): aborted subagents prefer the
trimmed abortReason; otherwise fall through error, stderr (trimmed),
abortReason, and the named-bridge default. Empty/whitespace strings no
longer mask anything. The failure-detection condition also accepts
result.aborted so an abort with exitCode 0 (theoretically) still flows
the abort reason out.

Added a regression test asserting that runtime-limit aborts, whitespace
stderr/error, and totally blank aborts all produce non-empty messages
matching the executor's abortReason text.

Fixes #2006
2026-06-06 18:15:04 +00:00
can1357 a7f5e83067 chore: bump version to 15.9.69 2026-06-06 20:13:59 +02:00
can1357 43b22e9564 fix(coding-agent/web): defaulted perplexity ask to experimental model
- Forced authenticated ask requests to `experimental`, matching the anonymous fallback since the cookie session ignores pro upgrades.
- Kept TUI collapsed search answers full; capping now only applies in compact mode via `maxAnswerLines`.
- Preserved full multiline task pending preview instead of bounding it.
2026-06-06 20:13:06 +02:00
roboomp 1c5cb37df5 style: bun run fix 2026-06-06 18:12:08 +00:00
roboomp ce60b6626b fix(coding-agent): harden todo renderer against malformed streaming args
The todo tool's renderCall ran args?.ops?.map(...) directly, which throws
TypeError on any non-array ops value. parseStreamingJson surfaces such
shapes mid-stream: a partial Anthropic input_json_delta buffer like
'{"ops":"[{' becomes { ops: '[{' }, and intermediate states can hand back
null entries before object fields arrive. Each crash spammed Tool
renderer failed warnings and starved the TUI render loop.

Guard against:
- ops being any non-array (string, object, primitive)
- entries being null / non-object
- entry.items being a non-array

The fix is in the TUI renderer only — schema validation in the agent
loop is unchanged, so any genuinely malformed model output still
surfaces an invalid-args tool error to the model.

Fixes #2005
2026-06-06 18:11:42 +00:00
can1357 246688e874 fix(coding-agent/web): unlocked perplexity pro via session cookie
- Sent the OAuth token as `__Secure-next-auth.session-token` cookie since the ask endpoint ignores bearer headers and silently downgrades to `turbo`.
- Fell back to `result.title` when web results omit `name`.
- Renamed `callPerplexityOAuth` to `callPerplexityAsk` and removed a stray brace.
- Added tests covering OAuth, API-key, and anonymous request shapes.
2026-06-06 20:01:51 +02:00
can1357 f73892d491 fix(coding-agent): bounded expanded single-file search results
- Stopped expanded view from dumping every match when all hits share one file.
- Applied an `EXPANDED_LINES × 2` budget while keeping context rows.
- Appended a `… N more matches` summary when truncated.
2026-06-06 20:01:21 +02:00
can1357 8a5b99a967 feat(coding-agent): enabled anonymous Perplexity fallback and updated web-search checks
- Added anonymous Perplexity authentication mode for unauthenticated web searches.
- Switched web-search setup checks to use `isExplicitlyAvailable` and removed key enforcement in doctor.
- Updated Perplexity OAuth flow to reuse auth handling for all non-key searches and anonymous responses.
- Updated CLI and provider option help text to mark the Perplexity key optional with fallback.
2026-06-06 19:42:16 +02:00
can1357 d63a91bf5e fix(coding-agent/web): sent bare query on perplexity OAuth path
- Stopped prepending system_prompt to the consumer ask endpoint, which lacks a system slot and refused the meta-instruction.
- Kept system_prompt as a proper system message on the API-key path.
2026-06-06 19:36:30 +02:00
can1357 1f3f19ab13 Merge remote-tracking branch 'origin/farm/438dde40/tool-use-flicker' 2026-06-06 19:27:31 +02:00
roboomp c43eac9d91 style: bun run fix 2026-06-06 17:10:25 +00:00
roboomp adcb8793b2 fix(tui): gated deccara fills on sync output
Prevented DECCARA background-fill optimization from shortening rows unless the active TUI paint is protected by synchronized output, preserving padded background bytes when sync output is disabled.

Added regression coverage for the synchronized-output opt-out path and kept existing DECCARA tests forced onto synchronized output so the optimized path remains covered.

Fixes #2000
2026-06-06 17:10:18 +00:00
can1357 2887eec373 feat(cli): added PNG screenshot support for the gallery CLI command
- Added `omp gallery --screenshot`, `--out`, `--font`, and `--font-size` flags.
- Added a VHS-based screenshot path that captures gallery output as PNG file(s).
- Added chunking and naming logic to split tall galleries into multiple numbered captures.
2026-06-06 19:05:53 +02:00
can1357 ead5cf6871 fix(coding-agent/scripts): fixed CLI startup by launching through a bunfig-free shim script
- Updated `install:dev` to symlink `packages/coding-agent/scripts/dev-launch` into Bun's global bin directory as `omp`.
- Added a `dev-launch` shell script that launches Bun from an isolated directory and preserves the caller's working directory for restoration.
- Added a preload shim that restores `OMP_LAUNCH_CWD` before CLI execution so external project `bunfig.toml` preloads are not used.
2026-06-06 19:05:20 +02:00
can1357 c642232266 ux(coding-agent/tools): improved tool error rendering with subordinate detail lines
- Added sanitizeErrorText in render-utils to normalize and truncate tool error messages.
- Introduced formatErrorDetail for indented subordinate error text without redundant icon or Error prefix.
- Updated goal and write tool renderers to use the new detail formatter, with write now handling isError results via a status header plus detail line.
2026-06-06 19:02:16 +02:00
can1357 9aa10dd92b ux(coding-agent): condensed web_search result rendering
- Showed answer text in full in the TUI; kept the `omp q` compact cap.
- Rendered each source as a single title/domain/age line with the URL linked on the title.
- Collapsed the metadata block to one Provider line plus Usage.
- Rendered search errors as a framed panel matching the success layout.
2026-06-06 18:52:46 +02:00
can1357 d1fbb28edc fix(coding-agent): removed redundant tool-name line in custom render
- Fixed custom-rendered tools with `mergeCallAndResult` (e.g. `lsp`) emitting a redundant tool-name line above the framed result.
- Collapsed the leading blank line for self-delimiting framed boxes.
- Added gallery fidelity routing `lsp`/`task` through the custom-tool branch via a `customRendered` fixture flag.
- Added gallery harness tests guarding state coverage and the custom-branch fallback label.
2026-06-06 18:40:52 +02:00
can1357 75e211415e feat(cli): added gallery CLI command for renderer previews and filtering options
- Added lazy-loaded `gallery` command registration and new filters for tool, state, width, expanded, and plain output.
- Implemented gallery state rendering with terminal-width defaults, state filtering, and unknown-tool fallback handling.
- Added shared fixture types and aggregated renderer fixtures for multiple tool families in `galleryFixtures`.
- Added tests for renderer state coverage, route-specific output (streaming/progress/success/error), and fixture fallback.
2026-06-06 18:28:19 +02:00
can1357 c49d5c99b1 fix(coding-agent): removed preview line capping on context lines
- Rendered full context lines instead of truncating via capPreviewLines.
2026-06-06 18:21:09 +02:00