Commit Graph
3995 Commits
Author SHA1 Message Date
can1357 2efbdeaab9 refactor(ai): deduplicated usage provider coercion helpers
- Deleted usage/shared.ts's duplicate toNumber; every consumer in the directory
  now resolves the single catalog implementation.
- Promoted HOUR_MS, DAY_MS, WEEK_MS, parseIsoTimestamp, parsePositiveTimestamp
  and usageStatus into usage/shared.ts and dropped the local copies.
- Left the epoch coercers, status thresholds and base-url normalizers alone:
  they disagree on the 1e12 boundary, non-positive values and warn levels, so
  merging them would change reported usage.
2026-08-08 06:32:00 +02:00
can1357 5fa578fa3b refactor(ai): migrated hand-rolled api-key logins to shared factory
- Nine providers reimplemented createApiKeyLogin's exact prompt/trim/abort flow
  verbatim; they now call the existing helper instead.
- Extended ApiKeyLoginConfig with optional authUrl/instructions and an
  emptyKeyFallback so the local-server family (llama.cpp, vLLM, LM Studio)
  collapses onto the same helper without changing observable behavior.
- Left ollama, ollama-cloud, nvidia, qwen-portal, github-copilot and both
  alibaba flows hand-written: their error classes, messages or callback
  ordering differ, so migrating them would change observable behavior.
2026-08-08 06:32:00 +02:00
can1357 055a5d4f26 chore: bump version to 17.2.11 2026-08-07 23:38:40 +02:00
can1357 a9dcf0f8d1 chore: update changelogs 2026-08-07 23:38:26 +02:00
roboompandcan1357 a09dfd0ba8 fix(extensions): restored provider unregistration
Added the upstream unregisterProvider lifecycle to queued and initialized extension runtimes. Provider removal now clears runtime model/auth state before replacement, while failed factories restore the prior registration queue.

Fixes #7914
2026-08-07 23:38:25 +02:00
can1357 850225a694 style: applied biome formatting after community fix merges 2026-08-07 13:40:52 +02:00
can1357 422f5d137f chore: normalized changelogs after merging community fixes 2026-08-07 13:40:09 +02:00
can1357 85b5d0ac1f fix(ai): keep Cursor session tokens on default origin 2026-08-07 13:37:55 +02:00
can1357 f1ad890d3f Merge PR #7613: feat(ai): add Cursor personal usage reporting (@chessl) 2026-08-07 13:37:55 +02:00
can1357 10863ef6b4 Merge PR #7892: fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat (@voonfoo) 2026-08-07 13:37:54 +02:00
can1357 f087a9050d fix(ai): keep plan per-minute limits transient 2026-08-07 13:37:54 +02:00
can1357 68dece477f Merge PR #7772: fix(session): handle subscription-cap retry exhaustion (@roboomp) 2026-08-07 13:37:54 +02:00
can1357 3d12722dba fix(ai): keep Chinese transient caps out of quota rotation 2026-08-07 13:37:54 +02:00
can1357 50cc615e45 Merge PR #7828: fix(ai): classify Simplified Chinese quota exhaustion as credential-rotatable (@IceCodeNew) 2026-08-07 13:37:54 +02:00
can1357 6be1b41d3b Merge PR #7821: fix(ai): preserve Copilot integrator entitlement errors (@roboomp) 2026-08-07 13:37:53 +02:00
can1357 3c35180de7 Merge PR #7878: fix(auth): purge pre-org OAuth tombstone on org-scoped re-login (@roboomp) 2026-08-07 13:37:53 +02:00
Voon Foo 61ff318e6d review: typed compat access, 0-disable lazy watchdog test, changelogs
- Replace the compat cast with an annotated CompatOf<Api> local narrowed
  via the in operator: assignment up-cast, compiler-checked field type,
  no as assertion.
- Cover the 0 sentinel end-to-end through the lazy wrapper with fake
  timers (mirrors the direct-Anthropic 0-disable test): advance 400s
  past the generic budget, assert no watchdog abort, then cancel
  cleanly.
- Add Unreleased changelog entries for pi-ai and pi-catalog with
  external attribution for #7892.
2026-08-07 15:39:04 +08:00
Voon Foo 2f24d4457e fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).

- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
  a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
  Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
  tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
  and lazy terminal errors carry the structural errorId classification
  so session auto-retry classifies stalls without text matching.
2026-08-07 14:18:46 +08:00
roboomp 648d858bb1 fix(auth): purged pre-org oauth tombstone on org-scoped re-login
#purgeSupersededDisabledRows matched disabled rows against active ones by
exact identity-key string equality, but the active-replacement path it
mirrors (matchesReplacementCredential) claims pre-org legacy rows
(`<b>` vs `<b>|org:<o>`). So a later org-scoped login of the same account
never purged the pre-org tombstone, which then rendered forever as a red
row in `omp usage` with no CLI/TUI escape. This is the OAuth half of the
class of bug #2943 fixed for api_key rows in the same function.

Reuse matchesReplacementCredential in the purge so an org-scoped login
claims and hard-deletes its pre-org tombstone, inheriting the one-way
upgrade and shared-workspace guards unchanged.

Fixes #7876
2026-08-07 04:17:10 +00:00
roboomp 65f0d5869b fix(anthropic): honor ANTHROPIC_BASE_URL for chat requests
resolveAnthropicBaseUrl() resolved the chat base URL from github-copilot,
FOUNDRY_BASE_URL, then model.baseUrl -> hardcoded api.anthropic.com, and
never read $env.ANTHROPIC_BASE_URL. The stock anthropic descriptor pins
model.baseUrl to api.anthropic.com, so the env fallback was unreachable:
gateway-scoped keys were sent to api.anthropic.com (401, credential leak)
regardless of ANTHROPIC_BASE_URL, contradicting docs and the web-search
fix in #1693.

- Chat resolver now returns ANTHROPIC_BASE_URL (after Foundry, ahead of the
  official default); an explicit non-official model.baseUrl still wins.
- resolveAnthropicCustomHeaders keys off the resolved base URL so
  ANTHROPIC_CUSTOM_HEADERS reach env-configured non-official gateways.
- stream.ts leaked-thinking heal exemption mirror updated to the same
  effective-endpoint precedence.

Fixes #7874
2026-08-07 03:54:45 +00:00
IceCodeNew e8e3fbc4c8 fix(ai): classify Simplified Chinese quota exhaustion as credential-rotatable
Zhipu Coding Plan returns '429 已达到 5 小时的使用上限。您的限额将在 … 重置。'
(type=1308) when the 5h window is spent. The error classifier only matched
English quota phrasing, so this message classified as UNKNOWN, Flag.UsageLimit
was never set, and multi-key sessions stayed pinned to the exhausted api_key
credential instead of rotating to a sibling key.

Add CN_QUOTA_EXHAUSTED_PATTERN (达到…使用上限, 已达上限, 额度/配额…耗尽/用完,
限额…重置, 余额不足) consulted by parseRateLimitReason before the transient
branches and by matchesUsageLimitText. Treat Simplified Chinese error bodies
as informative in isOpaqueStatusBody so a plain Chinese throttle (已达到速率限制)
does not rotate credentials via the opaque-429 fallback.

The 达到…使用上限 arm requires the 使用 token, so a concurrency or rate cap
phrased as 达到…上限 (without 使用) stays in the upstream-backoff lane instead
of being misclassified as a credential-exhausting quota.
2026-08-06 22:13:34 +08:00
can1357 3e2e715b09 test(ai): aligned deepseek flash ladder expectations with #7668
- deepseek-v4-flash bakes the wire-exact [low, high, max] ladder on
  every host since 736b496cc6; V4 Pro stays [high, max].
- The stale xhigh alias-filter assertions now expect the flash ladder.
2026-08-06 14:18:50 +02:00
can1357 43c1b245e7 chore: bump version to 17.2.10 2026-08-06 13:32:34 +02:00
can1357 9e738dc880 chore: reformat + rewrite changelogs 2026-08-06 13:30:08 +02:00
roboomp 1bcf08c27f fix(ai): preserved Copilot integrator entitlement errors
- Classified model_not_available_for_integrator as a permanent entitlement denial instead of transient fleet skew.

- Preserved the provider response and Available models list while retaining model_not_supported fleet retries.

Fixes #7819
2026-08-06 09:50:32 +00:00
roboomp f6c5a43a1f fix(session): handled subscription-cap retry exhaustion
- Classified subscription and plan rate caps as credential-rotatable usage limits while preserving transient per-minute throttles.

- Applied reason-specific backoff to transient rate limits and collapsed exhausted retry attempts behind one budget-labeled terminal error.

- Covered classification, delay selection, persisted transcript aggregation, and retry event propagation.

Fixes #7767
2026-08-06 01:31:22 +00:00
can1357 1b7b8bb0c7 Merge PR #7743: fix(agent): strip output statuses from remote compaction (@roboomp) 2026-08-05 22:15:46 +02:00
can1357 0b0e7dd530 chore: normalized changelogs after merging 20 pull requests 2026-08-05 21:50:46 +02:00
roboomp 86a856361a fix(agent): stripped output statuses from remote compaction
- Reused the Responses replay lifecycle policy for V1 and V2 compaction input.

- Covered persisted native history and converted assistant replay items.

Fixes #7742
2026-08-05 17:26:32 +00:00
can1357 e9888367d1 refactor: migrated packages to internal utility modules and removed external dependencies
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
2026-08-05 13:39:09 +02:00
can1357 f7f8e040ee chore: bump version to 17.2.9 2026-08-05 03:07:47 +02:00
can1357 b0a94a8fc0 chore: cleanup 2026-08-05 03:07:16 +02:00
can1357 eb56330c9c chore(changelog): normalized unreleased entries after merges 2026-08-05 02:32:48 +02:00
roboompandcan1357 99748bbe61 fix(ai): omitted codex optional response defaults
Stopped OpenAI Codex requests from sending reasoning.summary, reasoning.context, and text.verbosity unless explicitly configured, matching the safer native Codex request shape for GPT-5.x models.

Preserved explicit overrides and added regression coverage for transformer and settings-aware stream behavior.

Fixes #4949

(cherry picked from commit 15d86667d6572faf453e74b922eed33dd2d3f3c9)
2026-08-05 02:32:36 +02:00
can1357 a6eafcbf6b fix(ai): keep concurrency caps out of auth rotation
(cherry picked from commit bd6285ad9b16ae6f0a75e336a1c4bf961d3e43c7)
2026-08-05 02:32:36 +02:00
metaphoricsandcan1357 e322f28fcc fix(ai): exclude concurrency caps from 403 rotation and test the Copilot gate
403 concurrency caps bypass credential rotation in stream and auth-retry paths; snake-case concurrency codes classify; a session-level test covers the Copilot credential-removal gate.

(cherry picked from commit aca348e797ed987750a14ecf625952b6b971f7a5)
2026-08-05 02:32:35 +02:00
metaphoricsandcan1357 5726fac646 fix(ai): tighten concurrency-cap classification and account-cap gating
Reset-window rotation requires account-specific wording; concurrency caps require an actual cap signal; credential removal gated on AuthFailed without UsageLimit so a valid-but-blocked 403 credential is retained.

(cherry picked from commit 2f72752c2586352a4f7e9e814af1cdb0cf192af4)
2026-08-05 02:32:35 +02:00
metaphoricsandcan1357 7262d9762e fix(ai): honor account reset windows and statusless concurrency caps
Account-reset hint evaluated before short retry hints; account-scoped caps rotate on status 403 or undefined (Devin statusless trailer); statusless concurrency caps marked transient; transient same-model retries use the concurrency backoff.

Refuted: quota-worded concurrency caps were already excluded from rotation before the usage-limit text match.
(cherry picked from commit f2b9a18d715ddbcb6ae703670f2212da36bb2826)
2026-08-05 02:32:35 +02:00
metaphoricsandcan1357 9d51aabce9 fix(ai): address PR review feedback (#6958)
(cherry picked from commit 19327455a482ffdf8dfebcfb47b12fe0d4c888da)
2026-08-05 02:32:35 +02:00
metaphoricsandcan1357 216f863063 fix(ai): classified concurrency limits and rotated on account-scoped 403s
(cherry picked from commit 23bf29b62cfbb54757e67020ef45d870c1877b9e)
2026-08-05 02:32:35 +02:00
can1357 77d537add7 Merge PR #7618: fix(ai): respect explicit Codex usage allowance (@roboomp) 2026-08-05 02:22:00 +02:00
can1357 34ecb4a107 chore(changelog): normalized unreleased entries after merges 2026-08-05 01:13:13 +02:00
can1357 94c838faea Merge PR #7539: fix(coding-agent): complete usage-aware fallback integration (@eggpeat) 2026-08-05 01:12:02 +02:00
can1357 9a9c3b1541 Merge PR #7589: refactor(ai): isolate SQLite credential store (@metaphorics) 2026-08-05 01:12:02 +02:00
can1357 975d0a2e9d Merge PR #7501: fix(agent): classify thinking-only unexpected stops (@roboomp) 2026-08-05 01:12:00 +02:00
can1357 a0f203bfd4 Merge PR #7592: fix(cursor): correct inline read range metadata (@roboomp) 2026-08-05 01:11:58 +02:00
can1357 13d9a09999 Merge PR #7612: chore(ts): enforce noImplicitOverride (@metaphorics) 2026-08-05 01:11:31 +02:00
Magicien f3ad5ede6b fix(ai): scope the Copilot 8-attempt retry budget to model flaps
The 3 -> 8 bump in callWithCopilotModelRetry was shared by the generic
retryable branch, so a persistent status-less transport blip on Copilot
would ramp across 8 attempts (~11.2s of dead time) instead of the 3 it
took before, and a repeated Retry-After 429 could stretch the same way
on top of the transport's own fetchWithRetry budget.

Derive the budget from the failure kind: model-availability 400s keep the
8-attempt reroll, everything else caps at the previous 3.

Also read COPILOT_TRANSIENT_MODEL_CODES with Object.hasOwn — `code` is
provider-controlled, so a 400 body whose code was `__proto__` or
`toString` classified as transient through the prototype chain.
2026-08-04 21:51:35 +01:00
Magicien d33f2a1658 fix(ai): retry GitHub Copilot fleet-skew model 400s
Any Copilot model in the middle of a rollout (claude-sonnet-4.6,
claude-opus-4.6, gpt-5.4, gpt-5.3-codex, ...) returned a raw HTTP 400 on
roughly half of all turns. GET /models on api.githubcopilot.com returns
two different catalogs across repeated calls: part of the fleet serves
those ids, part rejects them with

  400 {"error":{"message":"The requested model is not available for
  integrator \"copilot-language-server\". ...",
  "code":"model_not_available_for_integrator", ...}}

The absorb machinery already existed and was correct
(isCopilotTransientModelError -> isProviderRetryableError -> the
Anthropic transport's PROVIDER_MAX_RETRIES). Only the classifier missed:
it matched the older model_not_supported code and probed err.code /
err.error.code, while the real code is model_not_available_for_integrator
sitting at err.error.error.code (the SDK stores the parsed body on
.error, and Copilot's body is itself {error:{code}}). So
isProviderRetryableError fell through to "4xx => terminal".

Fix the classifier: providerErrorCode() walks the error envelope up to
depth 3 instead of hardcoding a shape, both Copilot model-availability
codes are accepted, and a wire-body text match backs it up because SDK
envelope shapes drift between provider families while the stringified
message does not. This alone restores the retry path, because
isProviderRetryableError consults the provider hook before its
4xx short-circuit.

Retry shape, since a rejection is a per-request replica reroll rather
than upstream backpressure:

- both transports wait a flat delay between model-flap attempts instead
  of the growing backoff; generic retryable failures (429/5xx/transport)
  keep their linear ramp and Retry-After handling
- the OpenAI-transport budget goes 3 -> 8 attempts, because a measured
  ~70% flap window produced a turn that needed 6 wire attempts and
  exhausting the budget escalates to the agent-level retry, which
  restarts the whole turn

Absorbed attempts cost no tokens: rejections are gateway-side, carry no
usage block, and return in ~208ms median versus ~1884ms for a served
request.

Also refresh the exhausted-retry guidance text, which cited a
nonexistent model id and described the cause as a per-client rollout gap
rather than fleet skew.
2026-08-04 21:45:44 +01:00
metaphorics 3802d2dd80 chore(ts): enforce noImplicitOverride 2026-08-05 02:49:01 +09:00