- Replace the compat cast with an annotated CompatOf<Api> local narrowed
via the in operator: assignment up-cast, compiler-checked field type,
no as assertion.
- Cover the 0 sentinel end-to-end through the lazy wrapper with fake
timers (mirrors the direct-Anthropic 0-disable test): advance 400s
past the generic budget, assert no watchdog abort, then cancel
cleanly.
- Add Unreleased changelog entries for pi-ai and pi-catalog with
external attribution for #7892.
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).
- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
and lazy terminal errors carry the structural errorId classification
so session auto-retry classifies stalls without text matching.
#purgeSupersededDisabledRows matched disabled rows against active ones by
exact identity-key string equality, but the active-replacement path it
mirrors (matchesReplacementCredential) claims pre-org legacy rows
(`<b>` vs `<b>|org:<o>`). So a later org-scoped login of the same account
never purged the pre-org tombstone, which then rendered forever as a red
row in `omp usage` with no CLI/TUI escape. This is the OAuth half of the
class of bug #2943 fixed for api_key rows in the same function.
Reuse matchesReplacementCredential in the purge so an org-scoped login
claims and hard-deletes its pre-org tombstone, inheriting the one-way
upgrade and shared-workspace guards unchanged.
Fixes#7876
resolveAnthropicBaseUrl() resolved the chat base URL from github-copilot,
FOUNDRY_BASE_URL, then model.baseUrl -> hardcoded api.anthropic.com, and
never read $env.ANTHROPIC_BASE_URL. The stock anthropic descriptor pins
model.baseUrl to api.anthropic.com, so the env fallback was unreachable:
gateway-scoped keys were sent to api.anthropic.com (401, credential leak)
regardless of ANTHROPIC_BASE_URL, contradicting docs and the web-search
fix in #1693.
- Chat resolver now returns ANTHROPIC_BASE_URL (after Foundry, ahead of the
official default); an explicit non-official model.baseUrl still wins.
- resolveAnthropicCustomHeaders keys off the resolved base URL so
ANTHROPIC_CUSTOM_HEADERS reach env-configured non-official gateways.
- stream.ts leaked-thinking heal exemption mirror updated to the same
effective-endpoint precedence.
Fixes#7874
Zhipu Coding Plan returns '429 已达到 5 小时的使用上限。您的限额将在 … 重置。'
(type=1308) when the 5h window is spent. The error classifier only matched
English quota phrasing, so this message classified as UNKNOWN, Flag.UsageLimit
was never set, and multi-key sessions stayed pinned to the exhausted api_key
credential instead of rotating to a sibling key.
Add CN_QUOTA_EXHAUSTED_PATTERN (达到…使用上限, 已达上限, 额度/配额…耗尽/用完,
限额…重置, 余额不足) consulted by parseRateLimitReason before the transient
branches and by matchesUsageLimitText. Treat Simplified Chinese error bodies
as informative in isOpaqueStatusBody so a plain Chinese throttle (已达到速率限制)
does not rotate credentials via the opaque-429 fallback.
The 达到…使用上限 arm requires the 使用 token, so a concurrency or rate cap
phrased as 达到…上限 (without 使用) stays in the upstream-backoff lane instead
of being misclassified as a credential-exhausting quota.
- deepseek-v4-flash bakes the wire-exact [low, high, max] ladder on
every host since 736b496cc6; V4 Pro stays [high, max].
- The stale xhigh alias-filter assertions now expect the flash ladder.
- Classified model_not_available_for_integrator as a permanent entitlement denial instead of transient fleet skew.
- Preserved the provider response and Available models list while retaining model_not_supported fleet retries.
Fixes#7819
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
403 concurrency caps bypass credential rotation in stream and auth-retry paths; snake-case concurrency codes classify; a session-level test covers the Copilot credential-removal gate.
(cherry picked from commit aca348e797ed987750a14ecf625952b6b971f7a5)
Reset-window rotation requires account-specific wording; concurrency caps require an actual cap signal; credential removal gated on AuthFailed without UsageLimit so a valid-but-blocked 403 credential is retained.
(cherry picked from commit 2f72752c2586352a4f7e9e814af1cdb0cf192af4)
Account-reset hint evaluated before short retry hints; account-scoped caps rotate on status 403 or undefined (Devin statusless trailer); statusless concurrency caps marked transient; transient same-model retries use the concurrency backoff.
Refuted: quota-worded concurrency caps were already excluded from rotation before the usage-limit text match.
(cherry picked from commit f2b9a18d715ddbcb6ae703670f2212da36bb2826)
The 3 -> 8 bump in callWithCopilotModelRetry was shared by the generic
retryable branch, so a persistent status-less transport blip on Copilot
would ramp across 8 attempts (~11.2s of dead time) instead of the 3 it
took before, and a repeated Retry-After 429 could stretch the same way
on top of the transport's own fetchWithRetry budget.
Derive the budget from the failure kind: model-availability 400s keep the
8-attempt reroll, everything else caps at the previous 3.
Also read COPILOT_TRANSIENT_MODEL_CODES with Object.hasOwn — `code` is
provider-controlled, so a 400 body whose code was `__proto__` or
`toString` classified as transient through the prototype chain.
Any Copilot model in the middle of a rollout (claude-sonnet-4.6,
claude-opus-4.6, gpt-5.4, gpt-5.3-codex, ...) returned a raw HTTP 400 on
roughly half of all turns. GET /models on api.githubcopilot.com returns
two different catalogs across repeated calls: part of the fleet serves
those ids, part rejects them with
400 {"error":{"message":"The requested model is not available for
integrator \"copilot-language-server\". ...",
"code":"model_not_available_for_integrator", ...}}
The absorb machinery already existed and was correct
(isCopilotTransientModelError -> isProviderRetryableError -> the
Anthropic transport's PROVIDER_MAX_RETRIES). Only the classifier missed:
it matched the older model_not_supported code and probed err.code /
err.error.code, while the real code is model_not_available_for_integrator
sitting at err.error.error.code (the SDK stores the parsed body on
.error, and Copilot's body is itself {error:{code}}). So
isProviderRetryableError fell through to "4xx => terminal".
Fix the classifier: providerErrorCode() walks the error envelope up to
depth 3 instead of hardcoding a shape, both Copilot model-availability
codes are accepted, and a wire-body text match backs it up because SDK
envelope shapes drift between provider families while the stringified
message does not. This alone restores the retry path, because
isProviderRetryableError consults the provider hook before its
4xx short-circuit.
Retry shape, since a rejection is a per-request replica reroll rather
than upstream backpressure:
- both transports wait a flat delay between model-flap attempts instead
of the growing backoff; generic retryable failures (429/5xx/transport)
keep their linear ramp and Retry-After handling
- the OpenAI-transport budget goes 3 -> 8 attempts, because a measured
~70% flap window produced a turn that needed 6 wire attempts and
exhausting the budget escalates to the agent-level retry, which
restarts the whole turn
Absorbed attempts cost no tokens: rejections are gateway-side, carry no
usage block, and return in ~208ms median versus ~1884ms for a served
request.
Also refresh the exhausted-retry guidance text, which cited a
nonexistent model id and described the cause as a per-client rollout gap
rather than fleet skew.
Treated an explicit allowed=true and limit_reached=false verdict as authoritative when Codex rounds usage to 100 percent. AuthStorage now honors normalized non-exhausted statuses before numeric fallbacks, keeping a usable Team credential ahead of rejected siblings.
Fixes#7617
Detected path-embedded OMP line selectors when reporting Cursor read results and stopped treating ranged payload lengths as whole-file totals.
Exposed exact source line counts from EOF-reaching read results and covered both the wire response and read metadata contracts.
Fixes#7590
Moved the direct DeepSeek enabled toggle into the thinking-only compat variant and normalized stale cached compat metadata before request encoding.
Fixes#7559
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
- Warned when a K3 turn completed without thinking blocks so the source turn is visible before degraded replay.
- Reported one-based assistant-turn positions on first degraded replay and deduplicated subsequent warnings for the same messages.
- Covered completion and replay warning behavior.
Fixes#7516
shouldSendServiceTier/applyOpenAIServiceTier began forwarding every tier —
including `auto` — for openai/openai-codex after PR #7376. Legacy/default
sessions resolve to {openai:"auto"}, so Codex (ChatGPT OAuth) requests now
carry service_tier:"auto", which that endpoint rejects with a 400, breaking
every turn at default settings.
Never send `auto`: it is OpenAI's implicit default, so omitting service_tier
is identical where accepted and required where the tier is rejected. Explicit
default/flex/scale/priority are unchanged.
Fixes#7517
assertCursorKimiK3HistoryReplayable threw for any assistant turn lacking a
non-empty thinking block, including same-model kimi-k3 turns whose stream
carried no thinkingDelta events. Such a turn is persisted with only
text/toolCall blocks, so every subsequent turn failed locally before any
request, permanently bricking the session.
Split the two failure modes: foreign history still hard-errors (another
model's turns cannot replay K3 reasoning), but a same-model turn missing
thinking now degrades to a one-time warning and replays without the
reasoning part via buildCursorAssistantContent, which already omits it.
Fixes#7516
- Add `renderToolExamplesJsdoc` to generate `@example` comment lines for tool inventories.
- Extract `bareStringArg` helper to simplify single-argument checks across renderers.
- Update tool inventory rendering to use the new JSDoc example format.