- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
- Added `leaked-thinking-stream.ts` whose `wrapLeakedThinkingStream`/`LeakedThinkingProjector` re-projects any provider stream, splitting leaked ```thinking`/`<think>` fences out of the visible-text channel into structured thinking blocks live while preserving text/thinking/tool signatures.
- Wrapped the shared `withProviderInFlightLimit` dispatch (both the unlimited and queued paths) and the standalone GitLab Duo path in `stream.ts` so healing covers every provider exit idempotently.
- Reworked `google-gemini-cli.ts` visible-text emission onto `StreamMarkupHealing` with `feedVisibleText`/`flushVisibleText` and explicit block bookkeeping, lifting leaked Gemini fences ahead of native tool calls.
- Added `leaked-thinking-stream` coverage and a gemini-cli healing case, and updated `stream-auth-retry`/`google-gemini-cli-alignment` expectations for the healed event sequence and dropped empty-text residue.
- Recorded the changelog `### Fixed` entry (whose run also carries the adjacent Codex `all_turns` line).
The daily Cloud Code Assist backend (`daily-cloudcode-pa`) exposes Claude 4.6
asymmetrically: `claude-sonnet-4-6` has no `-thinking` twin and
`claude-opus-4-6` has only the `-thinking` twin. The shared
`thinkingPair("claude-sonnet-4-6", …)` family (with `preserveAbsentEffortRoutes`)
kept every effort routed to `claude-sonnet-4-6-thinking` even when discovery
only returned the bare id, so any reasoning-on request 404'd with
`Requested entity was not found`. Claude on Antigravity also caps
`maxOutputTokens` at 64000, while `ANTIGRAVITY_MODEL_WIRE_PROFILES` had no
Claude entries — discovery's 65536 propagated to the wire and 400'd with
`Request contains an invalid argument`.
- Replaced the two `thinkingPair` calls for Claude 4.6 in `SHARED_CCA_FAMILIES`
with bespoke single-wire families. Sonnet collapses to the bare wire id,
Opus collapses to the `-thinking` wire id, and per-effort thinking is
carried by the request body's `thinkingBudget` on the single shared wire id.
Listing both candidate ids in `members` (priority order) keeps the collapse
correct if the backend mix ever rebalances.
- Added `claude-sonnet-4-6` and `claude-opus-4-6-thinking` entries to
`ANTIGRAVITY_MODEL_WIRE_PROFILES` capping `maxOutputTokens` at 64000.
- Made `AntigravityModelWireProfile.modelEnum` optional — Anthropic-backed
wire ids are accepted without a captured `labels.model_enum` token. The
request builder now emits the label only when the profile defines one.
- Regression tests in `variant-collapse.test.ts` (routing/wire-id resolution
for both 4.6 families across all three discovery permutations) and
`google-gemini-cli-alignment.test.ts` (request builder caps Claude
`maxOutputTokens` at 64000 and omits the unset `model_enum` label).
Fixes#3067
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Adjusted image budget IDs to use a shared 24-bit seed and wrap safely so new budgets no longer collide.
- Added transmit-state reset and changed image output ordering so full paints can replay image data after terminal clears.
- Updated image and protocol tests to match the new wrapped-id behavior and cursor-save/restore-wrapped transmit sequences.
- Removed ANTIGRAVITY_NO_PREAMBLE_INSTRUCTION from the shared Gemini header constants.
- Updated Gemini CLI and web search request builders to prepend only ANTIGRAVITY_SYSTEM_INSTRUCTION before existing system parts.
- Adjusted the Gemini CLI alignment test to stop asserting the removed anti-preamble prompt lines.
- Enabled Antigravity request envelopes to include requestId/labels and per-attempt responseId state.
- Enabled budget transport to send `thinkingBudget` and capped `maxOutputTokens` 65_536 for Gemini calls.
- Added budget-aware model profiles for Gemini flash/pro variants and updated effort routing expectations.
- Refreshed variant collapse to separate Gemini and Antigravity tables and heal stale snapshots.
Stopped google-gemini-cli and google-antigravity requests from sending explicit AUTO toolConfig entries for toolChoice auto, matching the shared Google provider behavior.
Added regression coverage for both Gemini CLI provider identities.
Fixes#2830
- agent-loop: raise repetition-detection floor to 180 chars and clear thinking
replay anchors when collapsing a detected loop.
- providers/google: ignore empty text parts, retain terminal thoughtSignatures,
and stop function-call signatures clobbering the prior block.
- autolearn: capture goal-mode at the turn boundary; harden managed-skill writes
against hard-links/symlinks (O_NOFOLLOW + nlink); refuse minting managed skills
whose name an authored skill already claims.
- eager tasks: thread agentKind through the session so a custom top-level agentId
still gets always-mode delegation; split Eager Tasks prompt into hard vs soft.
- title-generator: race the online title model against a local tiny-model fallback.
- eager-todo: keep the soft reminder aligned with the todo init schema.
- mcp/stdio: keep close() detaching the read loop instead of awaiting it.
- stream loop: fix collapsing and tool-call thought-signature handling.
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
- Derived descriptors, default-model map, env keys, login list, and refresh dispatch from one ProviderDefinition per provider.
- Disabled OpenAI Codex stream obfuscation and interrupted whitespace-only tool-call argument deltas.
- Derived auth-broker callback ports and paste-code login set from the registry.
- Centralized OAuth access lifecycle in `AuthStorage`, returning identity metadata and new access-result types.
- Added 60-second skew and strict expiry checks, returning undefined/throws for stale or expired OAuth credentials.
- Removed provider-local token refresh flows from Gemini, Gemini CLI, Antigravity, Kimi, and related OAuth helpers.
- Migrated web-search providers from `AgentStorage` to `AuthStorage` session-aware lookup with `authStorage`/`sessionId`/`signal` flow.
- Replaced `findAnthropicAuth`/DB auth lookup with `buildAnthropicAuthConfig` and explicit base-url override/env fallback ordering.
shouldInjectAntigravitySystemInstruction checked for the string
"gemini-3-pro-high" (hyphen) but the deployed model IDs are
"gemini-3.1-pro-high" and "gemini-3.1-pro-low" (dot-versioned).
String.includes() never matched, so the required identity header was
silently omitted and Cloud Code Assist returned HTTP 400 for every
request targeting these models.
Broadened the guard to match.includes("gemini-3"), consistent with
the existing isGemini3Model convention in google-shared.ts, covering
all current and future Gemini 3.x variants on the antigravity provider.
Fixes#1274
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
- Removed local `abortableSleep` in favour of Node's built-in `scheduler.wait` from `node:timers/promises`.
- Consolidated per-provider retry/fetch loops into a shared `fetchWithRetry` utility in `packages/utils`.
- Moved `extractHttpStatusFromError`, `isRetryableError`, and related helpers out of `packages/ai` into `packages/utils`.
- Deleted `extractRetryDelay` in favour of `extractRetryHint` with unified header and body parsing.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Extracted fetch mocking logic into reusable `hookFetch()` utility function with middleware-style handler pattern.
- Replaced manual `globalThis.fetch` assignment and restoration across 10 test files with `hookFetch()` calls using `using` statement for automatic cleanup.
- Implemented Disposable pattern with Symbol.dispose for fetch hook resource management, eliminating try-finally blocks.
- Exported `hookFetch` from utils public API to enable consistent fetch mocking across packages.
- Fixed schema compatibility issue by sanitizing patternProperties in tool parameters during Antigravity format conversion.
- Added sanitizeSchemaForCCA utility to remove unsupported schema properties when rewriting tools to legacy parameters format.
- Added test case verifying patternProperties removal during tool parameter conversion.
- Added PEP 563 forward reference support to Python prelude for improved type annotation compatibility.
- Added exponential backoff retry mechanism with configurable max retries for transient failures. Implemented rate limit budget tracking to prevent excessive delays on 429 responses. Fixed HTTP status propagation to prevent network errors from being retried when explicit HTTP failures occur.
providers/google-gemini-cli:
- parseGeminiCliCredentials() handles legacy, alias (project_id/refresh/expires),
and enriched credential JSON formats
- shouldRefreshGeminiCliCredentials() + refreshGeminiCliCredentialsIfNeeded()
proactively refresh OAuth tokens 60s before expiry for both providers
- normalizeAntigravityTools() converts parametersJsonSchema -> parameters
in function declarations for Antigravity compatibility
- VALIDATED tool calling config applied for Antigravity + Claude model combos
- maxOutputTokens removed from generation config for Antigravity non-Claude models
- Antigravity system instruction injection scoped to Claude + gemini-3-pro-high models
- Antigravity session ID: signed decimal int63 derived from SHA-256 of first user
message (or random bounded int63), replacing truncated hex hash
- Antigravity requestId uses agent-{uuid}; non-Antigravity requests omit
requestId/userAgent/requestType from payload
- ANTIGRAVITY_DAILY_ENDPOINT corrected to daily-cloudcode-pa.googleapis.com;
sandbox kept as fallback
- ANTIGRAVITY_SYSTEM_INSTRUCTION exported
oauth/google-antigravity:
- PKCE removed from OAuth flow (no code_challenge)
- loadCodeAssist metadata ideType changed to ANTIGRAVITY
- discoverProject uses single production endpoint; falls back to onboardUser LRO
(up to 5 retries, 2s interval) instead of hardcoded default project ID
- ANTIGRAVITY_LOAD_CODE_ASSIST_METADATA exported
oauth/google-gemini-cli:
- PKCE removed from OAuth flow
oauth/index:
- getOAuthApiKey includes refreshToken, expiresAt, email, accountId in
Gemini/Antigravity JSON payload for proactive refresh support
discovery/antigravity:
- Tries production daily endpoint first, sandbox as fallback
- Removed recommended/agentModelSorts filter; applies denylist instead
- ANTIGRAVITY_DISCOVERY_DENYLIST filters low-quality/internal models
- Request body no longer includes project field
coding-agent:
- gemini_image: corrected responseModalities to uppercase IMAGE/TEXT
- Gemini web search: endpoint fallback (daily->sandbox) with retry on 429/5xx,
aligned Antigravity request metadata, ANTIGRAVITY_SYSTEM_INSTRUCTION injection
- buildGeminiRequestTools() helper for composable googleSearch/codeExecution/urlContext
- Web search schema: expose max_tokens, temperature, num_search_results params
- Web search: explicit provider falls back to auto chain when unavailable
Tests: google-antigravity-auth, google-gemini-cli-alignment, web-search-gemini