- Removed the Google Interactions transport and deleted interaction-specific request options from the shared AI stream typing/API surface.
- Simplified Google provider routing to eliminate interactions auto-selection logic and keep `streamGoogle` on the `:streamGenerateContent` path.
- Updated Vertex request handling to use resolved stream hosts without `/interactions`/`Api-Revision` and removed related interaction constants.
- Deleted obsolete Interactions tests and updated remaining Google stream tests to no longer reference `useInteractionsApi`/`storeInteraction`/`previousInteractionId`.
- Added `google-interactions.ts` streaming the Gemini Interactions API (model-mode steps, thought/tool/text deltas, usage, `previous_interaction_id` lineage) shared by the direct Google and Vertex providers.
- Routed `streamGoogle` and `streamGoogleVertex` through `resolveInteractionDispatch`, defaulting Gemini 3+ onto Interactions (official endpoint for direct Google, bearer/ADC for Vertex) with transparent `:streamGenerateContent` fallback and a `useInteractionsApi: false` opt-out.
- Added explicit Vertex bearer support in `google-auth.ts` via `GOOGLE_CLOUD_ACCESS_TOKEN`/`CLOUDSDK_AUTH_ACCESS_TOKEN` plus a `hasVertexBearerCredentialsHint` probe gating the auto-default.
- Added `useInteractionsApi`/`storeInteraction`/`previousInteractionId` to `StreamOptions` and `GoogleSharedStreamOptions`, threading them through `mapOptionsForApi`.
- Added `google-interactions` coverage and pinned the existing generateContent assertions with `useInteractionsApi: false`.
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
- Removed deprecated AWS/Google client and proxy dependencies from root and AI package manifests.
- Added AWS credential chaining, SigV4 signing, and Bedrock stream decoding with region fallback and CRC checks.
- Added local Google type mirrors and migrated providers to request plans with SSE fetch and token-based auth.
- Updated AI changelog with auth/stream fixes, and added fetch-stub and SigV4/event-stream tests.
- Introduced `FetchImpl` type with optional `preconnect` to accept non-Bun fetch implementations without type errors.
- Applied the new type across all providers and `StreamOptions.fetch`.
- Added tests verifying fetch override routing for openai-completions, openai-responses, and fetchWithRetry.
- Introduced a `fetch` option on `StreamOptions` and threaded it through providers to let callers supply a custom request transport.
- Updated provider clients and direct HTTP calls across Anthropic, OpenAI, Azure, Google, GitLab Duo, Gemini CLI, Ollama, and Codex flows to use the injected fetch implementation.
- Extended retry helper options to accept a fetch override and preserved preconnect support from the selected fetch function.
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
- Anthropic streaming now populated output.errorMessage from refusal stop_details, including category and explanation when available.
- Stream handlers for Bedrock, Azure OpenAI, Google, Google Vertex, Google Gemini CLI, and OpenAI now threw output.errorMessage when stopReason indicated aborted or error instead of a generic unknown message.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Extended `Usage` typing with `reasoningTokens`, `cttl`, and `server` fields for richer token accounting.
- Added conditional `reasoningTokens` output across OpenAI and Google providers when token counts are positive.
- Updated `parseChunkUsage` and OpenAI usage parsing to prevent `reasoning_tokens` double-counting.
- Added Anthropic usage extras mapping to emit TTL and server tool counters only when non-zero.
- Preserved Anthropic cache TTL and server counters when later usage events omit those fields.
- Added usage-attribution tests for parse helpers, missing fields, and zero-value merge scenarios.
When no GEMINI_API_KEY is set, the fallback `|| ""` passes an empty
string to `new GoogleGenAI({ apiKey: "" })`. The SDK treats this as a
user-provided API key and warns that it will take precedence over
Vertex AI project/location auth, breaking Vertex users.
Pass `undefined` instead so the SDK falls through to Vertex auth.
Co-authored-by: Muness Castle <munesscastle@artium.ai>
The context fullness gauge was driven by output token count, causing
erratic jumps between turns (e.g. 84% -> 64%) with no compaction.
Status bar and estimateContextTokens now use calculatePromptTokens()
which returns input + cacheRead + cacheWrite — the actual input context
size. Previously both used a formula that included the final output token
count, which fluctuates with response length and is not part of the
context window for the current request.
isContextOverflow's usage-based fallback (z.ai silent overflow) was
also missing cacheWrite (cache_creation_input_tokens). Per Anthropic
docs the threshold is input + cache_read + cache_creation — all three.
Ref: https://platform.claude.com/docs/en/about-claude/pricing#long-context-pricing
google.ts and google-vertex.ts were double-counting cached tokens.
Gemini's promptTokenCount already includes cachedContentTokenCount, so
assigning input = promptTokenCount and cacheRead = cachedContentTokenCount
overcounted by cachedContentTokenCount on every cached request. Fixed
by subtracting first, matching the OpenAI convention:
input = promptTokenCount - cachedContentTokenCount
cacheRead = cachedContentTokenCount
=> input + cacheRead = promptTokenCount (total prompt, no double-count)
Ref: https://ai.google.dev/api/generate-content#v1beta.GenerateContentResponse.UsageMetadata
All other providers validated: amazon-bedrock (inputTokens is uncached
by API contract), openai-completions/responses/azure (already subtract
cached), kimi/gitlab-duo (delegate to correct implementations), cursor
(API exposes output tokens only — input stays 0 by design).
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Replaced custom `sanitizeSurrogates()` utility with native `String.prototype.toWellFormed()` across all provider implementations.
- Removed `sanitize-unicode.ts` utility module and all imports of `sanitizeSurrogates` function from 10 provider files.
- Updated Unicode surrogate handling in message conversion, response processing, and system prompt handling to use built-in JavaScript API.
- Documented removal of `sanitizeSurrogates()` utility and migration to native `toWellFormed()` in CHANGELOG.
- Extracted schema utilities from typebox-helpers and google-shared into new modular utils/schema package with 17 exported functions.
- Consolidated OpenAI strict mode schema enforcement across codex, completions, and responses providers using unified adaptSchemaForStrict() helper.
- Refactored credential ranking from hardcoded Codex-specific logic to pluggable CredentialRankingStrategy pattern with provider implementations.
- Migrated 500+ lines of Google schema sanitization and normalization logic from google-shared.ts to dedicated utils/schema modules with expanded functionality.
- Added raw HTTP request dumps to error messages for 400 status codes.
- Implemented client-side retry for Anthropic streaming on transient errors.
- Preserved context by converting aborted assistant tool calls to text.
- Optimized model registry refresh in coding agent sessions.
- Removed Prettier configuration files (.prettierignore and .prettierrc) and migrated formatting to Biome.
- Updated Biome configuration from version 2.3.11 to 2.3.12 and changed arrowParentheses rule from 'always' to 'asNeeded'.
- Pinned @biomejs/biome dependency to exact version 2.3.12 in package.json and bun.lock.
- Applied consistent arrow function formatting across 489 files by removing unnecessary parentheses around single parameters.
- Removed blank lines after comment blocks and reorganized imports for consistency across the codebase.
- Removed WASM generation script; use Bun `wasm?raw` loader for imports.
- Added bunfig.toml with loaders for `.md`, `.py`, and `.wasm?raw` text imports.
- Added types/assets/index.d.ts for global TypeScript module declarations.
- Unified TypeScript configuration with tsgo-based checking across monorepo.
- Removed build and WASM steps from install and publish pipelines.
- Added tsconfig.publish.json files to all packages with optimized publish-time configuration.
- Updated all package.json scripts with prepublishOnly hooks for correct type checking during publish.
- Added @oh-my-pi/omp-stats path mappings to root tsconfig.json for consistent imports.
- Added WASM generation script for photon module and integrated into install:dev script.
- converted relative imports to path aliases ($c/*, $ai/*, $tui/*, etc.) across all packages
- added per-package tsconfig.json with complete path mappings for runtime resolution
- set importModuleSpecifier to non-relative for IDE auto-import preferences
- updated dev script to run from monorepo root for consistent path resolution
- Added duration and TTFT (time to first token) metrics to AssistantMessage interface.
- Implemented performance tracking across all AI providers for streaming responses.
- Added timing capture in both success and error code paths for comprehensive metrics.
- Updated changelog to document new performance tracking capabilities.
- Added headers option to all providers for custom request headers.
- Added onPayload hook to observe provider request payloads before sending.
- Added strictResponsesPairing option for Azure OpenAI Responses API compatibility.
- Added originator option to loginOpenAICodex for custom OAuth flow identification.
- Enhanced AWS credential detection to support ECS task roles and IRSA web identity tokens.
- Fixed tool parameter schema sanitization to only apply Google-specific transformations for Gemini models.
- Added model parameter to convertTools function to enable conditional schema sanitization.
- Preserved original tool schemas for non-Gemini models running through Google providers.
- Exported sanitizeSchemaForGoogle utility function for public use.
- Added filtering of unsupported JSON Schema fields ($schema, $ref, $defs, format, examples, etc.) in Google provider schema sanitization.
- Added handling to ignore unsupported additionalProperties: false in Google provider.
- Removed format: date-time from timestamp type conversion in JTD to JSON Schema transformation.
- Added schema validation tests to verify tool schemas don't contain problematic JSON Schema features.
- Reorganized system prompt to display context, environment, and tools sections before discipline guidelines.
- Added retry logic with exponential backoff and model fallback for auto-compaction failures.
- Added support for pi/<role> model aliases and automatic model inheritance for subtasks.
- Enhanced error messages with retry-after timing from rate limit headers across all providers.
- Fixed image attachments being dropped when steering messages are queued during streaming.
- Changed edit tool to merge call and result displays into single block.
- Changed model override behavior to persist in settings when explicitly set via CLI.
- Port pi-ai from upstream with Bun-first approach (source files, no dist)
- Add cli_ prefix to tool names in OAuth mode to avoid collisions
- Improve beta header handling with proper deduplication
- Convert codex-instructions.md to Bun text embed
- Strip .js extensions from all imports
- Add GoogleThinkingLevel type mirroring Google's ThinkingLevel enum
- Update GoogleGeminiCliOptions and GoogleOptions to use our type
- Cast to any when assigning to Google SDK's ThinkingConfig
- Add new API type 'google-cloud-code-assist' for Gemini CLI / Antigravity auth
- Extract shared Google utilities to google-shared.ts
- Implement streaming provider for Cloud Code Assist endpoint
- Add 7 models: gemini-3-pro-high/low, gemini-3-flash, claude-sonnet/opus, gpt-oss
Models use OAuth authentication and have sh cost (uses Google account quota).
OAuth flow will be implemented in coding-agent in a follow-up.
Previously, when using 'google-generative-ai' API with a custom baseUrl
in models.json, the baseUrl was ignored and requests always went to the
default Google endpoint.
Now the provider correctly passes model.baseUrl to the SDK's
httpOptions.baseUrl, enabling use of custom endpoints or API proxies.
Fixes#216
- Fix tool result format for Gemini 3 Flash Preview compatibility
- Use 'output' key for successful results (not 'result')
- Use 'error' key for error results (not 'isError')
- Per Google SDK documentation for FunctionResponse.response
- Improve type safety in google.ts provider
- Add ImageContent import and use proper type guards
- Replace 'as any' casts with proper typing
- Import and use Schema type for tool parameters
- Add proper typing for index deletion in error handler
- Add comprehensive test for Gemini 3 Flash tool calling
- Tests successful tool call and result handling
- Tests error tool result handling
- Verifies fix for issue #213Fixes#213
* use the correct Gemini 3 Flash Preview thinking levels
* fix a build error
* add changelog entry
* regenerate models
* make less assumptions about future models
- Added totalTokens field to Usage interface in pi-ai
- Anthropic: computed as input + output + cacheRead + cacheWrite
- OpenAI/Google: uses native total_tokens/totalTokenCount
- Fixed openai-completions to compute totalTokens when reasoning tokens present
- Updated calculateContextTokens() to use totalTokens field
- Added comprehensive test covering 13 providers
fixes#130
Fixes#39
- Added headers field to Model type (provider and model level)
- Model headers override provider headers when merged
- Supported in all APIs:
- Anthropic: defaultHeaders
- OpenAI (completions/responses): defaultHeaders
- Google: httpOptions.headers
- Enables bypassing Cloudflare bot detection for proxied endpoints
- Updated documentation with examples
Also fixed:
- Mistral/Chutes syntax error (iif -> if)
- process.env.ANTHROPIC_API_KEY bug (use delete instead of = undefined)
- Add NO_IMAGE to error finish reasons in Google provider
- Fix non-null assertion after optional chaining in Anthropic provider
- Migrate biome config to 2.3.5
- Ignore Tailwind CSS file from biome checks
- Bump all packages to version 0.6.0
Tool results now use content blocks and can include both text and images.
All providers (Anthropic, Google, OpenAI Completions, OpenAI Responses)
correctly pass images from tool results to LLMs.
- Update ToolResultMessage type to use content blocks
- Add placeholder text for image-only tool results in Google/Anthropic
- OpenAI providers send tool result + follow-up user message with images
- Fix Anthropic JSON parsing for empty tool arguments
- Add comprehensive tests for image-only and text+image tool results
- Update README with tool result content blocks API