Replaced Google Vertex project discovery with the models.dev catalog so bundled model selection includes current Vertex MaaS and Gemini entries while pruning retired fallbacks.
Fixes#1456
xai-oauth has no upstream catalog source (not in models.dev or
MODELS_DEV_PROVIDER_DESCRIPTORS) and dynamic discovery only fires after
the user authenticates and refresh() runs. Boot resolves the persisted
modelRoles.default synchronously from #loadModels(), which reads only
models.json — so a persisted "xai-oauth/<id>" default silently resets on
fresh starts because the bundle was empty.
Add XAI_OAUTH_CURATED_MODELS as the single source of truth for the
xai-oauth chat picker (grok-build, grok-4.3, grok-4.20-multi-agent-0309,
grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning) plus
buildXaiOAuthStaticSeed to render them as Model<"openai-responses">.
Wire the seed three ways:
- generate-models.ts pushes the seed into models.json so #loadModels()
sees it at boot.
- xaiOAuthModelManagerOptions hands the seed as staticModels so the
picker shows the catalog before fetchDynamicModels fires.
- test/xai-oauth-bundle.test.ts pins the bundle ⇔ seed invariant so
editing one without regenerating the other fails CI.
Op: correct
Restores: ref:feat/xai-grok-oauth@0f86695c1
Restores: ref:feat/xai-grok-oauth@1097373bc
Restores: ref:feat/xai-grok-oauth@eab1ed721
Centralizing OAuth refresh in AuthStorage (e6893515) introduced five
follow-on bugs surfaced by an audit of the commit; this fixes all of
them and updates the tests that relied on the old refresh seam.
1. packages/ai/src/auth-storage.ts (#tryOAuthCredential):
For built-in providers the path went directly to `getOAuthApiKey`
with the (possibly still-expired) selection.credential when the
pre-refresh at line 2587 caught a transient error. `getOAuthApiKey`
then threw the "expired … must be refreshed via AuthStorage"
precondition error, which the disable classifier matched against
`/expired.*refresh/` and soft-disabled the row. A single network
blip during refresh could permanently kill a still-valid Anthropic /
OpenAI / Gemini-CLI / Copilot credential. Built-in providers now
route through the broker-aware single-flighted
`#refreshOAuthCredential` first, so transient failures surface as
network errors (5-min temp block) instead of definitive auth
failures.
2. packages/ai/src/auth-storage.ts (#fetchUsageUncached):
The usage refresh check only fired once `Date.now() >= expiresAt`,
missing the 60-second skew that `getApiKey` honors. A token
expiring inside the skew window was posted to the usage endpoint
and 401'd mid-flight, briefly hiding quota in the UI. Aligned with
`OAUTH_REFRESH_SKEW_MS`.
3. packages/coding-agent/src/web/search/index.ts (webSearchCustomTool):
The CustomTool counterpart of WebSearchTool dropped sessionId so
SDK callers that opted into `web_search` via toolNames lost
per-session credential stickiness — multi-account users saw the
provider round-robin between searches in the same session. Threads
`ctx.sessionManager.getSessionId()` through to `executeSearch`.
4. packages/coding-agent/src/web/search/providers/perplexity.ts
(findOAuthToken):
`authStorage.getApiKey("perplexity")` returns runtime/config
overrides, stored api_key credentials, OAuth bearers, and env keys.
Filtering only env keys meant a config-pinned `pplx-…` API key was
POSTed to `www.perplexity.ai/rest/sse/perplexity_ask` (the OAuth
endpoint) instead of falling through to
`api.perplexity.ai/chat/completions`, producing 401s. Switched to
`getOAuthAccess` so only true OAuth bearers reach the OAuth
branch; api_key credentials/overrides correctly fall through.
5. packages/ai/scripts/generate-models.ts:
`getOAuthApiKey` was being called directly with possibly-expired
credentials. The new contract throws on expired, the broad catch
swallowed it, and the build silently fell back to bundled models
instead of refreshing. Both helpers now route through
AuthStorage's `getApiKey` / `getOAuthAccess`, which trigger the
full broker-aware refresh pipeline.
Test updates:
- auth-storage-credential-disabled-event.test.ts,
sdk-credential-disabled-bridge.test.ts: the `failOAuthRefresh`
helper used to spy on `getOAuthApiKey` to inject invalid_grant.
With refresh now happening before that helper, the spy never fired.
Switched to spying on `refreshOAuthToken` so the simulated failure
reaches the disable classifier.
- auth-storage-rotation.test.ts: stub `refreshOAuthToken` so the test
doesn't hit a real OAuth endpoint when the seeded credential lands
inside the 60s skew window.
- Added `AuthBrokerClient`, `RemoteAuthCredentialStore`, `AuthBrokerRefresher`, and `startAuthBroker` server in `packages/ai/src/auth-broker`.
- Renamed `AuthCredentialStore` class to `SqliteAuthCredentialStore`; extracted `AuthCredentialStore` as a persistence interface.
- Added `exportSnapshot`, `forceRefreshCredentialById`, `disableCredentialById`, and `upsertCredential` to `AuthStorage` for broker wire protocol.
- Added `omp auth-broker` CLI subcommand (serve, token, login, logout, import, status) and `discoverAuthStorage` broker-mode path keyed on `OMP_AUTH_BROKER_URL`.
- Updated Codex model pricing data to include non-zero input/output/cacheRead rates.
- Added generate-models fallback to copy billable OpenAI costs into bundled openai-codex models.
- Added catalog-cost helpers to fallback openai-codex pricing to openai and compute token totals.
- Added regression tests for openai-codex cost parity and db backfill, and documented the fix in changelog.
- Added Fireworks provider onboarding with API-key login and credential storage.
- Added Fireworks provider registrations across stream, types, descriptors, and OAuth registries.
- Mapped Fireworks public model IDs to wire IDs for OpenAI-completion requests.
- Extended OpenAI-compatible detection to treat api.fireworks.ai as supported for streaming and max_tokens.
- Updated model catalog handling with Fireworks models plus context-window/max-token inheritance and metadata updates.
- Added Fireworks model reference loading and fallback logic to preserve best context and token caps.
- Updated OpenAI context-promotion linking to handle `-spark` variants and base `gpt-5.5` models.
- Adjusted model generation data so OpenAI-like models now promote first to `gpt-5.5` and then to `gpt-5.4` where configured.
- Updated unit tests to assert context-promotion targets for both spark and `gpt-5.5` model variants.
Copilot's /models endpoint exposes token limits under capabilities.limits.
max_prompt_tokens is the actual prompt capacity (what OMP calls contextWindow),
while max_context_window_tokens is the total window (prompt + output budget).
Using the latter inflates contextWindow, which breaks compaction thresholds,
overflow detection, and context promotion.
Three changes:
1. mapModel reference selection: always prefer the Copilot-specific bundled
reference over the global cross-provider reference. Copilot imposes its
own limits that are strictly lower than native provider limits.
2. mapModel contextWindow chain: remove max_context_window_tokens from the
fallback. New chain: context_length -> max_prompt_tokens -> reference.
3. generate-models: stop overwriting contextWindow/maxTokens in
applyGlobalModelsDevFallback. These are provider-specific and should not
be replaced with cross-provider models.dev global references.
Also fixes bundled values: github-copilot/gpt-5.4 (400k -> 272k) and
github-copilot/gpt-5.2 (264k -> 128k) to match live API max_prompt_tokens.
Refs: #225, #226
- Fixed thinking configuration format by replacing `levels` array with `minLevel`/`maxLevel` properties across 100+ model definitions.
- Corrected GPT-5.4 mini/nano context window from 400000 to 272000 tokens for accurate token limit reporting.
- Normalized GPT-5.4 variant priority handling to use parsed variant instead of raw model IDs for consistent behavior.
- Added "mini" variant support to OpenAI model parsing regex and updated thinking mode configuration for Claude models.
- Fixed test robustness by replacing exact string matching with numeric range comparison to handle BSD seq notation on macOS.
- Corrected model generation script execution order to apply policy overrides before promotion target linking.
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Updated dev dependencies including biome, TypeScript native preview, and lint-staged to latest versions.
- Updated AI package dependencies for Anthropic SDK, AWS Bedrock, and Zod with relaxed version constraints.
- Updated models.json with new model entries, removed premiumMultiplier fields from GitHub Copilot models, and corrected maxTokens and cost values for various providers.
- Added appearance and projfs export paths to natives package and updated peer dependency version for swarm-extension.
- Extracted credential storage to shared @oh-my-pi/pi-ai package with AuthCredentialStore and AuthStorage classes.
- Consolidated UI formatting logic from ToolUIKit class into standalone utility functions across render-utils and output-meta modules.
- Moved utility functions (parseCommandArgs, substituteArgs, expandPath, normalizeUnicode) to dedicated modules for improved code reuse.
- Extracted JTD type definitions and type guards to jtd-utils module for shared use across schema conversion tools.
- Updated Claude model pricing and added cache read costs in models.json for accurate billing calculations.
- Refactored agent-storage to delegate credential management to AuthCredentialStore instead of direct SQLite operations.
- Added GitLab Duo provider with support for Claude, GPT-5, and Duo Chat models via GitLab AI Gateway.
- Added OAuth authentication for GitLab Duo with automatic token refresh, PKCE security, and 25-minute token caching.
- Added 16 new GitLab Duo models including Claude Opus/Sonnet/Haiku and GPT-5 variants with reasoning and multimodal support.
- Added `isOAuth` option to Anthropic provider for OAuth bearer token authentication mode.
- Exported `streamGitLabDuo`, `getGitLabDuoModels`, and `clearGitLabDuoDirectAccessCache` functions for GitLab Duo integration.
- Added configurable model refresh strategies (online, online-if-uncached) to control model discovery behavior.
- Implemented global model.dev fallback resolution with context window and token prioritization for improved model attribute consistency.
- Incremented model cache schema version to support improved global model fallback resolution.
- Enhanced model generation script with dedicated functions for building model reference maps and applying fallback attributes.
- Extracted provider descriptor helper functions (descriptor, catalog, catalogDescriptor, simpleModelsDevDescriptor, openAiCompletionsDescriptor, anthropicMessagesDescriptor) to reduce boilerplate across 47+ provider definitions.
- Refactored model generation script to simplify OAuth credential lifecycle by moving storage cleanup to finally block and parallelizing special discovery sources with Promise.all().
- Consolidated provider validation logic in model registry into reusable validateProviderConfiguration() function with context-aware validation modes.
- Reorganized models.json structure to place contextWindow and maxTokens before compat field for consistency across all provider entries.
- Unified provider descriptors into single source of truth in descriptors.ts module for runtime and catalog discovery.
- Consolidated model generation script to use declarative CatalogProviderDescriptor interface, reducing code duplication.
- Refactored model registry and selector to use descriptor pattern with priority-based sorting and version extraction.
- Added priority field to Model interface enabling provider-assigned model prioritization in discovery and selection.
- Added support for Synthetic model provider in web search command and improved model sorting by priority and version.
- Extracted model post-processing policies into dedicated model-policies module for improved testability and maintainability.
- Refactored model generation script to use declarative provider descriptors instead of 620+ lines of inline provider-specific logic.
- Consolidated provider model manager initialization to use descriptor-driven iteration, eliminating 29+ individual conditional blocks.
- Removed static bundled models for Ollama and vLLM from models.json to rely on dynamic discovery instead.
- Added support for 11 new AI providers (Hugging Face, NVIDIA, Together, Ollama, LiteLLM, Xiaomi, Moonshot, Venice, Qwen Portal, vLLM, Cloudflare AI Gateway) with API key authentication and login flows.
- Implemented $pickenv() utility for environment variable fallback chains, enabling multi-key resolution for providers with alternative credential names.
- Extended KnownProvider and OAuthProvider types to include all 11 new providers with corresponding model manager functions and OAuth handlers.
- Expanded models.json with thousands of new model entries across all new providers and replaced deprecated opencode provider with cloudflare-ai-gateway.
- Refactored model generation script to use unified fetchProviderModelsFromCatalog() and centralized API key resolution for all providers.
- Added ModelManager API with createModelManager() factory for managing bundled and dynamically discovered models with configurable refresh strategies.
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints with provider-specific model manager configuration helpers.
- Renamed public API functions for clarity: getModel() -> getBundledModel(), getModels() -> getBundledModels(), getProviders() -> getBundledProviders().
- Added on-disk model caching with TTL-based invalidation and resolveProviderModels() function for runtime model resolution with source precedence.
- Refactored model discovery script to dynamically fetch models from Codex, Cursor, and Antigravity using OAuth credentials instead of hardcoded lists.
- Added Claude Sonnet 4.6 and Claude Sonnet 4.6 Thinking models to Antigravity provider.
- Added GLM-5 Free model via OpenCode provider.
- Added GLM-4.7-FlashX model via ZAI provider.
- Added MiniMax-M2.5-highspeed model across four providers (minimax-code, minimax-code-cn, minimax, minimax-cn).
- Added Claude Sonnet 4.6 model to OpenRouter and Vercel AI Gateway providers.
- Updated pricing and token limits for deepseek-v3, mistral-large-2411, and Qwen models across OpenRouter and Together AI providers.
- Expanded EU cross-region inference variant support to all Claude models on Bedrock (previously limited to Haiku, Sonnet, and Opus 4.5).
- Updated context window handling for Sonnet 4.6 models to enforce 200K limit across all providers.
- Added contextPromotionTarget model property to specify preferred fallback model when context promotion is triggered.
- Added automatic context promotion target assignment for Spark models to their base model equivalents.
- Updated Qwen model context window and max token limits for improved accuracy.
- Updated o1 model context window from 256000 to 262144 tokens and max tokens from 64000 to 65536 tokens.
- Implemented context promotion logic to use configured contextPromotionTarget when available instead of role-based model resolution.
- Added DeepSeek-V3.2 model support via Amazon Bedrock.
- Added GLM-5 model support via OpenCode.
- Added MiniMax M2.5 model support via OpenCode.
- Updated GLM models to use anthropic-messages API instead of openai-completions and changed base URL from https://api.z.ai/api/coding/paas/v4 to https://api.z.ai/api/anthropic.
- Removed compat field with supportsDeveloperRole and thinkingFormat properties from GLM models.
- Updated pricing and context window specifications for multiple models including Mistral, Moonshot, and Qwen variants.
- Added sorting of models by ID across multiple model fetch functions (OpenRouter, AI Gateway, Kimi Code, and dev data) using localeCompare for consistent ordering.
- Added WebSocket transport support for OpenAI Codex responses with automatic fallback to SSE on connection failure.
- Added preferWebsockets option to Agent and Model configurations to hint that WebSocket transport should be preferred when supported by provider implementations.
- Added prewarmOpenAICodexResponses() function to pre-establish WebSocket connections for improved performance.
- Added getProviderDetails() function and getOpenAICodexTransportDetails() function to expose transport state and provider configuration information.
- Added provider details display in session info showing active provider configuration and authentication details.
- Added OpenAI websockets setting to enable WebSocket transport preference for OpenAI Codex models in coding agent configuration.
- Added GPT-5.3 Codex Spark model with 128K context window and extended reasoning capabilities.
- Added MiniMax M2.5 and M2.5 Lightning models via OpenAI-compatible API (minimax-code and minimax-code-cn providers).
- Added MiniMax M2.5 and M2.5 Lightning models via Anthropic API (minimax and minimax-cn providers).
- Added Llama 3.1 8B model via Cerebras API.
- Added MiniMax M2.5 model via OpenRouter and Vercel AI Gateway.
- Added Qwen3 VL 32B Instruct multimodal model via OpenRouter.
- show help instead of crashing on `omp setup` with no args
- show runtime-discovered MCP servers in `/mcp list`
- remove deprecated Anthropic model entries from models.json
- sort models by recency in model selector
- Sorted Antigravity models alphabetically in generated models.json and generation script.
- Updated Antigravity model configurations with corrected model IDs, names, and parameters.
- Fixed qwen/qwen3-max-thinking maxTokens from 4096 to 202752 to match context window.
Add support for MiniMax Coding Plan with OpenAI-compatible API:
- New providers: minimax-code (international) and minimax-code-cn (China)
- Environment variables: MINIMAX_CODE_API_KEY and MINIMAX_CODE_CN_API_KEY
- Uses thinkingFormat: 'zai' for reasoning compatibility
- Models: MiniMax-M2, MiniMax-M2.1, MiniMax-M2.1-lightning
- Base URLs: https://api.minimax.io/v1 (intl), https://api.minimaxi.com/v1 (CN)
The Coding Plan is a subscription-based service separate from regular MiniMax API.
- Migrated models export from TypeScript module to JSON format, changing the public API from importing MODELS from './models.generated' to importing from './models.json' with JSON import assertion.
- Updated @anthropic-ai/sdk dependency from ^0.72.1 to ^0.74.0.
- Simplified model generation script by replacing 49 lines of TypeScript code generation with direct JSON serialization.
- Updated @types/bun devDependency from ^1.3.8 to ^1.3.9 across all packages.
- Removed models.generated.ts exclusion from biome.json linting configuration.
- Added support for Kimi K2, K2 Turbo Preview, and K2.5 models with reasoning capabilities.
- Fixed Claude Opus 4.6 context window to 200K across all providers (was incorrectly set to 1M).
- Fixed Claude Sonnet 4 context window to 200K across multiple providers (was incorrectly set to 1M).
- Improved Kimi Code model fetching to merge fallback models not returned by the API endpoint.
- Added dynamic Antigravity model fetching from API when credentials are available, with hardcoded fallback models for offline use.
- Updated Antigravity models to use free tier pricing (0 cost) across all models.
- Extracted Antigravity model loading into separate functions (fetchAntigravityModels, getAntigravityFallbackModels, getAntigravityToken) for better maintainability.
- Added support for fetching recommended models from Antigravity API response and filtering internal models.
- Added Claude Opus 4.6 Thinking model for Antigravity provider.
- Added Gemini 2.5 Flash, Gemini 2.5 Flash Thinking, and Gemini 2.5 Pro models for Antigravity provider.
- Added Pony Alpha model via OpenRouter.
- Updated Claude Opus 4.6 context window from 200,000 to 1,000,000 tokens across Bedrock regions.
- Updated Claude Opus 4.6 cache pricing and Antigravity model pricing to free tier across multiple models.
- Fixed Claude Opus 4.6 model ID format by removing version suffix (:0) in Bedrock configurations.
- Extracted stream parsing utilities into reusable functions in pi-utils package (readLines, readJsonl, parseJsonlLenient, readSseJson).
- Replaced manual buffer management with ConcatSink class for efficient stream handling across multiple modules.
- Simplified SSE stream parsing in google-gemini-cli provider by using readSseJson utility instead of manual reader setup.
- Refactored MCP stdio transport to use readLines utility for cleaner line-based stream processing.
- Updated RPC client and mode to use readJsonl utility, removing duplicate JSONL parsing logic.
- Optimized buffer allocations by using allocUnsafe where data is immediately populated and reusing empty buffer instances.
- Migrated environment variable access from function-based `getEnv()` API to object property-based `$env` API across 43 files in the monorepo.
- Simplified `packages/utils/src/env.ts` by removing scoped environment variable logic (EnvScope, getEnvMap, getEnv functions) and replacing with direct `process.env` reference exported as `$env`.
- Updated `packages/ai/src/stream.ts` to remove environment parameter from KeyResolver functions and use global `$env` object instead of passed-in env parameter.
- Migrated environment variable access from direct process.env to centralized getEnv() utility function across all packages.
- Renamed environment variable prefix from OMP_ to PI_ throughout codebase (e.g., OMP_CODING_AGENT_DIR -> PI_CODING_AGENT_DIR).
- Removed automatic environment variable migration from PI_ to OMP_ prefixes via migrate-env.ts module.
- Removed env setting from configuration schema and applyEnvironmentVariables() method from settings.
- Updated CI/CD build configuration to use PI_COMPILED flag instead of OMP_COMPILED.
- Changed venvPath property in PythonRuntime from nullable (string | null) to optional (string | undefined).
- Added default generic type parameter to Model interface, allowing Model to be used without explicit type argument.
- Removed explicit <any> generic type parameters from Model type annotations throughout codebase, leveraging new default parameter.
- Added Kimi Code provider integration with OAuth device authorization flow and token management.
- Added four new Kimi Code models (kimi-for-coding, kimi-k2, kimi-k2-turbo-preview, kimi-k2.5) with reasoning support and 262K context window.
- Added kimiUsageProvider for fetching and caching Kimi Code API usage quota information.
- Added kimi-code login command to CLI for OAuth authentication with Kimi Code.
- Updated openai-completions provider to support Kimi-specific headers and cached token formats.
- Updated MiniMax-M2 model pricing: input 1.2->0.6, output 1.2->3, cacheRead 0.6->0.1.
- Added ToolChoice type and toolChoice parameter support across all AI providers (OpenAI, Azure OpenAI, Anthropic, Google) enabling fine-grained control over tool/function selection during LLM calls.
- Added toolChoice override capability to Agent.prompt() method and session prompt options allowing callers to control tool selection behavior per request.
- Added provider-specific tool choice mapping functions (mapAnthropicToolChoice, mapGoogleToolChoice, mapOpenAiToolChoice) to normalize tool choice formats across different LLM APIs.
- Removed kernel heartbeat/ping mechanism from PythonKernel, simplifying health monitoring by relying on direct isAlive() checks instead of periodic HTTP requests.
- Updated model data generation script to reference 'zai-coding-plan' instead of 'zai' for model configuration.
- Regenerated models configuration with updated pricing data and floating-point precision corrections.
- Added GLM-4.7-Flash model to generated models list.
- Reformatted code indentation and spacing across multiple TypeScript files.
- Standardized export statements to single-line format in various modules.
- Fixed async/await precedence and corrected indentation in benchmark files.
- Added benchmark report files for GPT-5.1-codex-mini model performance.
- Added Vercel AI Gateway model fetching with 100+ model configurations.
- Updated Amazon Bedrock models with new Claude 4.5 EU variants.
- Added GPT-5.2-codex model across OpenAI, GitHub Copilot, and OpenCode providers.
- Changed MiniMax models to use anthropic-messages API format.
- Added thinkingFormat support to ZAI provider models.
- Improved LSP diagnostic output by stripping clippy URLs and noise.
Provider additions:
- Added Amazon Bedrock provider with bedrock-converse-stream API
- Added MiniMax provider with OpenAI-compatible API
- Added EU cross-region inference model variants for Bedrock
Provider fixes:
- Fixed Gemini CLI retries with header parsing and empty stream retry logic
- Fixed Bedrock tool call transforms via transformMessages
- Fixed z.ai thinking/reasoning params
- Fixed OpenRouter+Anthropic cache control
- Fixed OpenAI responses timeout and service tier options
- Fixed tool call ID normalization for cross-provider switches
- Fixed thought signature validation for Google providers
- Fixed prompt cache key using session ID
TUI improvements:
- Added OverlayOptions API with CSS-like positioning (SizeValue, percentages)
- Added OverlayHandle for programmatic visibility control
- Added visible callback for responsive overlays
- Added pad parameter to truncateToWidth
- Added pageUp/pageDown key support
- Fixed numbered list items showing 1. when code blocks break continuity
- Fixed overlay width overflow crash with complex ANSI sequences
- Fixed light theme colors for WCAG AA compliance
Coding agent fixes:
- Fixed /new command to create new session file
- Fixed session selector to stay open when folder has no sessions
- Added session header emission in JSON print mode
- Added queued message hint with theme.tree.hook
- Exported highlightCode and getLanguageFromPath for extensions
Also renamed transorm-messages.ts to transform-messages.ts (typo fix)