Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
- Seeded Anthropic curated fallback models into generation before discovery sources.
- Added claude-fable-5 and claude-mythos-5 model entries with 1,000,000 context and 128k max tokens.
- Enabled reasoning, text,image input, thinking controls, and larger limits across many models.
- Fixed OpenRouter Anthropic tool-call behavior for empty cache_control payloads.
- Derived descriptors, default-model map, env keys, login list, and refresh dispatch from one ProviderDefinition per provider.
- Disabled OpenAI Codex stream obfuscation and interrupted whitespace-only tool-call argument deltas.
- Derived auth-broker callback ports and paste-code login set from the registry.
- Added stripFireworksDeepSeekThinkingToggle to strip DeepSeek V4 `thinking` metadata before model compatibility is used.
- Applied the stripping helper in Fireworks model mapping and generation so produced models exclude the incompatible flag.
- Updated completion param construction to remove `thinking` when `reasoning_effort` is sent and added a regression test for Fireworks DeepSeek V4.
Fireworks /v1/models reports max_completion_tokens: 65536 generically for the
Kimi K2 family, but Kimi-on-Fireworks is documented to produce runaway
reasoning traces unless the output budget is bounded. The inflated value
flowed straight through fireworksModelManagerOptions.mapModel (forwarded
verbatim via toPositiveNumber), ended up in the bundled models.json for
fireworks/kimi-k2.5, fireworks/kimi-k2.6, and firepass/kimi-k2.6-turbo,
and was preserved across regenerations by prevModelsJson — so callers (and
the openai-completions default-injection safety net) could ship a budget the
router cannot honor.
Add a Fireworks-family Kimi cap (FIREWORKS_KIMI_MAX_TOKENS = 32_768) plus
isFireworksKimiK2ModelId / clampFireworksKimiMaxTokens helpers that recognize
both the public catalog ids (kimi-k2.5, kimi-k2.6, kimi-k2.6-turbo,
kimi-k2-thinking) and the canonical wire ids
(accounts/fireworks/{models,routers}/kimi-k2…). The Fireworks resolver
clamps every discovered Kimi K2.x model at runtime, and a new
applyFireworksKimiMaxTokensCap pass in generate-models.ts applies the same
ceiling to the fireworks/firepass slice of the assembled catalog so the
firepass static fallback and any future regens stay in sync.
Fixes#1849
Replaced Google Vertex project discovery with the models.dev catalog so bundled model selection includes current Vertex MaaS and Gemini entries while pruning retired fallbacks.
Fixes#1456
xai-oauth has no upstream catalog source (not in models.dev or
MODELS_DEV_PROVIDER_DESCRIPTORS) and dynamic discovery only fires after
the user authenticates and refresh() runs. Boot resolves the persisted
modelRoles.default synchronously from #loadModels(), which reads only
models.json — so a persisted "xai-oauth/<id>" default silently resets on
fresh starts because the bundle was empty.
Add XAI_OAUTH_CURATED_MODELS as the single source of truth for the
xai-oauth chat picker (grok-build, grok-4.3, grok-4.20-multi-agent-0309,
grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning) plus
buildXaiOAuthStaticSeed to render them as Model<"openai-responses">.
Wire the seed three ways:
- generate-models.ts pushes the seed into models.json so #loadModels()
sees it at boot.
- xaiOAuthModelManagerOptions hands the seed as staticModels so the
picker shows the catalog before fetchDynamicModels fires.
- test/xai-oauth-bundle.test.ts pins the bundle ⇔ seed invariant so
editing one without regenerating the other fails CI.
Op: correct
Restores: ref:feat/xai-grok-oauth@0f86695c1
Restores: ref:feat/xai-grok-oauth@1097373bc
Restores: ref:feat/xai-grok-oauth@eab1ed721
Centralizing OAuth refresh in AuthStorage (e6893515) introduced five
follow-on bugs surfaced by an audit of the commit; this fixes all of
them and updates the tests that relied on the old refresh seam.
1. packages/ai/src/auth-storage.ts (#tryOAuthCredential):
For built-in providers the path went directly to `getOAuthApiKey`
with the (possibly still-expired) selection.credential when the
pre-refresh at line 2587 caught a transient error. `getOAuthApiKey`
then threw the "expired … must be refreshed via AuthStorage"
precondition error, which the disable classifier matched against
`/expired.*refresh/` and soft-disabled the row. A single network
blip during refresh could permanently kill a still-valid Anthropic /
OpenAI / Gemini-CLI / Copilot credential. Built-in providers now
route through the broker-aware single-flighted
`#refreshOAuthCredential` first, so transient failures surface as
network errors (5-min temp block) instead of definitive auth
failures.
2. packages/ai/src/auth-storage.ts (#fetchUsageUncached):
The usage refresh check only fired once `Date.now() >= expiresAt`,
missing the 60-second skew that `getApiKey` honors. A token
expiring inside the skew window was posted to the usage endpoint
and 401'd mid-flight, briefly hiding quota in the UI. Aligned with
`OAUTH_REFRESH_SKEW_MS`.
3. packages/coding-agent/src/web/search/index.ts (webSearchCustomTool):
The CustomTool counterpart of WebSearchTool dropped sessionId so
SDK callers that opted into `web_search` via toolNames lost
per-session credential stickiness — multi-account users saw the
provider round-robin between searches in the same session. Threads
`ctx.sessionManager.getSessionId()` through to `executeSearch`.
4. packages/coding-agent/src/web/search/providers/perplexity.ts
(findOAuthToken):
`authStorage.getApiKey("perplexity")` returns runtime/config
overrides, stored api_key credentials, OAuth bearers, and env keys.
Filtering only env keys meant a config-pinned `pplx-…` API key was
POSTed to `www.perplexity.ai/rest/sse/perplexity_ask` (the OAuth
endpoint) instead of falling through to
`api.perplexity.ai/chat/completions`, producing 401s. Switched to
`getOAuthAccess` so only true OAuth bearers reach the OAuth
branch; api_key credentials/overrides correctly fall through.
5. packages/ai/scripts/generate-models.ts:
`getOAuthApiKey` was being called directly with possibly-expired
credentials. The new contract throws on expired, the broad catch
swallowed it, and the build silently fell back to bundled models
instead of refreshing. Both helpers now route through
AuthStorage's `getApiKey` / `getOAuthAccess`, which trigger the
full broker-aware refresh pipeline.
Test updates:
- auth-storage-credential-disabled-event.test.ts,
sdk-credential-disabled-bridge.test.ts: the `failOAuthRefresh`
helper used to spy on `getOAuthApiKey` to inject invalid_grant.
With refresh now happening before that helper, the spy never fired.
Switched to spying on `refreshOAuthToken` so the simulated failure
reaches the disable classifier.
- auth-storage-rotation.test.ts: stub `refreshOAuthToken` so the test
doesn't hit a real OAuth endpoint when the seeded credential lands
inside the 60s skew window.
- Added `AuthBrokerClient`, `RemoteAuthCredentialStore`, `AuthBrokerRefresher`, and `startAuthBroker` server in `packages/ai/src/auth-broker`.
- Renamed `AuthCredentialStore` class to `SqliteAuthCredentialStore`; extracted `AuthCredentialStore` as a persistence interface.
- Added `exportSnapshot`, `forceRefreshCredentialById`, `disableCredentialById`, and `upsertCredential` to `AuthStorage` for broker wire protocol.
- Added `omp auth-broker` CLI subcommand (serve, token, login, logout, import, status) and `discoverAuthStorage` broker-mode path keyed on `OMP_AUTH_BROKER_URL`.
- Updated Codex model pricing data to include non-zero input/output/cacheRead rates.
- Added generate-models fallback to copy billable OpenAI costs into bundled openai-codex models.
- Added catalog-cost helpers to fallback openai-codex pricing to openai and compute token totals.
- Added regression tests for openai-codex cost parity and db backfill, and documented the fix in changelog.
- Added Fireworks provider onboarding with API-key login and credential storage.
- Added Fireworks provider registrations across stream, types, descriptors, and OAuth registries.
- Mapped Fireworks public model IDs to wire IDs for OpenAI-completion requests.
- Extended OpenAI-compatible detection to treat api.fireworks.ai as supported for streaming and max_tokens.
- Updated model catalog handling with Fireworks models plus context-window/max-token inheritance and metadata updates.
- Added Fireworks model reference loading and fallback logic to preserve best context and token caps.
- Updated OpenAI context-promotion linking to handle `-spark` variants and base `gpt-5.5` models.
- Adjusted model generation data so OpenAI-like models now promote first to `gpt-5.5` and then to `gpt-5.4` where configured.
- Updated unit tests to assert context-promotion targets for both spark and `gpt-5.5` model variants.
Copilot's /models endpoint exposes token limits under capabilities.limits.
max_prompt_tokens is the actual prompt capacity (what OMP calls contextWindow),
while max_context_window_tokens is the total window (prompt + output budget).
Using the latter inflates contextWindow, which breaks compaction thresholds,
overflow detection, and context promotion.
Three changes:
1. mapModel reference selection: always prefer the Copilot-specific bundled
reference over the global cross-provider reference. Copilot imposes its
own limits that are strictly lower than native provider limits.
2. mapModel contextWindow chain: remove max_context_window_tokens from the
fallback. New chain: context_length -> max_prompt_tokens -> reference.
3. generate-models: stop overwriting contextWindow/maxTokens in
applyGlobalModelsDevFallback. These are provider-specific and should not
be replaced with cross-provider models.dev global references.
Also fixes bundled values: github-copilot/gpt-5.4 (400k -> 272k) and
github-copilot/gpt-5.2 (264k -> 128k) to match live API max_prompt_tokens.
Refs: #225, #226
- Fixed thinking configuration format by replacing `levels` array with `minLevel`/`maxLevel` properties across 100+ model definitions.
- Corrected GPT-5.4 mini/nano context window from 400000 to 272000 tokens for accurate token limit reporting.
- Normalized GPT-5.4 variant priority handling to use parsed variant instead of raw model IDs for consistent behavior.
- Added "mini" variant support to OpenAI model parsing regex and updated thinking mode configuration for Claude models.
- Fixed test robustness by replacing exact string matching with numeric range comparison to handle BSD seq notation on macOS.
- Corrected model generation script execution order to apply policy overrides before promotion target linking.
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Updated dev dependencies including biome, TypeScript native preview, and lint-staged to latest versions.
- Updated AI package dependencies for Anthropic SDK, AWS Bedrock, and Zod with relaxed version constraints.
- Updated models.json with new model entries, removed premiumMultiplier fields from GitHub Copilot models, and corrected maxTokens and cost values for various providers.
- Added appearance and projfs export paths to natives package and updated peer dependency version for swarm-extension.
- Extracted credential storage to shared @oh-my-pi/pi-ai package with AuthCredentialStore and AuthStorage classes.
- Consolidated UI formatting logic from ToolUIKit class into standalone utility functions across render-utils and output-meta modules.
- Moved utility functions (parseCommandArgs, substituteArgs, expandPath, normalizeUnicode) to dedicated modules for improved code reuse.
- Extracted JTD type definitions and type guards to jtd-utils module for shared use across schema conversion tools.
- Updated Claude model pricing and added cache read costs in models.json for accurate billing calculations.
- Refactored agent-storage to delegate credential management to AuthCredentialStore instead of direct SQLite operations.
- Added GitLab Duo provider with support for Claude, GPT-5, and Duo Chat models via GitLab AI Gateway.
- Added OAuth authentication for GitLab Duo with automatic token refresh, PKCE security, and 25-minute token caching.
- Added 16 new GitLab Duo models including Claude Opus/Sonnet/Haiku and GPT-5 variants with reasoning and multimodal support.
- Added `isOAuth` option to Anthropic provider for OAuth bearer token authentication mode.
- Exported `streamGitLabDuo`, `getGitLabDuoModels`, and `clearGitLabDuoDirectAccessCache` functions for GitLab Duo integration.
- Added configurable model refresh strategies (online, online-if-uncached) to control model discovery behavior.
- Implemented global model.dev fallback resolution with context window and token prioritization for improved model attribute consistency.
- Incremented model cache schema version to support improved global model fallback resolution.
- Enhanced model generation script with dedicated functions for building model reference maps and applying fallback attributes.
- Extracted provider descriptor helper functions (descriptor, catalog, catalogDescriptor, simpleModelsDevDescriptor, openAiCompletionsDescriptor, anthropicMessagesDescriptor) to reduce boilerplate across 47+ provider definitions.
- Refactored model generation script to simplify OAuth credential lifecycle by moving storage cleanup to finally block and parallelizing special discovery sources with Promise.all().
- Consolidated provider validation logic in model registry into reusable validateProviderConfiguration() function with context-aware validation modes.
- Reorganized models.json structure to place contextWindow and maxTokens before compat field for consistency across all provider entries.
- Unified provider descriptors into single source of truth in descriptors.ts module for runtime and catalog discovery.
- Consolidated model generation script to use declarative CatalogProviderDescriptor interface, reducing code duplication.
- Refactored model registry and selector to use descriptor pattern with priority-based sorting and version extraction.
- Added priority field to Model interface enabling provider-assigned model prioritization in discovery and selection.
- Added support for Synthetic model provider in web search command and improved model sorting by priority and version.
- Extracted model post-processing policies into dedicated model-policies module for improved testability and maintainability.
- Refactored model generation script to use declarative provider descriptors instead of 620+ lines of inline provider-specific logic.
- Consolidated provider model manager initialization to use descriptor-driven iteration, eliminating 29+ individual conditional blocks.
- Removed static bundled models for Ollama and vLLM from models.json to rely on dynamic discovery instead.
- Added support for 11 new AI providers (Hugging Face, NVIDIA, Together, Ollama, LiteLLM, Xiaomi, Moonshot, Venice, Qwen Portal, vLLM, Cloudflare AI Gateway) with API key authentication and login flows.
- Implemented $pickenv() utility for environment variable fallback chains, enabling multi-key resolution for providers with alternative credential names.
- Extended KnownProvider and OAuthProvider types to include all 11 new providers with corresponding model manager functions and OAuth handlers.
- Expanded models.json with thousands of new model entries across all new providers and replaced deprecated opencode provider with cloudflare-ai-gateway.
- Refactored model generation script to use unified fetchProviderModelsFromCatalog() and centralized API key resolution for all providers.
- Added ModelManager API with createModelManager() factory for managing bundled and dynamically discovered models with configurable refresh strategies.
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints with provider-specific model manager configuration helpers.
- Renamed public API functions for clarity: getModel() -> getBundledModel(), getModels() -> getBundledModels(), getProviders() -> getBundledProviders().
- Added on-disk model caching with TTL-based invalidation and resolveProviderModels() function for runtime model resolution with source precedence.
- Refactored model discovery script to dynamically fetch models from Codex, Cursor, and Antigravity using OAuth credentials instead of hardcoded lists.
- Added Claude Sonnet 4.6 and Claude Sonnet 4.6 Thinking models to Antigravity provider.
- Added GLM-5 Free model via OpenCode provider.
- Added GLM-4.7-FlashX model via ZAI provider.
- Added MiniMax-M2.5-highspeed model across four providers (minimax-code, minimax-code-cn, minimax, minimax-cn).
- Added Claude Sonnet 4.6 model to OpenRouter and Vercel AI Gateway providers.
- Updated pricing and token limits for deepseek-v3, mistral-large-2411, and Qwen models across OpenRouter and Together AI providers.
- Expanded EU cross-region inference variant support to all Claude models on Bedrock (previously limited to Haiku, Sonnet, and Opus 4.5).
- Updated context window handling for Sonnet 4.6 models to enforce 200K limit across all providers.
- Added contextPromotionTarget model property to specify preferred fallback model when context promotion is triggered.
- Added automatic context promotion target assignment for Spark models to their base model equivalents.
- Updated Qwen model context window and max token limits for improved accuracy.
- Updated o1 model context window from 256000 to 262144 tokens and max tokens from 64000 to 65536 tokens.
- Implemented context promotion logic to use configured contextPromotionTarget when available instead of role-based model resolution.
- Added DeepSeek-V3.2 model support via Amazon Bedrock.
- Added GLM-5 model support via OpenCode.
- Added MiniMax M2.5 model support via OpenCode.
- Updated GLM models to use anthropic-messages API instead of openai-completions and changed base URL from https://api.z.ai/api/coding/paas/v4 to https://api.z.ai/api/anthropic.
- Removed compat field with supportsDeveloperRole and thinkingFormat properties from GLM models.
- Updated pricing and context window specifications for multiple models including Mistral, Moonshot, and Qwen variants.
- Added sorting of models by ID across multiple model fetch functions (OpenRouter, AI Gateway, Kimi Code, and dev data) using localeCompare for consistent ordering.
- Added WebSocket transport support for OpenAI Codex responses with automatic fallback to SSE on connection failure.
- Added preferWebsockets option to Agent and Model configurations to hint that WebSocket transport should be preferred when supported by provider implementations.
- Added prewarmOpenAICodexResponses() function to pre-establish WebSocket connections for improved performance.
- Added getProviderDetails() function and getOpenAICodexTransportDetails() function to expose transport state and provider configuration information.
- Added provider details display in session info showing active provider configuration and authentication details.
- Added OpenAI websockets setting to enable WebSocket transport preference for OpenAI Codex models in coding agent configuration.
- Added GPT-5.3 Codex Spark model with 128K context window and extended reasoning capabilities.
- Added MiniMax M2.5 and M2.5 Lightning models via OpenAI-compatible API (minimax-code and minimax-code-cn providers).
- Added MiniMax M2.5 and M2.5 Lightning models via Anthropic API (minimax and minimax-cn providers).
- Added Llama 3.1 8B model via Cerebras API.
- Added MiniMax M2.5 model via OpenRouter and Vercel AI Gateway.
- Added Qwen3 VL 32B Instruct multimodal model via OpenRouter.
- show help instead of crashing on `omp setup` with no args
- show runtime-discovered MCP servers in `/mcp list`
- remove deprecated Anthropic model entries from models.json
- sort models by recency in model selector
- Sorted Antigravity models alphabetically in generated models.json and generation script.
- Updated Antigravity model configurations with corrected model IDs, names, and parameters.
- Fixed qwen/qwen3-max-thinking maxTokens from 4096 to 202752 to match context window.
Add support for MiniMax Coding Plan with OpenAI-compatible API:
- New providers: minimax-code (international) and minimax-code-cn (China)
- Environment variables: MINIMAX_CODE_API_KEY and MINIMAX_CODE_CN_API_KEY
- Uses thinkingFormat: 'zai' for reasoning compatibility
- Models: MiniMax-M2, MiniMax-M2.1, MiniMax-M2.1-lightning
- Base URLs: https://api.minimax.io/v1 (intl), https://api.minimaxi.com/v1 (CN)
The Coding Plan is a subscription-based service separate from regular MiniMax API.
- Migrated models export from TypeScript module to JSON format, changing the public API from importing MODELS from './models.generated' to importing from './models.json' with JSON import assertion.
- Updated @anthropic-ai/sdk dependency from ^0.72.1 to ^0.74.0.
- Simplified model generation script by replacing 49 lines of TypeScript code generation with direct JSON serialization.
- Updated @types/bun devDependency from ^1.3.8 to ^1.3.9 across all packages.
- Removed models.generated.ts exclusion from biome.json linting configuration.
- Added support for Kimi K2, K2 Turbo Preview, and K2.5 models with reasoning capabilities.
- Fixed Claude Opus 4.6 context window to 200K across all providers (was incorrectly set to 1M).
- Fixed Claude Sonnet 4 context window to 200K across multiple providers (was incorrectly set to 1M).
- Improved Kimi Code model fetching to merge fallback models not returned by the API endpoint.
- Added dynamic Antigravity model fetching from API when credentials are available, with hardcoded fallback models for offline use.
- Updated Antigravity models to use free tier pricing (0 cost) across all models.
- Extracted Antigravity model loading into separate functions (fetchAntigravityModels, getAntigravityFallbackModels, getAntigravityToken) for better maintainability.
- Added support for fetching recommended models from Antigravity API response and filtering internal models.
- Added Claude Opus 4.6 Thinking model for Antigravity provider.
- Added Gemini 2.5 Flash, Gemini 2.5 Flash Thinking, and Gemini 2.5 Pro models for Antigravity provider.
- Added Pony Alpha model via OpenRouter.
- Updated Claude Opus 4.6 context window from 200,000 to 1,000,000 tokens across Bedrock regions.
- Updated Claude Opus 4.6 cache pricing and Antigravity model pricing to free tier across multiple models.
- Fixed Claude Opus 4.6 model ID format by removing version suffix (:0) in Bedrock configurations.
- Extracted stream parsing utilities into reusable functions in pi-utils package (readLines, readJsonl, parseJsonlLenient, readSseJson).
- Replaced manual buffer management with ConcatSink class for efficient stream handling across multiple modules.
- Simplified SSE stream parsing in google-gemini-cli provider by using readSseJson utility instead of manual reader setup.
- Refactored MCP stdio transport to use readLines utility for cleaner line-based stream processing.
- Updated RPC client and mode to use readJsonl utility, removing duplicate JSONL parsing logic.
- Optimized buffer allocations by using allocUnsafe where data is immediately populated and reusing empty buffer instances.
- Migrated environment variable access from function-based `getEnv()` API to object property-based `$env` API across 43 files in the monorepo.
- Simplified `packages/utils/src/env.ts` by removing scoped environment variable logic (EnvScope, getEnvMap, getEnv functions) and replacing with direct `process.env` reference exported as `$env`.
- Updated `packages/ai/src/stream.ts` to remove environment parameter from KeyResolver functions and use global `$env` object instead of passed-in env parameter.