- Added GPT-5.3 Codex Spark model with 128K context window and extended reasoning capabilities.
- Added MiniMax M2.5 and M2.5 Lightning models via OpenAI-compatible API (minimax-code and minimax-code-cn providers).
- Added MiniMax M2.5 and M2.5 Lightning models via Anthropic API (minimax and minimax-cn providers).
- Added Llama 3.1 8B model via Cerebras API.
- Added MiniMax M2.5 model via OpenRouter and Vercel AI Gateway.
- Added Qwen3 VL 32B Instruct multimodal model via OpenRouter.
- show help instead of crashing on `omp setup` with no args
- show runtime-discovered MCP servers in `/mcp list`
- remove deprecated Anthropic model entries from models.json
- sort models by recency in model selector
- Sorted Antigravity models alphabetically in generated models.json and generation script.
- Updated Antigravity model configurations with corrected model IDs, names, and parameters.
- Fixed qwen/qwen3-max-thinking maxTokens from 4096 to 202752 to match context window.
Add support for MiniMax Coding Plan with OpenAI-compatible API:
- New providers: minimax-code (international) and minimax-code-cn (China)
- Environment variables: MINIMAX_CODE_API_KEY and MINIMAX_CODE_CN_API_KEY
- Uses thinkingFormat: 'zai' for reasoning compatibility
- Models: MiniMax-M2, MiniMax-M2.1, MiniMax-M2.1-lightning
- Base URLs: https://api.minimax.io/v1 (intl), https://api.minimaxi.com/v1 (CN)
The Coding Plan is a subscription-based service separate from regular MiniMax API.
- Migrated models export from TypeScript module to JSON format, changing the public API from importing MODELS from './models.generated' to importing from './models.json' with JSON import assertion.
- Updated @anthropic-ai/sdk dependency from ^0.72.1 to ^0.74.0.
- Simplified model generation script by replacing 49 lines of TypeScript code generation with direct JSON serialization.
- Updated @types/bun devDependency from ^1.3.8 to ^1.3.9 across all packages.
- Removed models.generated.ts exclusion from biome.json linting configuration.
- Added support for Kimi K2, K2 Turbo Preview, and K2.5 models with reasoning capabilities.
- Fixed Claude Opus 4.6 context window to 200K across all providers (was incorrectly set to 1M).
- Fixed Claude Sonnet 4 context window to 200K across multiple providers (was incorrectly set to 1M).
- Improved Kimi Code model fetching to merge fallback models not returned by the API endpoint.
- Added dynamic Antigravity model fetching from API when credentials are available, with hardcoded fallback models for offline use.
- Updated Antigravity models to use free tier pricing (0 cost) across all models.
- Extracted Antigravity model loading into separate functions (fetchAntigravityModels, getAntigravityFallbackModels, getAntigravityToken) for better maintainability.
- Added support for fetching recommended models from Antigravity API response and filtering internal models.
- Added Claude Opus 4.6 Thinking model for Antigravity provider.
- Added Gemini 2.5 Flash, Gemini 2.5 Flash Thinking, and Gemini 2.5 Pro models for Antigravity provider.
- Added Pony Alpha model via OpenRouter.
- Updated Claude Opus 4.6 context window from 200,000 to 1,000,000 tokens across Bedrock regions.
- Updated Claude Opus 4.6 cache pricing and Antigravity model pricing to free tier across multiple models.
- Fixed Claude Opus 4.6 model ID format by removing version suffix (:0) in Bedrock configurations.
- Extracted stream parsing utilities into reusable functions in pi-utils package (readLines, readJsonl, parseJsonlLenient, readSseJson).
- Replaced manual buffer management with ConcatSink class for efficient stream handling across multiple modules.
- Simplified SSE stream parsing in google-gemini-cli provider by using readSseJson utility instead of manual reader setup.
- Refactored MCP stdio transport to use readLines utility for cleaner line-based stream processing.
- Updated RPC client and mode to use readJsonl utility, removing duplicate JSONL parsing logic.
- Optimized buffer allocations by using allocUnsafe where data is immediately populated and reusing empty buffer instances.
- Migrated environment variable access from function-based `getEnv()` API to object property-based `$env` API across 43 files in the monorepo.
- Simplified `packages/utils/src/env.ts` by removing scoped environment variable logic (EnvScope, getEnvMap, getEnv functions) and replacing with direct `process.env` reference exported as `$env`.
- Updated `packages/ai/src/stream.ts` to remove environment parameter from KeyResolver functions and use global `$env` object instead of passed-in env parameter.
- Migrated environment variable access from direct process.env to centralized getEnv() utility function across all packages.
- Renamed environment variable prefix from OMP_ to PI_ throughout codebase (e.g., OMP_CODING_AGENT_DIR -> PI_CODING_AGENT_DIR).
- Removed automatic environment variable migration from PI_ to OMP_ prefixes via migrate-env.ts module.
- Removed env setting from configuration schema and applyEnvironmentVariables() method from settings.
- Updated CI/CD build configuration to use PI_COMPILED flag instead of OMP_COMPILED.
- Changed venvPath property in PythonRuntime from nullable (string | null) to optional (string | undefined).
- Added default generic type parameter to Model interface, allowing Model to be used without explicit type argument.
- Removed explicit <any> generic type parameters from Model type annotations throughout codebase, leveraging new default parameter.
- Added Kimi Code provider integration with OAuth device authorization flow and token management.
- Added four new Kimi Code models (kimi-for-coding, kimi-k2, kimi-k2-turbo-preview, kimi-k2.5) with reasoning support and 262K context window.
- Added kimiUsageProvider for fetching and caching Kimi Code API usage quota information.
- Added kimi-code login command to CLI for OAuth authentication with Kimi Code.
- Updated openai-completions provider to support Kimi-specific headers and cached token formats.
- Updated MiniMax-M2 model pricing: input 1.2->0.6, output 1.2->3, cacheRead 0.6->0.1.
- Added ToolChoice type and toolChoice parameter support across all AI providers (OpenAI, Azure OpenAI, Anthropic, Google) enabling fine-grained control over tool/function selection during LLM calls.
- Added toolChoice override capability to Agent.prompt() method and session prompt options allowing callers to control tool selection behavior per request.
- Added provider-specific tool choice mapping functions (mapAnthropicToolChoice, mapGoogleToolChoice, mapOpenAiToolChoice) to normalize tool choice formats across different LLM APIs.
- Removed kernel heartbeat/ping mechanism from PythonKernel, simplifying health monitoring by relying on direct isAlive() checks instead of periodic HTTP requests.
- Updated model data generation script to reference 'zai-coding-plan' instead of 'zai' for model configuration.
- Regenerated models configuration with updated pricing data and floating-point precision corrections.
- Added GLM-4.7-Flash model to generated models list.
- Reformatted code indentation and spacing across multiple TypeScript files.
- Standardized export statements to single-line format in various modules.
- Fixed async/await precedence and corrected indentation in benchmark files.
- Added benchmark report files for GPT-5.1-codex-mini model performance.
- Added Vercel AI Gateway model fetching with 100+ model configurations.
- Updated Amazon Bedrock models with new Claude 4.5 EU variants.
- Added GPT-5.2-codex model across OpenAI, GitHub Copilot, and OpenCode providers.
- Changed MiniMax models to use anthropic-messages API format.
- Added thinkingFormat support to ZAI provider models.
- Improved LSP diagnostic output by stripping clippy URLs and noise.
Provider additions:
- Added Amazon Bedrock provider with bedrock-converse-stream API
- Added MiniMax provider with OpenAI-compatible API
- Added EU cross-region inference model variants for Bedrock
Provider fixes:
- Fixed Gemini CLI retries with header parsing and empty stream retry logic
- Fixed Bedrock tool call transforms via transformMessages
- Fixed z.ai thinking/reasoning params
- Fixed OpenRouter+Anthropic cache control
- Fixed OpenAI responses timeout and service tier options
- Fixed tool call ID normalization for cross-provider switches
- Fixed thought signature validation for Google providers
- Fixed prompt cache key using session ID
TUI improvements:
- Added OverlayOptions API with CSS-like positioning (SizeValue, percentages)
- Added OverlayHandle for programmatic visibility control
- Added visible callback for responsive overlays
- Added pad parameter to truncateToWidth
- Added pageUp/pageDown key support
- Fixed numbered list items showing 1. when code blocks break continuity
- Fixed overlay width overflow crash with complex ANSI sequences
- Fixed light theme colors for WCAG AA compliance
Coding agent fixes:
- Fixed /new command to create new session file
- Fixed session selector to stay open when folder has no sessions
- Added session header emission in JSON print mode
- Added queued message hint with theme.tree.hook
- Exported highlightCode and getLanguageFromPath for extensions
Also renamed transorm-messages.ts to transform-messages.ts (typo fix)
- Added conversation state caching to persist context across multiple Cursor API requests in the same session.
- Added cursor-log.py script for filtering and displaying Cursor debug logs with follow mode, verbose mode, and delta coalescing support.
- Changed Cursor debug logging to use structured JSONL format with automatic MCP argument decoding.
- Moved proto-extractor.py from provider-specific directory to centralized scripts directory.
- Added model generation script to automatically fetch and update AI model definitions from models.dev and OpenRouter APIs.
- Implemented support for multiple providers including Anthropic, Google, OpenAI, Groq, Cerebras, xAI, Mistral, OpenCode, GitHub Copilot, OpenRouter, Google Vertex, and Cursor.
- Updated generated file comment to reference bun instead of npm for running the script.
- Co-located tool renderers with their respective tool implementations across 14 tool modules.
- Added render-utils module with shared formatting helpers for diagnostics, diffs, paths, and metadata.
- Removed ~800 lines of inline renderer implementations from centralized renderers.ts file.
- Added Google Vertex AI provider with 11 Gemini model configurations.
- Added OpenAI Codex provider with 13 GPT-5 model configurations.
Merged 53 commits from pi/main including:
- Export HTML rewrite with tree sidebar and /share command
- Hooks API enhancements (custom status, editor, session management)
- Edit diff preview before tool execution
- TUI fixes (Unicode format chars, OSC 8 hyperlinks, emoji optimization)
- Gemini CLI rate limit retry with server-provided delay
- Various bug fixes and documentation updates
Resolved conflicts by preserving Bun migration (spawn→Bun.spawn,
fs→Bun.file) while accepting upstream's new features and APIs.
- Migrated from Node.js to Bun runtime to leverage native TypeScript execution and faster startup times.
- Replaced npm with bun package manager and updated lockfile format.
- Converted Node.js APIs (child_process, fs, readline) to Bun equivalents across all packages.
- Updated shebang lines from node/tsx to bun for CLI entry points.
- Removed TypeScript build step by using source files directly as package entry points.
- Updated CI workflows to use Bun instead of Node.js for installation and testing.
- Migrate glm-4.5, glm-4.5-air, glm-4.5-flash, glm-4.6, glm-4.7 from anthropic-messages to openai-completions API
- Updated baseUrl from https://api.z.ai/api/anthropic to https://api.z.ai/api/coding/paas/v4
- Added compat setting to disable developer role for zai models
- Filter empty text blocks in openai-completions to avoid zai API validation errors
- Fixed zai provider tests to use OpenAI-style options (reasoningEffort)
- Add new API type 'google-cloud-code-assist' for Gemini CLI / Antigravity auth
- Extract shared Google utilities to google-shared.ts
- Implement streaming provider for Cloud Code Assist endpoint
- Add 7 models: gemini-3-pro-high/low, gemini-3-flash, claude-sonnet/opus, gpt-oss
Models use OAuth authentication and have sh cost (uses Google account quota).
OAuth flow will be implemented in coding-agent in a follow-up.
- Auto-enable all models after /login via POST /models/{model}/policy
- Use openai-responses API for gpt-5/o3/o4 models (not accessible via completions)
- Normalize tool call IDs when switching between github-copilot models with different APIs
(fixes#198: openai-responses generates 450+ char IDs with special chars that break other models)
- Update README with streamlined GitHub Copilot docs
- OAuth login for GitHub Copilot via /login command
- Support for github.com and GitHub Enterprise
- Models sourced from models.dev (Claude, GPT, Gemini, Grok, etc.)
- Dynamic base URL from token's proxy-ep field
- Use vscode-chat integration ID for API compatibility
- Documentation for model enablement at github.com/settings/copilot/features
Co-authored-by: cau1k <cau1k@users.noreply.github.com>
- add GitHub Copilot model discovery (env token fallback, headers,
compat) plus fallback list and quoted provider keys in generated map
- surface Copilot provider end-to-end (KnownProvider/default, env+OAuth
token refresh/save, enterprise base URL swap, available only when
creds/env exist)
- tweak interactive OAuth UI to render instruction text and prompt
placeholders
gpt-5.2-high took about 35 minutes. It had a lot of trouble with `npm
check` and went off on a "let's adjust every tsconfig" side quest.
Device code flow works, but the ai/scripts/generate-models.ts impl is
wrong as models from months ago are missing and only those deprecated
are accessible in the /models picker.
- Add Mistral to KnownProvider type and model generation
- Implement Mistral-specific compat handling in openai-completions:
- requiresToolResultName: tool results need name field
- requiresAssistantAfterToolResult: synthetic assistant message between tool/user
- requiresThinkingAsText: thinking blocks as <thinking> text
- requiresMistralToolIds: tool IDs must be exactly 9 alphanumeric chars
- Add MISTRAL_API_KEY environment variable support
- Add Mistral tests across all test files
- Update documentation (README, CHANGELOG) for both ai and coding-agent packages
- Remove client IDs from gemini.md, reference upstream source instead
Closes#165
Major improvements to mom's logging and cost reporting:
Centralized Logging System:
- Add src/log.ts with type-safe logging functions
- Colored console output (green=user, yellow=mom, dim=details)
- Consistent format: [HH:MM:SS] [context] message
- Replace scattered console.log/error calls throughout codebase
Usage Tracking & Cost Reporting:
- Track tokens (input, output, cache read/write) and costs per run
- Display summary at end of each run in console and Slack thread
- Example: 💰 Usage: 12,543 in + 847 out (5,234 cache read) = $0.0234
Prompt Caching Optimization:
- Move recent messages from system prompt to user message
- System prompt now mostly static (only changes with memory files)
- Enables effective use of Anthropic's prompt caching
- Significantly reduces costs on subsequent requests
Model & Cost Improvements:
- Switch from Claude Opus 4.5 to Sonnet 4.5 (~40% cost reduction)
- Fix Claude Opus 4.5 cache pricing in ai package (was 3x too expensive)
- Add manual override in generate-models.ts until upstream fix merges
- Submitted PR to models.dev: https://github.com/sst/models.dev/pull/439
UI/UX Improvements:
- Extract actual text from tool results instead of JSON wrapper
- Cleaner Slack thread formatting with duration and labels
- Tool args formatting shows paths with offset:limit notation
- Add chalk for colored terminal output
Dependencies:
- Add chalk package for terminal colors
- Created shared wrapTextWithAnsi() function in utils.ts
- Handles word-based wrapping while preserving ANSI escape codes
- Properly tracks active ANSI codes across wrapped lines
- Supports multi-byte characters (emoji, surrogate pairs)
- Updated Markdown and Text components to use shared wrapping
- Removed duplicate wrapping logic (158 lines total)