- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
- Added `qwenTemplateReasoningEffort` model compatibility option to disable Qwen chat template kwargs for strict local servers.
- Implemented fallback handling to strip rejected `chat_template_kwargs.reasoning_effort` and hoist values to top-level fields.
- Added comprehensive test coverage for Qwen reasoning effort fallback and keyword rejections.
Logged one actionable warning per LiteLLM management base when rich metadata discovery fails, while keeping missing 404 routes silent.
Documented metadata-route permissions and covered forbidden versus absent endpoints.
Fixes#5801
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
Aligned the context promotion documentation with the explicit contextPromotionTarget runtime behavior and the false default in settings-schema.
Fixes#5163
Loaded default custom model configuration from models.yaml when models.yml is absent while preserving yml precedence over yaml and legacy json migration.
Fixes#5145
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Added awaited `onTurnEnd` and `setOnTurnEnd` wiring for turn-end callbacks.
- Added `advisor.syncBacklog` settings (off/1/3/5) and documented 30-second catch-up caps.
- Fixed advisor runtime backlog handling with failure counters, waiters, and retry requeue.
- Updated agent sessions to enqueue advisor updates on turn end and removed direct `turn_end` branch logic.
- Added a new built-in `title` model role with `hidden` metadata and updated role definitions and schema.
- Updated title generation to resolve models in `title`, `commit`, then `smol` order and added test coverage for that precedence.
- Filtered hidden roles from selector badges and documented the new built-in role in model/settings docs.
Read vLLM max_model_len and OpenAI-compatible context_length metadata during model discovery, route providers.vllm.baseUrl into built-in discovery before cached models exist, and avoid sending local placeholder bearer tokens.
Scope the vLLM model cache to the discovery base URL so endpoint changes refetch immediately, and add focused regression coverage for configured and built-in vLLM discovery.
- docs/models.md: retitled the '/model and --list-models' section and
updated the bullet to describe 'omp models' plus 'omp models canonical'.
- docs/providers.md: updated the troubleshooting note to validate
models.yml with 'omp models' (and 'omp models find <substr>').
- packages/coding-agent/CHANGELOG.md: noted the doc fix under Unreleased.
Fixes#2458
Wrapped non-string singleton values when an array field is expected so malformed Anthropic-compatible tool calls can validate instead of looping on errors.
Fixes#2026
- Added `discovery.type: proxy` that hits `GET /v1/models` and routes each model via `supported_endpoint_types` (`anthropic` → `/v1/messages`, `openai` → `/v1/chat/completions`).
- Made provider-level `api` optional when `discovery.type` is `proxy`, since wire protocol is derived per-model.
- Increased discovery fetch timeout from 250ms to 10s to accommodate remote proxies.
- Documented proxy discovery configuration in `docs/models.md`.
Exposes model.compat.disableStrictTools (already supported by the anthropic
transport since #826) via models.yml so users can configure it without code
changes.
Set disableStrictTools: true at the provider level to disable strict tool
schemas for third-party Anthropic-compatible endpoints (AWS Bedrock, Vertex
AI proxies, custom gateways) that reject the strict field.
- Add disableStrictTools to ProviderConfigSchema
- Merge { disableStrictTools: true } into provider compat override when set,
flowing through the existing compat pipeline to model.compat.disableStrictTools
- disableStrictTools alone is sufficient for an override-only provider entry
- Update docs/models.md with field reference, Bedrock example, and proxy note
- Add tests covering provider-level propagation, built-in override, and
overlay merge
- Added canonical model equivalence types, cache helpers, and registry APIs for provider variant lookup.
- Changed model resolution to apply canonical ID overrides/excludes with provider order before fallback matching.
- Added canonical and provider model views in list-models and selector UI with canonical sorting/persistence.
- Updated role/model persistence to store selectors while runtime now resolves concrete canonical-backed provider models.
* add llama.cpp as local provider
* use responses api instead of messages
* use api-keys correctly for llama.cpp provider
---------
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
* feat: Add LM Studio as a supported model provider with OpenAI-compatible API fetching and discovery.
Add LM Studio as a supported AI provider with optional API key, environment variables, and discovery.
* feat: Add rustup as a dev dependency.
* fix: Refine LM Studio API key handling to conditionally send authorization headers during model discovery based on whether the key is a default local token or a custom key, and add new tests.
* feat: Enhance OAuth token and account ID resolution for model providers in the model registry.
* rebase for packages/ai/CHANGELOG.md
* feat: improve implicit model discovery to independently auto-detect Ollama and LM Studio, and refine LM Studio base URL handling.
---------
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
- Added ModelManager API with createModelManager() factory for managing bundled and dynamically discovered models with configurable refresh strategies.
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints with provider-specific model manager configuration helpers.
- Renamed public API functions for clarity: getModel() -> getBundledModel(), getModels() -> getBundledModels(), getProviders() -> getBundledProviders().
- Added on-disk model caching with TTL-based invalidation and resolveProviderModels() function for runtime model resolution with source precedence.
- Refactored model discovery script to dynamically fetch models from Codex, Cursor, and Antigravity using OAuth credentials instead of hardcoded lists.
- Changed context promotion to trigger on context overflow errors instead of a configurable threshold percentage.
- Removed the contextPromotion.thresholdPercent configuration setting.
- Updated context promotion to retry immediately on the promoted model without requiring compaction.
- Refactored context promotion logic to attempt promotion before compaction in the overflow handling flow.
- Updated agent session to merge context promotion checks into the compaction method for unified overflow handling.
- Updated tests to reflect overflow-based promotion triggering instead of threshold-based promotion.
- Moved documentation files from packages/coding-agent/docs/ to root docs/ directory to flatten the documentation structure.
- Updated all internal documentation links to account for the new file locations, adjusting relative paths to maintain correct references across the monorepo.
- Updated README.md and issue template configuration to reference documentation at the new root docs/ location instead of packages/coding-agent/docs/.