Commit Graph

45 Commits

Author SHA1 Message Date
can1357 0cdd37fc15 feat(pi-natives/tools): implemented utok tokenizer for multiple models
- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
2026-08-20 01:45:04 +02:00
can1357 74bc1f442e feat(ai): implemented fallback handling and option for qwen reasoning effort
- Added `qwenTemplateReasoningEffort` model compatibility option to disable Qwen chat template kwargs for strict local servers.
- Implemented fallback handling to strip rejected `chat_template_kwargs.reasoning_effort` and hoist values to top-level fields.
- Added comprehensive test coverage for Qwen reasoning effort fallback and keyword rejections.
2026-08-19 17:40:39 +02:00
Ethan Cawse 76b1169858 fix(coding-agent): drop stale native image replay 2026-08-12 09:16:57 -04:00
can1357 1a8caad23e chore: update docs + rename reset to clear 2026-08-03 18:37:23 +02:00
can1357 ebd5e3f86f chore: update stale docs 2026-08-03 16:39:23 +02:00
Wolfgang Schoenberger 031beb1751 docs(auth): correct credential precedence
(cherry picked from commit 33ede5efbc52cedc9819cb65c642cfeab82c337c)
2026-07-29 23:09:09 +02:00
roboomp 63cac8dfd5 fix(catalog): warned on LiteLLM metadata fallback
Logged one actionable warning per LiteLLM management base when rich metadata discovery fails, while keeping missing 404 routes silent.

Documented metadata-route permissions and covered forbidden versus absent endpoints.

Fixes #5801
2026-07-17 07:38:06 +00:00
can1357 f786fe4a6f Merge PR #5147: fix(coding-agent): support models.yaml custom config (@roboomp) 2026-07-14 18:39:26 +02:00
can1357 f9f6ed9e8d feat(coding-agent): replaced legacy pi/ role alias prefix with
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
2026-07-13 23:26:33 +02:00
roboomp cbe0832246 docs(coding-agent): corrected context promotion docs
Aligned the context promotion documentation with the explicit contextPromotionTarget runtime behavior and the false default in settings-schema.

Fixes #5163
2026-07-11 07:30:43 +00:00
roboomp 2d4d445f95 fix(coding-agent): supported models yaml config
Loaded default custom model configuration from models.yaml when models.yml is absent while preserving yml precedence over yaml and legacy json migration.

Fixes #5145
2026-07-11 03:52:45 +00:00
can1357 d435385ab1 feat: introduced max reasoning effort tier across model and rpc systems
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
2026-07-10 13:39:42 +02:00
can1357 f0f7a5ba89 feat(coding-agent): introduced tiny model role for background tasks
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
2026-06-27 07:56:27 +02:00
Jean-Luc Davern 3dc9099058 feat(coding-agent): discover rich LiteLLM metadata 2026-06-26 11:26:27 -05:00
can1357 d4317d3d20 feat(ai): consolidated OpenAI-family streaming and add OpenRouter API support
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.

Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
    - Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
    - Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
    - Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
2026-06-17 21:36:48 +02:00
Can Bölük b7fe6f61a1 Merge branch 'main' into fix/litellm-base-url 2026-06-17 11:37:28 +02:00
can1357 d5b3c78132 chore: update docs 2026-06-17 04:37:39 +02:00
Alexander Kirilin 11e6eb0c2c fix(litellm): support LITELLM_BASE_URL 2026-06-16 20:54:55 -04:00
can1357 dbf4c734fc feat(coding-agent): added advisor backlog sync and end-of-turn callback support
- Added awaited `onTurnEnd` and `setOnTurnEnd` wiring for turn-end callbacks.
- Added `advisor.syncBacklog` settings (off/1/3/5) and documented 30-second catch-up caps.
- Fixed advisor runtime backlog handling with failure counters, waiters, and retry requeue.
- Updated agent sessions to enqueue advisor updates on turn end and removed direct `turn_end` branch logic.
2026-06-15 17:25:04 +02:00
can1357 5a910863c3 feat(coding-agent): added hidden title role for prioritized title model selection
- Added a new built-in `title` model role with `hidden` metadata and updated role definitions and schema.
- Updated title generation to resolve models in `title`, `commit`, then `smol` order and added test coverage for that precedence.
- Filtered hidden roles from selector badges and documented the new built-in role in model/settings docs.
2026-06-15 00:23:29 +02:00
Gerben Meijer 4a6eec624b fix(coding-agent): preserve vLLM discovered context windows
Read vLLM max_model_len and OpenAI-compatible context_length metadata during model discovery, route providers.vllm.baseUrl into built-in discovery before cached models exist, and avoid sending local placeholder bearer tokens.

Scope the vLLM model cache to the discovery base URL so endpoint changes refetch immediately, and add focused regression coverage for configured and built-in vLLM discovery.
2026-06-14 17:50:16 +02:00
roboomp c1a8ff1a81 docs: replaced --list-models references with the omp models subcommand
- docs/models.md: retitled the '/model and --list-models' section and
  updated the bullet to describe 'omp models' plus 'omp models canonical'.
- docs/providers.md: updated the troubleshooting note to validate
  models.yml with 'omp models' (and 'omp models find <substr>').
- packages/coding-agent/CHANGELOG.md: noted the doc fix under Unreleased.

Fixes #2458
2026-06-13 17:33:29 +00:00
can1357 371846167b docs: update docs 2026-06-12 14:43:35 +02:00
can1357 bcf978a266 Merge pull request #2205: feat(models): resolve secrets from commands 2026-06-10 08:26:59 +02:00
danzaio 9c4d62b6ba docs(models): clarify oMLX local discovery 2026-06-10 08:26:00 +02:00
danzaio b36f5c0785 docs(models): document oMLX OpenAI-compatible setup 2026-06-10 08:26:00 +02:00
danzaio 887864fef0 feat(models): resolve secrets from commands 2026-06-10 08:26:00 +02:00
roboomp 0d6913e926 fix(ai): coerced singleton array arguments
Wrapped non-string singleton values when an array field is expected so malformed Anthropic-compatible tool calls can validate instead of looping on errors.

Fixes #2026
2026-06-07 06:31:30 +00:00
can1357 78700a2c77 feat(ollama): added OLLAMA_HOST and OLLAMA_CONTEXT_LENGTH support
- Used OLLAMA_HOST for implicit discovery when OLLAMA_BASE_URL is unset.
- Applied OLLAMA_CONTEXT_LENGTH override to discovered context budgeting.
2026-06-05 22:47:40 +02:00
can1357 1dba122c53 chore: updated docs 2026-05-31 04:36:14 +02:00
can1357 375d100555 feat(model-registry): added proxy discovery type for mixed Anthropic/OpenAI proxies
- Added `discovery.type: proxy` that hits `GET /v1/models` and routes each model via `supported_endpoint_types` (`anthropic` → `/v1/messages`, `openai` → `/v1/chat/completions`).
- Made provider-level `api` optional when `discovery.type` is `proxy`, since wire protocol is derived per-model.
- Increased discovery fetch timeout from 250ms to 10s to accommodate remote proxies.
- Documented proxy discovery configuration in `docs/models.md`.
2026-05-30 03:53:36 +02:00
can1357 c049613cab docs: added auth-broker, schema-normalize, install-id, and eval docs
- Added auth-broker-gateway.md covering remote OAuth vault, gateway forward-proxy, usage cache layering, and env surface.
- Added ai-schema-normalize.md documenting the unified tool-schema normalization pipeline and strict-mode edge cases.
- Added install-id.md describing the per-install UUID persistence and consumer contract.
- Rewrote eval.md to reflect structured JSON cells schema, removing the legacy `*** Cell` parser and Lark grammar.
- Updated environment-variables.md, models.md, sdk.md, secrets.md, lsp.md, session-tree-plan.md, ttsr-injection-lifecycle.md, and natives docs to match code changes.
2026-05-17 02:15:55 +02:00
can1357 d438c4e8a0 feat(coding-agent): support path-scoped model config
Fixes #947
2026-05-06 17:13:11 +02:00
can1357 5e67c50eb4 docs: refresh models.md compat section
Fixes #893
2026-05-02 07:46:52 +02:00
Christoph Gross f908ee9496 feat(coding-agent): add disableStrictTools provider option for anthropic-messages endpoints
Exposes model.compat.disableStrictTools (already supported by the anthropic
transport since #826) via models.yml so users can configure it without code
changes.

Set disableStrictTools: true at the provider level to disable strict tool
schemas for third-party Anthropic-compatible endpoints (AWS Bedrock, Vertex
AI proxies, custom gateways) that reject the strict field.

- Add disableStrictTools to ProviderConfigSchema
- Merge { disableStrictTools: true } into provider compat override when set,
  flowing through the existing compat pipeline to model.compat.disableStrictTools
- disableStrictTools alone is sufficient for an override-only provider entry
- Update docs/models.md with field reference, Bedrock example, and proxy note
- Add tests covering provider-level propagation, built-in override, and
  overlay merge
2026-04-30 08:28:03 +02:00
can1357 9865a4ce6c docs: update docs 2026-04-30 06:47:01 +02:00
can1357 5277e44139 feat(coding-agent): added canonical aliases for model role resolution
- Added canonical model equivalence types, cache helpers, and registry APIs for provider variant lookup.
- Changed model resolution to apply canonical ID overrides/excludes with provider order before fallback matching.
- Added canonical and provider model views in list-models and selector UI with canonical sorting/persistence.
- Updated role/model persistence to store selectors while runtime now resolves concrete canonical-backed provider models.
2026-04-11 08:28:50 +02:00
can1357 2a20bd302e feat(ai): support extra body fields in openai-completions requests
Fixes #363
2026-03-14 13:52:12 +01:00
Gregor 15c7429ad0 add llama.cpp as local provider (#370)
* add llama.cpp as local provider

* use responses api instead of messages

* use api-keys correctly for llama.cpp provider

---------

Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
2026-03-13 15:16:53 +01:00
can1357 4ed4cfb27c fix: honor per-role thinking in modelRoles helpers
Fixes #186
2026-03-11 00:03:48 +01:00
Sage Grigull 02ebc5e8e1 add LM Studio support (#259)
* feat: Add LM Studio as a supported model provider with OpenAI-compatible API fetching and discovery.

Add LM Studio as a supported AI provider with optional API key, environment variables, and discovery.

* feat: Add rustup as a dev dependency.

* fix: Refine LM Studio API key handling to conditionally send authorization headers during model discovery based on whether the key is a default local token or a custom key, and add new tests.

* feat: Enhance OAuth token and account ID resolution for model providers in the model registry.

* rebase for packages/ai/CHANGELOG.md

* feat: improve implicit model discovery to independently auto-detect Ollama and LM Studio, and refine LM Studio base URL handling.

---------

Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
2026-03-03 03:26:57 +01:00
can1357 bc69fd207d feat: implemented dynamic model resolution across all providers with ModelManager API
- Added ModelManager API with createModelManager() factory for managing bundled and dynamically discovered models with configurable refresh strategies.
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints with provider-specific model manager configuration helpers.
- Renamed public API functions for clarity: getModel() -> getBundledModel(), getModels() -> getBundledModels(), getProviders() -> getBundledProviders().
- Added on-disk model caching with TTL-based invalidation and resolveProviderModels() function for runtime model resolution with source precedence.
- Refactored model discovery script to dynamically fetch models from Codex, Cursor, and Antigravity using OAuth credentials instead of hardcoded lists.
2026-02-18 01:38:05 +01:00
can1357 0bbe52dee7 feat(coding-agent): changed context promotion to trigger on overflow errors instead of threshold
- Changed context promotion to trigger on context overflow errors instead of a configurable threshold percentage.
- Removed the contextPromotion.thresholdPercent configuration setting.
- Updated context promotion to retry immediately on the promoted model without requiring compaction.
- Refactored context promotion logic to attempt promotion before compaction in the overflow handling flow.
- Updated agent session to merge context promotion checks into the compaction method for unified overflow handling.
- Updated tests to reflect overflow-based promotion triggering instead of threshold-based promotion.
2026-02-16 18:33:05 +01:00
can1357 ce0919bdf0 docs(models): documented context promotion fallback chains 2026-02-16 18:33:04 +01:00
can1357 2e45297c43 docs(docs): moved documentation to root docs directory and updated all references
- Moved documentation files from packages/coding-agent/docs/ to root docs/ directory to flatten the documentation structure.
- Updated all internal documentation links to account for the new file locations, adjusting relative paths to maintain correct references across the monorepo.
- Updated README.md and issue template configuration to reference documentation at the new root docs/ location instead of packages/coding-agent/docs/.
2026-02-16 18:33:03 +01:00