chore: update stale docs
This commit is contained in:
+40
-111
@@ -6,11 +6,11 @@ This document describes how the coding-agent currently loads models, applies ove
|
||||
|
||||
Primary implementation files:
|
||||
|
||||
- `src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration
|
||||
- `src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models
|
||||
- `src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences)
|
||||
- `src/session/auth-storage.ts` — re-exports `AuthStorage` from `@oh-my-pi/pi-ai` (`packages/ai/src/auth-storage.ts`); API key + OAuth resolution order
|
||||
- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models (`getBundledModels` / `getBundledProviders`) and `Model`/`compat` types
|
||||
- `packages/coding-agent/src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration
|
||||
- `packages/coding-agent/src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models
|
||||
- `packages/coding-agent/src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences)
|
||||
- `packages/coding-agent/src/session/auth-storage.ts` — re-exports `AuthStorage` from `@oh-my-pi/pi-ai`; API key + OAuth resolution order
|
||||
- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models and public model types
|
||||
|
||||
## Config file location and legacy behavior
|
||||
|
||||
@@ -30,19 +30,11 @@ Legacy behavior still present:
|
||||
providers:
|
||||
<provider-id>:
|
||||
# provider-level config
|
||||
equivalence:
|
||||
overrides:
|
||||
<provider-id>/<model-id>: <canonical-model-id>
|
||||
exclude:
|
||||
- <provider-id>/<model-id>
|
||||
```
|
||||
|
||||
`provider-id` is the canonical provider key used across selection and auth lookup.
|
||||
|
||||
`equivalence` is optional and configures canonical model grouping on top of concrete provider models:
|
||||
|
||||
- `overrides` maps an exact concrete selector (`provider/modelId`) to an official upstream canonical id
|
||||
- `exclude` opts a concrete selector out of canonical grouping
|
||||
The root object currently contains only `providers`; unknown root keys fail schema validation.
|
||||
|
||||
## Provider-level fields
|
||||
|
||||
@@ -98,6 +90,7 @@ providers:
|
||||
- `openai-codex-responses`
|
||||
- `azure-openai-responses`
|
||||
- `anthropic-messages`
|
||||
- `bedrock-converse-stream`
|
||||
- `google-generative-ai`
|
||||
- `google-gemini-cli`
|
||||
- `google-vertex`
|
||||
@@ -178,78 +171,22 @@ the static + dynamic merge is bypassed entirely. The fingerprint is
|
||||
memoized per process by tagging the static-models array with a symbol
|
||||
property, so repeated cold-start calls do not re-hash.
|
||||
|
||||
## Canonical model equivalence and coalescing
|
||||
## Provider and model identity
|
||||
|
||||
The registry keeps every concrete provider model and then builds a canonical layer above them.
|
||||
|
||||
Canonical ids are official upstream ids only, for example:
|
||||
|
||||
- `claude-opus-4-6`
|
||||
- `claude-haiku-4-5`
|
||||
- `gpt-5.3-codex`
|
||||
|
||||
### `models.yml` equivalence config
|
||||
|
||||
Example:
|
||||
|
||||
```yaml
|
||||
providers:
|
||||
zenmux:
|
||||
baseUrl: https://api.zenmux.example/v1
|
||||
apiKey: ZENMUX_API_KEY
|
||||
api: openai-codex-responses
|
||||
models:
|
||||
- id: codex
|
||||
name: Zenmux Codex
|
||||
reasoning: true
|
||||
input: [text]
|
||||
cost:
|
||||
input: 0
|
||||
output: 0
|
||||
cacheRead: 0
|
||||
cacheWrite: 0
|
||||
contextWindow: 200000
|
||||
maxTokens: 32768
|
||||
|
||||
equivalence:
|
||||
overrides:
|
||||
zenmux/codex: gpt-5.3-codex
|
||||
p-codex/codex: gpt-5.3-codex
|
||||
exclude:
|
||||
- demo/codex-preview
|
||||
```
|
||||
|
||||
Build order for canonical grouping:
|
||||
|
||||
1. exact user override from `equivalence.overrides`
|
||||
2. bundled official-id matches from built-in model metadata
|
||||
3. conservative heuristic normalization for gateway/provider variants
|
||||
4. fallback to the concrete model's own id
|
||||
|
||||
Current heuristics are intentionally narrow:
|
||||
|
||||
- embedded upstream prefixes can be stripped when present, for example `anthropic/...` or `openai/...`
|
||||
- dotted and dashed version variants can normalize only when they map to an existing official id, for example `4.6 -> 4-6`
|
||||
- ambiguous families or versions are not merged without a bundled match or explicit override
|
||||
|
||||
### Canonical resolution behavior
|
||||
|
||||
When multiple concrete variants share a canonical id, resolution uses:
|
||||
|
||||
1. availability and auth
|
||||
2. `config.yml` `modelProviderOrder`
|
||||
3. existing registry/provider order if `modelProviderOrder` is unset
|
||||
|
||||
Disabled or unauthenticated providers are skipped.
|
||||
|
||||
Session state and transcripts continue to record the concrete provider/model that actually executed the turn.
|
||||
The registry retains concrete `provider` + `id` identities. Use an exact
|
||||
`provider/modelId` selector when the same model id exists under multiple providers. Session state
|
||||
and transcripts record the concrete provider/model that executed the turn.
|
||||
|
||||
Provider defaults vs per-model overrides:
|
||||
|
||||
- Provider `headers` are baseline.
|
||||
- Provider `headers`, `compat`, and `remoteCompaction` are baselines.
|
||||
- Model `headers` override provider header keys.
|
||||
- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`, `supportsTools`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`, `omitMaxOutputTokens`, `headers`, `compat`, `contextPromotionTarget`).
|
||||
- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`, `extraBody`).
|
||||
- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`,
|
||||
`supportsTools`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`,
|
||||
`omitMaxOutputTokens`, `headers`, `compat`, `contextPromotionTarget`, `compactionModel`, and
|
||||
`remoteCompaction`).
|
||||
- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`,
|
||||
`extraBody`, and `whenThinking`).
|
||||
|
||||
## Runtime discovery integration
|
||||
|
||||
@@ -420,20 +357,13 @@ So a model can exist in registry but not be selectable until auth is available.
|
||||
`model-resolver.ts` supports:
|
||||
|
||||
- exact `provider/modelId`
|
||||
- exact canonical model id
|
||||
- exact model id (provider inferred)
|
||||
- fuzzy/substring matching
|
||||
- glob scope patterns in `--models` (e.g. `openai/*`, `*sonnet*`)
|
||||
- optional `:thinkingLevel` suffix (`off|minimal|low|medium|high|xhigh|max`)
|
||||
|
||||
`--provider` is legacy; `--model` is preferred.
|
||||
|
||||
Resolution precedence for exact selectors:
|
||||
|
||||
1. exact `provider/modelId` bypasses coalescing
|
||||
2. exact canonical id resolves through the canonical index
|
||||
3. exact bare concrete id still works
|
||||
4. fuzzy and glob matching run after the exact paths
|
||||
`--provider` is legacy; `--model` is preferred. An exact `provider/modelId` is unambiguous; bare ids
|
||||
and fuzzy patterns are resolved against the available concrete models.
|
||||
|
||||
### Initial model selection priority
|
||||
|
||||
@@ -461,20 +391,12 @@ Related settings:
|
||||
|
||||
- `modelRoles` (record)
|
||||
- `enabledModels` (scoped pattern list)
|
||||
- `modelProviderOrder` (global canonical-provider precedence)
|
||||
- `modelProviderOrder` (provider precedence when equivalent concrete choices share an id)
|
||||
- `providers.kimiApiFormat` (`openai` or `anthropic` request format)
|
||||
- `providers.openaiWebsockets` (`auto|off|on` websocket preference for OpenAI Codex transport)
|
||||
|
||||
`modelRoles` may store either:
|
||||
|
||||
- `provider/modelId` to pin a concrete provider variant
|
||||
- a canonical id such as `gpt-5.3-codex` to allow provider coalescing
|
||||
|
||||
For `enabledModels` and CLI `--models`:
|
||||
|
||||
- exact canonical ids expand to all concrete variants in that canonical group
|
||||
- explicit `provider/modelId` entries stay exact
|
||||
- globs and fuzzy matches still operate on concrete models
|
||||
`modelRoles` stores model selectors such as `provider/modelId`; `enabledModels` and CLI `--models`
|
||||
accept exact selectors, globs, and fuzzy matches.
|
||||
|
||||
Global `enabledModels` and `disabledProviders` entries may also be scoped to a path prefix:
|
||||
|
||||
@@ -495,14 +417,8 @@ String entries apply everywhere. Scoped entries apply when the current working d
|
||||
|
||||
## `/model` and `omp models`
|
||||
|
||||
Both surfaces keep provider-prefixed models visible and selectable.
|
||||
|
||||
They now also expose canonical/coalesced models:
|
||||
|
||||
- `/model` includes a canonical view alongside provider tabs
|
||||
- `omp models` prints provider-grouped tables of every concrete model, and `omp models canonical` prints the coalesced canonical view
|
||||
|
||||
Selecting a canonical entry stores the canonical selector. Selecting a provider row stores the explicit `provider/modelId`.
|
||||
Both surfaces keep provider-prefixed concrete models visible and selectable. Selecting a provider
|
||||
row stores its explicit `provider/modelId`.
|
||||
|
||||
## Context promotion (model-level fallback chains)
|
||||
|
||||
@@ -580,6 +496,8 @@ Request shaping:
|
||||
- `streamIdleTimeoutMs` — stream-watchdog idle-timeout floor in ms for slow reasoning hosts. Default: auto (GLM coding-plan hosts, direct DeepSeek reasoning).
|
||||
- `cacheControlFormat` — `"anthropic"` to include Anthropic-style prompt-cache markers in chat-completions payloads. Default: auto (OpenRouter `anthropic/*` models).
|
||||
- `supportsLongPromptCacheRetention` — host honors `prompt_cache_retention: "24h"` on the Responses API. Default: auto (api.openai.com).
|
||||
- `supportsImageDetailOriginal` — allow the Responses API's nonstandard `detail: "original"` image
|
||||
mode where the endpoint supports it.
|
||||
- `extraBody` — extra top-level fields merged into every request body (gateway hints, controller selectors, etc.).
|
||||
|
||||
Reasoning / thinking:
|
||||
@@ -608,11 +526,22 @@ Gateway routing (only applied when `baseUrl` matches the gateway):
|
||||
- `openRouterRouting.only` / `openRouterRouting.order` — provider routing on `openrouter.ai` (see <https://openrouter.ai/docs/provider-routing>).
|
||||
- `vercelGatewayRouting.only` / `vercelGatewayRouting.order` — provider routing on `ai-gateway.vercel.sh` (see <https://vercel.com/docs/ai-gateway/models-and-providers/provider-options>).
|
||||
|
||||
Provider-level `compat` is the baseline; per-model `compat` is deep-merged on top, with `openRouterRouting`, `vercelGatewayRouting`, and `extraBody` merged as nested objects.
|
||||
Provider-level `compat` is the baseline; per-model `compat` is deep-merged on top, with
|
||||
`openRouterRouting`, `vercelGatewayRouting`, `extraBody`, and `whenThinking` merged as nested objects.
|
||||
|
||||
### Anthropic compatibility (`anthropic-messages`)
|
||||
|
||||
For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a top-level provider field (see below) plus two Anthropic-side flags in the same `compat` slot — `requiresToolResultId` (non-standard `id` alias on `tool_result` blocks for Z.AI-style proxies) and `replayUnsignedThinking` (replay unsigned thinking blocks as native thinking instead of demoting them to text); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`, `supportsForcedToolChoice`, `supportsSamplingParams`, `escapeBuiltinToolNames`) are set by built-in catalog metadata and are not user-configurable from `models.yml`.
|
||||
For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape
|
||||
(`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a
|
||||
top-level provider field plus `requiresToolResultId`, `replayUnsignedThinking`,
|
||||
`supportsEagerToolInputStreaming`, and `allowAnthropicHeaderOverrides` in `compat`. Other
|
||||
Anthropic-side knobs are supplied by built-in catalog metadata and are not configurable here.
|
||||
|
||||
### Bedrock compatibility (`bedrock-converse-stream`)
|
||||
|
||||
The same `compat` slot accepts `promptCacheMode` (`none`, `automatic`, or `explicit`),
|
||||
`supportsLongPromptCacheRetention`, `promptCacheMinimumTokens`, and
|
||||
`promptCacheMaximumCheckpoints` for Bedrock models.
|
||||
|
||||
### Strict tool schemas (`disableStrictTools`)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user