chore: update stale docs

This commit is contained in:
can1357
2026-08-03 16:37:05 +02:00
parent fc04aa6fa7
commit ebd5e3f86f
120 changed files with 5246 additions and 4691 deletions
+40 -111
View File
@@ -6,11 +6,11 @@ This document describes how the coding-agent currently loads models, applies ove
Primary implementation files:
- `src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration
- `src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models
- `src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences)
- `src/session/auth-storage.ts` — re-exports `AuthStorage` from `@oh-my-pi/pi-ai` (`packages/ai/src/auth-storage.ts`); API key + OAuth resolution order
- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models (`getBundledModels` / `getBundledProviders`) and `Model`/`compat` types
- `packages/coding-agent/src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration
- `packages/coding-agent/src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models
- `packages/coding-agent/src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences)
- `packages/coding-agent/src/session/auth-storage.ts` — re-exports `AuthStorage` from `@oh-my-pi/pi-ai`; API key + OAuth resolution order
- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models and public model types
## Config file location and legacy behavior
@@ -30,19 +30,11 @@ Legacy behavior still present:
providers:
<provider-id>:
# provider-level config
equivalence:
overrides:
<provider-id>/<model-id>: <canonical-model-id>
exclude:
- <provider-id>/<model-id>
```
`provider-id` is the canonical provider key used across selection and auth lookup.
`equivalence` is optional and configures canonical model grouping on top of concrete provider models:
- `overrides` maps an exact concrete selector (`provider/modelId`) to an official upstream canonical id
- `exclude` opts a concrete selector out of canonical grouping
The root object currently contains only `providers`; unknown root keys fail schema validation.
## Provider-level fields
@@ -98,6 +90,7 @@ providers:
- `openai-codex-responses`
- `azure-openai-responses`
- `anthropic-messages`
- `bedrock-converse-stream`
- `google-generative-ai`
- `google-gemini-cli`
- `google-vertex`
@@ -178,78 +171,22 @@ the static + dynamic merge is bypassed entirely. The fingerprint is
memoized per process by tagging the static-models array with a symbol
property, so repeated cold-start calls do not re-hash.
## Canonical model equivalence and coalescing
## Provider and model identity
The registry keeps every concrete provider model and then builds a canonical layer above them.
Canonical ids are official upstream ids only, for example:
- `claude-opus-4-6`
- `claude-haiku-4-5`
- `gpt-5.3-codex`
### `models.yml` equivalence config
Example:
```yaml
providers:
zenmux:
baseUrl: https://api.zenmux.example/v1
apiKey: ZENMUX_API_KEY
api: openai-codex-responses
models:
- id: codex
name: Zenmux Codex
reasoning: true
input: [text]
cost:
input: 0
output: 0
cacheRead: 0
cacheWrite: 0
contextWindow: 200000
maxTokens: 32768
equivalence:
overrides:
zenmux/codex: gpt-5.3-codex
p-codex/codex: gpt-5.3-codex
exclude:
- demo/codex-preview
```
Build order for canonical grouping:
1. exact user override from `equivalence.overrides`
2. bundled official-id matches from built-in model metadata
3. conservative heuristic normalization for gateway/provider variants
4. fallback to the concrete model's own id
Current heuristics are intentionally narrow:
- embedded upstream prefixes can be stripped when present, for example `anthropic/...` or `openai/...`
- dotted and dashed version variants can normalize only when they map to an existing official id, for example `4.6 -> 4-6`
- ambiguous families or versions are not merged without a bundled match or explicit override
### Canonical resolution behavior
When multiple concrete variants share a canonical id, resolution uses:
1. availability and auth
2. `config.yml` `modelProviderOrder`
3. existing registry/provider order if `modelProviderOrder` is unset
Disabled or unauthenticated providers are skipped.
Session state and transcripts continue to record the concrete provider/model that actually executed the turn.
The registry retains concrete `provider` + `id` identities. Use an exact
`provider/modelId` selector when the same model id exists under multiple providers. Session state
and transcripts record the concrete provider/model that executed the turn.
Provider defaults vs per-model overrides:
- Provider `headers` are baseline.
- Provider `headers`, `compat`, and `remoteCompaction` are baselines.
- Model `headers` override provider header keys.
- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`, `supportsTools`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`, `omitMaxOutputTokens`, `headers`, `compat`, `contextPromotionTarget`).
- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`, `extraBody`).
- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`,
`supportsTools`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`,
`omitMaxOutputTokens`, `headers`, `compat`, `contextPromotionTarget`, `compactionModel`, and
`remoteCompaction`).
- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`,
`extraBody`, and `whenThinking`).
## Runtime discovery integration
@@ -420,20 +357,13 @@ So a model can exist in registry but not be selectable until auth is available.
`model-resolver.ts` supports:
- exact `provider/modelId`
- exact canonical model id
- exact model id (provider inferred)
- fuzzy/substring matching
- glob scope patterns in `--models` (e.g. `openai/*`, `*sonnet*`)
- optional `:thinkingLevel` suffix (`off|minimal|low|medium|high|xhigh|max`)
`--provider` is legacy; `--model` is preferred.
Resolution precedence for exact selectors:
1. exact `provider/modelId` bypasses coalescing
2. exact canonical id resolves through the canonical index
3. exact bare concrete id still works
4. fuzzy and glob matching run after the exact paths
`--provider` is legacy; `--model` is preferred. An exact `provider/modelId` is unambiguous; bare ids
and fuzzy patterns are resolved against the available concrete models.
### Initial model selection priority
@@ -461,20 +391,12 @@ Related settings:
- `modelRoles` (record)
- `enabledModels` (scoped pattern list)
- `modelProviderOrder` (global canonical-provider precedence)
- `modelProviderOrder` (provider precedence when equivalent concrete choices share an id)
- `providers.kimiApiFormat` (`openai` or `anthropic` request format)
- `providers.openaiWebsockets` (`auto|off|on` websocket preference for OpenAI Codex transport)
`modelRoles` may store either:
- `provider/modelId` to pin a concrete provider variant
- a canonical id such as `gpt-5.3-codex` to allow provider coalescing
For `enabledModels` and CLI `--models`:
- exact canonical ids expand to all concrete variants in that canonical group
- explicit `provider/modelId` entries stay exact
- globs and fuzzy matches still operate on concrete models
`modelRoles` stores model selectors such as `provider/modelId`; `enabledModels` and CLI `--models`
accept exact selectors, globs, and fuzzy matches.
Global `enabledModels` and `disabledProviders` entries may also be scoped to a path prefix:
@@ -495,14 +417,8 @@ String entries apply everywhere. Scoped entries apply when the current working d
## `/model` and `omp models`
Both surfaces keep provider-prefixed models visible and selectable.
They now also expose canonical/coalesced models:
- `/model` includes a canonical view alongside provider tabs
- `omp models` prints provider-grouped tables of every concrete model, and `omp models canonical` prints the coalesced canonical view
Selecting a canonical entry stores the canonical selector. Selecting a provider row stores the explicit `provider/modelId`.
Both surfaces keep provider-prefixed concrete models visible and selectable. Selecting a provider
row stores its explicit `provider/modelId`.
## Context promotion (model-level fallback chains)
@@ -580,6 +496,8 @@ Request shaping:
- `streamIdleTimeoutMs` — stream-watchdog idle-timeout floor in ms for slow reasoning hosts. Default: auto (GLM coding-plan hosts, direct DeepSeek reasoning).
- `cacheControlFormat` — `"anthropic"` to include Anthropic-style prompt-cache markers in chat-completions payloads. Default: auto (OpenRouter `anthropic/*` models).
- `supportsLongPromptCacheRetention` — host honors `prompt_cache_retention: "24h"` on the Responses API. Default: auto (api.openai.com).
- `supportsImageDetailOriginal` — allow the Responses API's nonstandard `detail: "original"` image
mode where the endpoint supports it.
- `extraBody` — extra top-level fields merged into every request body (gateway hints, controller selectors, etc.).
Reasoning / thinking:
@@ -608,11 +526,22 @@ Gateway routing (only applied when `baseUrl` matches the gateway):
- `openRouterRouting.only` / `openRouterRouting.order` — provider routing on `openrouter.ai` (see <https://openrouter.ai/docs/provider-routing>).
- `vercelGatewayRouting.only` / `vercelGatewayRouting.order` — provider routing on `ai-gateway.vercel.sh` (see <https://vercel.com/docs/ai-gateway/models-and-providers/provider-options>).
Provider-level `compat` is the baseline; per-model `compat` is deep-merged on top, with `openRouterRouting`, `vercelGatewayRouting`, and `extraBody` merged as nested objects.
Provider-level `compat` is the baseline; per-model `compat` is deep-merged on top, with
`openRouterRouting`, `vercelGatewayRouting`, `extraBody`, and `whenThinking` merged as nested objects.
### Anthropic compatibility (`anthropic-messages`)
For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a top-level provider field (see below) plus two Anthropic-side flags in the same `compat` slot — `requiresToolResultId` (non-standard `id` alias on `tool_result` blocks for Z.AI-style proxies) and `replayUnsignedThinking` (replay unsigned thinking blocks as native thinking instead of demoting them to text); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`, `supportsForcedToolChoice`, `supportsSamplingParams`, `escapeBuiltinToolNames`) are set by built-in catalog metadata and are not user-configurable from `models.yml`.
For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape
(`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a
top-level provider field plus `requiresToolResultId`, `replayUnsignedThinking`,
`supportsEagerToolInputStreaming`, and `allowAnthropicHeaderOverrides` in `compat`. Other
Anthropic-side knobs are supplied by built-in catalog metadata and are not configurable here.
### Bedrock compatibility (`bedrock-converse-stream`)
The same `compat` slot accepts `promptCacheMode` (`none`, `automatic`, or `explicit`),
`supportsLongPromptCacheRetention`, `promptCacheMinimumTokens`, and
`promptCacheMaximumCheckpoints` for Bedrock models.
### Strict tool schemas (`disableStrictTools`)