feat: implemented long-context pricing and configuration support

- Added long-context pricing tiers and billing policies for subscription Codex models in the catalog.
- Introduced the `extendedContext` configuration setting to control premium long-context windows.
- Implemented runtime policy refresh and model re-binding when context settings change.
- Added comprehensive unit tests for pricing tiers, context capping, and policy toggling behavior.
This commit is contained in:
can1357
2026-08-20 02:47:32 +02:00
parent 577c0d795f
commit 6d2bae2c41
11 changed files with 2842 additions and 899 deletions
+1
View File
@@ -6,6 +6,7 @@
- Added `qwenTemplateReasoningEffort` to the `models.yml` `compat` schema, so the auto-enabled Qwen 3.8+ template effort dialect (`chat_template_kwargs.reasoning_effort`) can be switched off per provider/model for strict local servers that reject unknown `chat_template_kwargs`.
- Added `tokenizer` to custom model and `modelOverrides` configuration. It overrides the catalog-resolved local tokenizer family for a model when a proxy serves a known model id with a different tokenizer.
- Added `extendedContext` setting (`/settings` → Context → General, default on). When off, models with a premium long-context price tier (OpenAI GPT-5.6 Sol/Terra/Luna bill 2x input / 1.5x output above 272K input tokens, on both the API and subscription Codex) are capped at the standard-pricing threshold — they appear as 272K again and compaction fires before a request crosses into premium billing. Toggling mid-session re-clamps or restores the active model's window immediately. Anthropic Claude 4.6+ serves its full 1M window at standard pricing, so no Anthropic model is affected.
### Changed