Commit Graph

83 Commits

Author SHA1 Message Date
can1357 6d2bae2c41 feat: implemented long-context pricing and configuration support
- Added long-context pricing tiers and billing policies for subscription Codex models in the catalog.
- Introduced the `extendedContext` configuration setting to control premium long-context windows.
- Implemented runtime policy refresh and model re-binding when context settings change.
- Added comprehensive unit tests for pricing tiers, context capping, and policy toggling behavior.
2026-08-20 02:47:32 +02:00
can1357 d8c5659d9a feat(catalog): updated context window floor and pricing parameters for gpt models
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
2026-08-17 10:47:30 +03:00
can1357 5fde7547f7 Merge PR #8244: fix(catalog): omit forced tool choice for go responses (@roboomp)
# Conflicts:
#	packages/catalog/src/models.json
2026-08-16 02:44:24 +02:00
can1357 a8b01fb560 Merge PR #8614: fix(catalog): price Codex Daybreak aliases (@SJY051) 2026-08-16 02:03:14 +02:00
can1357 f1095fca75 Merge PR #7454: feat(catalog): route paid xAI through Responses like SuperGrok (@geraint0923) 2026-08-16 01:17:07 +02:00
ASQi 682680e186 fix(catalog): price Codex Daybreak aliases 2026-08-15 15:06:29 +09:00
Yang Yang 02eaee09bd style(catalog): sort generated-policies imports after rebase
Keep applyXaiResponsesThinkingPolicy in the existing openai-compat
import so biome organizeImports stays clean on origin/main.
2026-08-14 22:04:53 -07:00
Yang Yang 76faddf886 fix(catalog): omit reasoningEffortMap on no-dial xAI rows
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
2026-08-14 22:02:52 -07:00
Yang Yang b49b5b88d2 fix(catalog): strip stale xAI Responses effort dials from generated rows
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
2026-08-14 22:02:07 -07:00
oldschoola e49ee4b4e2 feat(catalog): add GLM-5.3 support with uniform low/high/max effort ladder and mandatory thinking
GLM-5.3 introduces three key API changes from GLM-5.2:
- Uniform wire-exact low/high/max reasoning_effort ladder on every host
  (replacing GLM-5.2's host-specific dialects)
- Thinking can no longer be disabled (thinking.type must always be "enabled")
- Default effort is max

Changes:
- Add isGlm53ReasoningEffortModelId classifier (>=5.3, base/air/turbo, non-vision)
- getModelDefinedEfforts: GLM-5.3 returns LOW_HIGH_MAX uniformly
- impliesMandatoryReasoning: GLM-5.3 floors thinking-off to lowest effort
- deriveThinking/fillThinkingWireDefaults: defaultLevel=max for GLM-5.3
- generated-policies: pin glm-5.3 to 1M context (zai + zhipu-coding-plan)
- descriptors: zai defaultModel -> glm-5.3
- generate-models: curated seed (glm-5.3 is live but not in /models discovery)
- models.json: bundled glm-5.3 entry
- Tests: catalog thinking-metadata + AI wire-mapping (5 new tests)
2026-08-13 23:18:17 -07:00
pickpocket e7e280a6fb fix(catalog): complete GPT-5.6 off and pricing support
(cherry picked from commit fc034d61ff69e8ad1870f674c72ff3f761741862)
2026-08-13 02:00:59 +02:00
pickpocket 37af70b086 fix(catalog): classify Codex Daybreak aliases
(cherry picked from commit eba697926e9fc030d357d170e47856f9e3141dcc)
2026-08-13 02:00:59 +02:00
pickpocket 685055356a feat: add OpenAI Daybreak model support
(cherry picked from commit 889b55bbca0309e9eee6ce0c4287659bfc4fccb3)
2026-08-13 02:00:59 +02:00
roboomp f13ee010d0 fix(catalog): omitted forced tool choice for go responses
Applied OpenCode Go DeepSeek tool-choice compat to both OpenAI APIs, so the Responses route drops named selectors while preserving tools.

Added generated-policy and request-payload regression coverage.

Fixes #8243
2026-08-11 11:17:36 +00:00
roboomp 4d2c6e37f1 fix(catalog): preserved token plan preview vision
Applied curated Alibaba Token Plan seeds after generic models.dev fallback so bundled capabilities cannot be overwritten by incomplete upstream metadata.
2026-08-08 14:45:44 +00:00
roboomp 155fdaedba fix(catalog): corrected qwen3.8 max discovery metadata
Curated reasoning, multimodal input, context limits, and the provider-specific effort ladder for the discovered Alibaba Token Plan model.

Fixes #8019
2026-08-08 14:37:24 +00:00
can1357 e06ccbd907 Merge PR #7080: fix(ai): add authenticated Bedrock Mantle routing (@anatoli-tsinovoy)
# Conflicts:
#	packages/ai/src/registry/registry.ts
#	packages/catalog/scripts/generated-policies.ts
#	packages/catalog/src/models.json
2026-08-03 14:36:52 +02:00
can1357 e154c8ed0b fix(catalog): sourced antigravity claude pricing from google vertex
- Claude ids now alias to google-vertex suffixed entries (claude-opus-4-6@default etc.) so Antigravity follows Google's price if it diverges from Anthropic's list price; plain-id anthropic lookup remains as dangling-alias fallback.
- Regenerated models.json via gen:models.
2026-08-01 21:27:51 +02:00
can1357 d0d15f1a55 fix(catalog): implemented fallback pricing for unpriced antigravity models
- Added `applyAntigravityPricingFallback` to back-fill unpriced `google-antigravity` models using first-party list prices from `google` and `anthropic`.
- Mapped Antigravity model IDs to preview-id aliases (e.g., `gemini-3.1-pro` to `gemini-3.1-pro-preview`) for correct pricing resolution.
- Added test coverage in `packages/catalog/test/generated-policies.test.ts` verifying fallback pricing and preservation of billable costs.
2026-08-01 21:23:17 +02:00
can1357 4428110979 Merge PR #7267: fix(catalog): inherit base model limits for ollama tag variants (@roboomp) 2026-08-01 20:14:39 +02:00
can1357 09a7c86563 fix(catalog): unioned codex generator discovery across all oauth accounts
- Rewrote fetchCodexDiscoveryModels to resolve every stored openai-codex
  OAuth account via getOAuthAccesses and reuse openaiCodexModelManagerOptions'
  tested union/fail-closed path, so a single narrow account can no longer
  authoritatively wipe sibling-account models from the bundle.
- Restored openai-codex gpt-5.4, gpt-5.6-sol, and gpt-5.3-codex-spark bundle
  entries (and gpt-5.5's contextPromotionTarget) dropped by the previous
  single-account regen.
- Updated the live Codex image tool-result tests off the retired
  gpt-5.2-codex id to gpt-5.5.

Fixes #6265
2026-08-01 17:39:43 +02:00
roboomp c00790fa3a fix(catalog): cap ollama cloud deepseek-v4 output at 65536
Ollama Cloud's deepseek-v4-pro and deepseek-v4-flash deployments reject any
output budget above 65536 with HTTP 400, despite advertising a 1M context /
384K output (ollama/ollama#16890). Ollama's /api/show never reports this cap,
so the catalog left the base models at the full context window and the dated
tag deepseek-v4-flash:0731 at a stale 8192 fallback. Pin these ids (base plus
tag variants) to min(contextWindow, 65536) at both runtime discovery and
generation; other cloud models keep their discovered limits.

Fixes #7266
2026-08-01 15:30:17 +00:00
Anatoli Tsinovoy e5541a577a fix(ai): address Bedrock Mantle review feedback 2026-08-01 12:29:07 +03:00
can1357 71c755d5f6 feat: introduced aiand provider registry entry and model catalog
- Implement the ai& provider registry entry with API-key authentication and login support.
- Add model descriptors, static model seeding, and openai-compatible model discovery for the ai& provider.
- Update the model catalog with ai& provider models, pricing, and updated provider model names.
- Add unit tests for the ai& provider environment resolution, metadata, and dynamic model mapping.
2026-08-01 08:30:37 +02:00
can1357 c565e6fb44 feat(catalog): migrated model catalog endpoints to stencil.so with ETag and ZStd
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
2026-07-31 20:38:49 +02:00
can1357 685311807c fix(catalog,ai): seeded GMI Cloud default model into bundled catalog
- Added GMI_CLOUD_STATIC_MODELS seed wired into gen:models so a fresh
  install resolves deepseek-ai/DeepSeek-V4-Flash synchronously at boot,
  before async /v1/models discovery fires; live discovery stays
  authoritative and replaces the seed.
- Regenerated models.json with the seeded gmi-cloud slice.
- Dropped the GeneratedProvider cast now that gmi-cloud is bundled.
- Moved gmi-cloud to the end of the /login provider order.
- Moved changelog entries from released 17.0.3 sections to Unreleased.
- Added a regression test asserting the seed covers the descriptor's
  defaultModel.
2026-07-31 18:40:09 +02:00
Anatoli Tsinovoy ef6d4fb119 fix(ai): authenticate Bedrock Mantle responses 2026-07-30 15:21:52 +03:00
roboomp b09b82946b fix(catalog): derive kimi-code output caps per family
kimi-code discovery mapModel and the bundled catalog hardcoded
maxTokens=32000 for every model, truncating k3/k3-256k output at ~4x
below their real 131072 ceiling and kimi-for-coding[-highspeed] below
32768. Add kimiCodeMaxTokens() (k3* -> 131072, kimi-for-coding* ->
32768, else fallback), wire it into the discovery mapper and the
generator cap policy, and correct the bundled kimi-code entries.

Fixes #6711
2026-07-26 14:41:32 +00:00
k1riiiii a481a89e62 feat(catalog): add Claude Opus 5 entries for Amazon Bedrock
Changes:
- Add Amazon Bedrock catalog entries for Claude Opus 5
- Exclude unsupported jp. inference profiles during generation
- Add Bedrock prompt cache compatibility info, AWS model card sources, and regression tests

Reason / background:
- Incorporate the new model info from models.dev and align with the official AWS spec

Scope of impact:
- Model catalog generation, Bedrock compatibility info, catalog tests
2026-07-25 10:12:32 +00:00
can1357 1e1d2e91ea style: applied biome formatting 2026-07-24 02:27:35 +02:00
can1357 dcd902f3bf fix: scoped empty-success catalog authority to alibaba-token-plan
fetchProviderModelsFromCatalog returning succeeded=true for an empty
discovery made every dynamicModelsAuthoritative provider drop its
models.dev and previous-snapshot rows after a flaky empty-but-200
response. Restored the fetched-models requirement for other providers;
only alibaba-token-plan treats an empty success as authoritative so the
subscribed-edition allowlist is not widened by the curated seed. Also
dropped the stray trailing newline in models.json that the generator
does not emit.
2026-07-24 02:25:25 +02:00
Brent e184fe8f8d feat: add native Alibaba Token Plan provider 2026-07-23 23:04:00 +00:00
can1357 a511a78fc9 fix(catalog/discovery): removed context window floor for gpt-5.6 skus
- Removed the logic that forced a minimum context window of 372k for GPT-5.6 SKUs when upstream actively reported lower values.
- Updated Codex discovery to treat 372k as a fallback only when upstream omits the context window.

Fixes #6371
2026-07-23 22:08:13 +02:00
Brent 3b1ba087fe feat: add native Meta Model API provider 2026-07-23 14:15:40 +00:00
roboomp 930bb33f4a fix(catalog): floored gpt-5.6 codex context window at 372k
Codex discovery under-reports the gpt-5.6 sol/terra/luna window: some
accounts omit `context_window`, others actively return 272000. The #5707
`?? fallback` only fired on absence, so an actively-reported 272000 passed
through and `preferDiscoveryLimit` overwrote the bundled 372K pin at
runtime, dropping compaction from 279000 to 204000 tokens.

Treat GPT_5_6_CONTEXT_WINDOW as a floor for these SKUs via Math.max so
neither omission nor active under-report regresses the real capacity;
other models keep honoring the reported value. Corrected the stale
"omits" comments in codex.ts and generated-policies.ts.

Fixes #6259
2026-07-22 04:03:25 +00:00
can1357 ee26bec5a5 fix(catalog): drop redundant swe-1-7 static seed
main already bundles devin/swe-1-7 (discovered live, with image input);
the text-only seed wins the earlier-sources dedup in generate-models and
downgrades the bundled entry to text-only. Carry-over from the previous
snapshot already preserves the model across keyless regens.
2026-07-20 22:50:16 +02:00
can1357 024497995c Merge PR #4882: fix(catalog): collapse Devin GLM-5.2 variants so free 200K model works when quota is exhausted (@oldschoola)
# Conflicts:
#	packages/catalog/src/models.json
2026-07-20 22:50:16 +02:00
can1357 ffab5689a0 Merge PR #5497: fix(catalog): prune unsupported Codex account models (@roboomp)
# Conflicts:
#	packages/catalog/src/discovery/codex.ts
#	packages/catalog/src/models.json
#	packages/catalog/test/codex-discovery.test.ts
2026-07-18 21:02:46 +02:00
can1357 f2959f255e merge PR #5735 via eval/pr-5735: fix(catalog): sourced Umans PAYG model pricing 2026-07-17 04:40:04 +02:00
roboomp 0becfbe197 fix(catalog): sourced Umans PAYG model pricing
- Repointed the Umans models.dev descriptor to published PAYG rates.
- Backfilled authoritative discovery rows and the Flash technical alias.
- Added runtime, descriptor, and bundled catalog regression coverage.

Fixes #5733
2026-07-16 17:39:22 +00:00
roboomp 62c164d256 fix(catalog): pinned gpt-5.6 codex context window to 372k
Codex discovery falls back to DEFAULT_CONTEXT_WINDOW (272000) when upstream
omits context_window, which overwrote the previously-bundled 372000 hard
capacity for openai-codex gpt-5.6 luna/sol/terra on regen. OpenAI's Codex
model registry declares context_window = max_context_window = 372000, and a
direct Responses request with 350,317 input tokens completes, proving 272K is
not the route's hard cap.

Pinned these SKUs to 372000 in applyOpenAICatalogPolicy and refreshed the
bundled entries.

Fixes #5705
2026-07-16 13:39:04 +00:00
roboomp 51ab2fcf23 fix(catalog): made codex discovery authoritative
Replaced stale bundled OpenAI Codex entries after successful account-scoped discovery in both runtime resolution and catalog generation.

Fixes #5364
2026-07-14 19:17:43 +00:00
can1357 faa70100ea feat: enabled openai reasoning mode and integrated new model catalog
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
2026-07-09 22:20:32 +02:00
oldschoola 5ebafcd9f7 fix(catalog): collapse Devin GLM-5.2 variants so free 200K model works when quota is exhausted
Devin exposes 6 GLM-5.2 wire UIDs. Without a collapse family, all 6
appeared as separate catalog entries and the user could inadvertently
select a quota-gated variant (glm-5-2-max, glm-5-2-none) that silently
fails with 'weekly usage quota exhausted' even though the base glm-5-2
is free.

Live verification via streamDevin confirmed:
- glm-5-2 (base, 200K): FREE — works with quota exhausted
- glm-5-2-none (200K): quota-gated — fails with 'usage quota exhausted'
- glm-5-2-max (200K): quota-gated — fails with 'usage quota exhausted'
- swe-1-6, swe-1-7, kimi-k2-7: FREE — all work with quota exhausted

Changes:
- Add GLM-5.2 collapse family routing all efforts (high/xhigh) to the
  free glm-5-2 wire UID — never to the quota-gated variants
- Add GLM-5.2 1M collapse family for paid variants (glm-5-2-1m,
  glm-5-2-none-1m, glm-5-2-max-1m)
- Add swe-1-7 static fallback seed (missing from catalog, free on Devin)
- Wire seed in generate-models.ts with authoritative-discovery guard
- 3 unit tests for GLM-5.2 collapse routing
2026-07-08 19:37:12 -07:00
roboomp 32d12dde9e fix(catalog): set opencode go deepseek max_tokens
- Updated generated OpenCode Go DeepSeek V4 catalog policy to use max_tokens instead of max_completion_tokens.

- Added regression coverage for deepseek-v4-flash:xhigh tool requests carrying max_tokens and reasoning_effort:max.

Fixes #4647
2026-07-06 00:49:06 +00:00
can1357 3458b037ae chore: update tests 2026-07-05 16:53:07 +02:00
can1357 4c18cc1a1a feat(catalog): integrated baseten provider and updated model definitions
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
2026-07-03 06:04:31 +02:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 490662ab2e feat(ai/utils): increased the Gemini header runaway threshold
- Raised `GEMINI_HEADER_RUNAWAY_THRESHOLD` from 10 to 24 to avoid false-positive interrupts on legitimate, complex reasoning blocks.
- Added a regression test verifying that 10 distinct, progressing headers do not trip the detector while 24 headers still trigger it.
2026-06-30 20:27:52 +02:00
can1357 f1453d72ae refactor(catalog): restructured model generation to prune redundant compatibility fields
- Introduced a model canonicalization helper to strip redundant model compatibility fields that match defaults.
- Regenerated the models catalog JSON to eliminate over eight hundred lines of redundant compatibility specifications.
- Updated the variant collapse logic to rebuild models using the projected compatibility configurations.
- Added a missing type annotation to a test environment variable to resolve a compilation warning.
2026-06-30 20:26:37 +02:00