Commit Graph

807 Commits

Author SHA1 Message Date
can1357 e5ebb2aee0 chore: bump version to 17.2.14 2026-08-11 20:43:02 +02:00
can1357 2157becbe9 chore: bump version to 17.2.13 2026-08-11 16:03:05 +02:00
can1357 b524dfe36f refactor: standardized outbound User-Agent headers on shared utility constant
- Define a centralized `USER_AGENT` constant in `@oh-my-pi/pi-utils` formatted as `omp/<version>`.
- Replace hardcoded and platform-specific user agent strings across AI providers, catalog scrapers, tools, and search providers with the unified `USER_AGENT`.
- Add unit tests for update-cli binary release distribution gating.
2026-08-11 15:38:32 +02:00
can1357 64baa7c1bd chore(format): applied biome formatting and removed dead code from merged prs 2026-08-11 15:14:15 +02:00
can1357 1340ce6d18 Merge PR #8200: fix(catalog): Fix reasoning levels of GLM-5.2 models; add Baseten GLM 5.2 Fast (@jcfrancisco) 2026-08-11 15:09:34 +02:00
Carlo Francisco aecd4c76e5 fix(catalog): mark Baseten GLM-5.2-Fast as reasoning with high/max effort
Adds zai-org/GLM-5.2-Fast to the Baseten reasoning allowlists, matching
the sibling zai-org/GLM-5.2. Also fixes parseGlmModel to handle uppercase
GLM model IDs (used by Baseten, CoreWeave, HuggingFace, etc.), so the
identity deriver correctly classifies them as GLM-5.2 reasoning models
during catalog generation — previously the case-sensitive regex caused
rebakeModelThinking to fall back to the generic effort ladder.

Regenerated models.json with a live Baseten API key: GLM-5.2 and
GLM-5.2-Fast now bundle reasoning:true with the correct high/max effort
ladder, and other uppercase GLM-5.2 resellers (CoreWeave, HuggingFace,
Synthetic, Together, Wafer) get the corrected minimal..max ladder.
2026-08-10 22:30:11 -04:00
Chen Buskilla 7d7964c14b address review: fix test, changelog and models.json newline
- Update packages/catalog/test/meta-provider.test.ts to expect
  three META_MUSE_STATIC_MODELS entries (1.1, 1.2, contributor)
  with input ["text","image"] — fixes blocking failure.

- Add ## [Unreleased] entry in packages/catalog/CHANGELOG.md per
  AGENTS.md.

- Remove trailing newline from models.json to match
  generate-models.ts (Bun.write without \n).

Co-authored-by: roboomp <roboomp@users.noreply.github.com>
2026-08-09 18:50:43 +03:00
Chen Buskilla c8007f18cd fix(catalog): mark muse-spark-1.2 as image capable
Meta Muse Spark 1.2 and its contributor variant support vision
inputs (text, image) like 1.1, but were missing from
META_MUSE_STATIC_MODELS and bundled models.json. Discovery
without a bundled reference falls back to text-only with zero
cost, so omp models listed them as not image enabled.

Add both models to META_MUSE_STATIC_MODELS with the same
cost/thinking/vision metadata as 1.1 (contributor uses its
discounted pricing) and regenerate the bundled catalog.
2026-08-09 18:45:48 +03:00
can1357 6fb07028fd chore: bump version to 17.2.12 2026-08-08 20:57:55 +02:00
can1357 60d4cb997e test(catalog): pinned copilot grok-4.5 migration to the responses route
- The regenerated bundle (merged with PR #8021) now ships a
  responses-route github-copilot/grok-4.5, so the id legitimately
  resurfaces from the bundle when the migration refresh fails; the
  contract worth defending is that the stale cached completions route
  never returns and the unbundled long-context variant stays dropped.
2026-08-08 20:57:10 +02:00
can1357 bf04fbfc8d fix(catalog): routed opencode-go deepseek-v4-flash through the responses api
- The OpenCode Go gateway does not serve DSV4-Flash at
  /zen/go/v1/chat/completions; /zen/go/v1/responses works (user-verified
  against the live gateway). Added a per-id override in
  OPENCODE_GO_API_RESOLUTION so both bundled generation and the runtime
  /v1/models refresh route it to openai-responses; deepseek-v4-pro keeps
  chat completions.
- Regenerated models.json from the resolver source.
2026-08-08 20:57:10 +02:00
can1357 ac55ea7697 fix(catalog): toggle qwen3.8 max thinking on wire 2026-08-08 19:38:31 +02:00
roboomp c0eda613c8 fix(catalog): kept qwen3.8 max preview on enable_thinking
Restored the preview to its main compat so the reasoning_effort dialect is scoped to qwen3.8-max, leaving the preview ladder unchanged.
2026-08-08 14:54:29 +00:00
roboomp 4d2c6e37f1 fix(catalog): preserved token plan preview vision
Applied curated Alibaba Token Plan seeds after generic models.dev fallback so bundled capabilities cannot be overwritten by incomplete upstream metadata.
2026-08-08 14:45:44 +00:00
roboomp 155fdaedba fix(catalog): corrected qwen3.8 max discovery metadata
Curated reasoning, multimodal input, context limits, and the provider-specific effort ladder for the discovered Alibaba Token Plan model.

Fixes #8019
2026-08-08 14:37:24 +00:00
can1357 c80a531226 refactor(catalog): deduplicated openai-compatible manager builders
- Added one internal generic builder covering the repeated apiKey/baseUrl
  resolution, bundled reference map, providerId and fetchDynamicModels closure.
- Migrated openai, cerebras, novita, aimlApi, alibabaCodingPlan, venice,
  baseten and moonshot, and routed createSimpleOpenAICompletionsOptions
  through it; deleted the duplicate responses-side helper.
- Preserved fetchDynamicModels key PRESENCE per site, since the apiKey spread
  guard omits the key entirely rather than setting it undefined.
2026-08-08 06:32:00 +02:00
can1357 055a5d4f26 chore: bump version to 17.2.11 2026-08-07 23:38:40 +02:00
can1357 a9dcf0f8d1 chore: update changelogs 2026-08-07 23:38:26 +02:00
can1357 9ab6ea6c8f test(catalog): cover discovered Token Plan limits 2026-08-07 13:37:55 +02:00
can1357 ded3bbfce0 Merge PR #7849: fix(catalog): enrich Alibaba Token Plan discovered model limits (@Mustaqeem66) 2026-08-07 13:37:55 +02:00
can1357 0394bf4a29 Merge PR #7865: fix(catalog): update Devin reasoning family routing (@will-bogusz) 2026-08-07 13:37:55 +02:00
Voon Foo 61ff318e6d review: typed compat access, 0-disable lazy watchdog test, changelogs
- Replace the compat cast with an annotated CompatOf<Api> local narrowed
  via the in operator: assignment up-cast, compiler-checked field type,
  no as assertion.
- Cover the 0 sentinel end-to-end through the lazy wrapper with fake
  timers (mirrors the direct-Anthropic 0-disable test): advance 400s
  past the generic budget, assert no watchdog abort, then cancel
  cleanly.
- Add Unreleased changelog entries for pi-ai and pi-catalog with
  external attribution for #7892.
2026-08-07 15:39:04 +08:00
Voon Foo 2f24d4457e fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).

- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
  a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
  Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
  tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
  and lazy terminal errors carry the structural errorId classification
  so session auto-retry classifies stalls without text matching.
2026-08-07 14:18:46 +08:00
Will 7d6433cffe docs(catalog): correct Devin effort invariant 2026-08-06 21:45:15 -04:00
Will 1ad85b5584 fix(catalog): collapse current Devin reasoning families 2026-08-06 21:36:11 -04:00
Will 7e95b61ffb fix(catalog): route Devin GPT-5.6 fast max effort 2026-08-06 21:35:56 -04:00
Mustaqeem66 b80a888dc8 fix(catalog): enrich Alibaba Token Plan discovered model limits 2026-08-06 17:12:21 +00:00
can1357 43c1b245e7 chore: bump version to 17.2.10 2026-08-06 13:32:34 +02:00
can1357 9e738dc880 chore: reformat + rewrite changelogs 2026-08-06 13:30:08 +02:00
can1357 c5cce0f325 chore: normalized changelogs after merging 14 pull requests 2026-08-05 22:17:01 +02:00
can1357 a6918d3397 Merge PR #7669: fix(catalog): exposed low effort tier for deepseek-v4-flash (@roboomp) 2026-08-05 22:16:29 +02:00
can1357 0b0e7dd530 chore: normalized changelogs after merging 20 pull requests 2026-08-05 21:50:46 +02:00
can1357 e9888367d1 refactor: migrated packages to internal utility modules and removed external dependencies
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
2026-08-05 13:39:09 +02:00
roboomp e97d1fd21a test(catalog): removed tautological effort assertions 2026-08-05 02:34:04 +00:00
roboomp 736b496cc6 fix(catalog): exposed low effort tier for deepseek-v4-flash
DeepSeek's API accepts reasoning_effort low/high/max and only
deepseek-v4-flash supports all three tiers (V4 Pro is high/max). The
identity-derived effort fallback blanket-applied high/max to every
direct-DeepSeek reasoning model, hiding the low tier flash accepts.

Added isDeepseekV4FlashModelId and route flash to the low/high/max ladder
on every host; non-flash DeepSeek keeps high/max (high-only on OpenRouter).

Fixes #7668
2026-08-05 02:22:27 +00:00
can1357 f7f8e040ee chore: bump version to 17.2.9 2026-08-05 03:07:47 +02:00
Pete Samwel fce059930e fix(catalog): emit AWS GovCloud us-gov Bedrock Claude inference profiles
Bare anthropic.claude-* Bedrock rows already derived eu.* selectors; also
emit us-gov.* so GovCloud accounts can resolve system inference profiles
without requiring a full partition ARN.
2026-08-04 13:13:18 -05:00
can1357 003bb5548c chore: bump version to 17.2.8 2026-08-04 05:53:35 +02:00
can1357 a5090f1f81 chore: bump version to 17.2.7 2026-08-04 01:19:36 +02:00
can1357 60acbee44a docs(changelog): normalized and regenerated unreleased changelogs 2026-08-04 01:19:08 +02:00
roboomp 58ced35dc0 fix(catalog): honored disabled DeepSeek thinking
Moved the direct DeepSeek enabled toggle into the thinking-only compat variant and normalized stale cached compat metadata before request encoding.

Fixes #7559
2026-08-03 21:48:40 +00:00
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
can1357 01c1f91ff5 chore: bump version to 17.2.6 2026-08-03 16:44:19 +02:00
can1357 4ddc2d2cd4 chore: rewrite changelogs + fix stale tests 2026-08-03 15:32:19 +02:00
can1357 9fdb989c63 docs(changelog): restored released sections and normalized unreleased entries
- Union changelog merges interleaved stale pre-17.2.5 PR-branch entries into
  released sections; released bodies are restored byte-for-byte from the
  pre-merge main state.
- [Unreleased] now carries exactly the entries for PR #7080 and the nine
  merged fixes (#7495, #7460, #7466, #7468, #7473, #7481, #7477, #7368, #7453).
2026-08-03 14:51:01 +02:00
can1357 35ed5db9e1 Merge PR #7473: fix(catalog): use live Copilot default-tier prices (@roboomp) 2026-08-03 14:46:23 +02:00
can1357 11559395c2 Merge PR #7468: fix(catalog): add reasoning config for deepseek-v4 family in alibaba-token-plan (@21307369) 2026-08-03 14:46:23 +02:00
can1357 c48376d8f0 fix(catalog): made bedrock-mantle dynamic discovery authoritative
- Account-scoped bearer /v1/models responses now replace the static seed
  instead of merging, so models disabled for the account are not selectable.
- Extended the catalog regression to run a real online refresh and assert
  the static seeds are pruned to the fetched IDs.
2026-08-03 14:36:57 +02:00
can1357 e06ccbd907 Merge PR #7080: fix(ai): add authenticated Bedrock Mantle routing (@anatoli-tsinovoy)
# Conflicts:
#	packages/ai/src/registry/registry.ts
#	packages/catalog/scripts/generated-policies.ts
#	packages/catalog/src/models.json
2026-08-03 14:36:52 +02:00
roboomp 27c1d6e8d4 fix(catalog): used live copilot default-tier prices
Applied GitHub Copilot's discovered default token-price tier to base models while preserving the provider fallback for unreported cache-write costs. Added regression coverage for GPT-5.6 Luna base and long-context pricing.

Fixes #7471
2026-08-03 08:38:14 +00:00