Commit Graph

799 Commits

Author SHA1 Message Date
can1357 6fb07028fd chore: bump version to 17.2.12 2026-08-08 20:57:55 +02:00
can1357 60d4cb997e test(catalog): pinned copilot grok-4.5 migration to the responses route
- The regenerated bundle (merged with PR #8021) now ships a
  responses-route github-copilot/grok-4.5, so the id legitimately
  resurfaces from the bundle when the migration refresh fails; the
  contract worth defending is that the stale cached completions route
  never returns and the unbundled long-context variant stays dropped.
2026-08-08 20:57:10 +02:00
can1357 bf04fbfc8d fix(catalog): routed opencode-go deepseek-v4-flash through the responses api
- The OpenCode Go gateway does not serve DSV4-Flash at
  /zen/go/v1/chat/completions; /zen/go/v1/responses works (user-verified
  against the live gateway). Added a per-id override in
  OPENCODE_GO_API_RESOLUTION so both bundled generation and the runtime
  /v1/models refresh route it to openai-responses; deepseek-v4-pro keeps
  chat completions.
- Regenerated models.json from the resolver source.
2026-08-08 20:57:10 +02:00
can1357 ac55ea7697 fix(catalog): toggle qwen3.8 max thinking on wire 2026-08-08 19:38:31 +02:00
roboomp c0eda613c8 fix(catalog): kept qwen3.8 max preview on enable_thinking
Restored the preview to its main compat so the reasoning_effort dialect is scoped to qwen3.8-max, leaving the preview ladder unchanged.
2026-08-08 14:54:29 +00:00
roboomp 4d2c6e37f1 fix(catalog): preserved token plan preview vision
Applied curated Alibaba Token Plan seeds after generic models.dev fallback so bundled capabilities cannot be overwritten by incomplete upstream metadata.
2026-08-08 14:45:44 +00:00
roboomp 155fdaedba fix(catalog): corrected qwen3.8 max discovery metadata
Curated reasoning, multimodal input, context limits, and the provider-specific effort ladder for the discovered Alibaba Token Plan model.

Fixes #8019
2026-08-08 14:37:24 +00:00
can1357 c80a531226 refactor(catalog): deduplicated openai-compatible manager builders
- Added one internal generic builder covering the repeated apiKey/baseUrl
  resolution, bundled reference map, providerId and fetchDynamicModels closure.
- Migrated openai, cerebras, novita, aimlApi, alibabaCodingPlan, venice,
  baseten and moonshot, and routed createSimpleOpenAICompletionsOptions
  through it; deleted the duplicate responses-side helper.
- Preserved fetchDynamicModels key PRESENCE per site, since the apiKey spread
  guard omits the key entirely rather than setting it undefined.
2026-08-08 06:32:00 +02:00
can1357 055a5d4f26 chore: bump version to 17.2.11 2026-08-07 23:38:40 +02:00
can1357 a9dcf0f8d1 chore: update changelogs 2026-08-07 23:38:26 +02:00
can1357 9ab6ea6c8f test(catalog): cover discovered Token Plan limits 2026-08-07 13:37:55 +02:00
can1357 ded3bbfce0 Merge PR #7849: fix(catalog): enrich Alibaba Token Plan discovered model limits (@Mustaqeem66) 2026-08-07 13:37:55 +02:00
can1357 0394bf4a29 Merge PR #7865: fix(catalog): update Devin reasoning family routing (@will-bogusz) 2026-08-07 13:37:55 +02:00
Voon Foo 61ff318e6d review: typed compat access, 0-disable lazy watchdog test, changelogs
- Replace the compat cast with an annotated CompatOf<Api> local narrowed
  via the in operator: assignment up-cast, compiler-checked field type,
  no as assertion.
- Cover the 0 sentinel end-to-end through the lazy wrapper with fake
  timers (mirrors the direct-Anthropic 0-disable test): advance 400s
  past the generic budget, assert no watchdog abort, then cancel
  cleanly.
- Add Unreleased changelog entries for pi-ai and pi-catalog with
  external attribution for #7892.
2026-08-07 15:39:04 +08:00
Voon Foo 2f24d4457e fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).

- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
  a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
  Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
  tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
  and lazy terminal errors carry the structural errorId classification
  so session auto-retry classifies stalls without text matching.
2026-08-07 14:18:46 +08:00
Will 7d6433cffe docs(catalog): correct Devin effort invariant 2026-08-06 21:45:15 -04:00
Will 1ad85b5584 fix(catalog): collapse current Devin reasoning families 2026-08-06 21:36:11 -04:00
Will 7e95b61ffb fix(catalog): route Devin GPT-5.6 fast max effort 2026-08-06 21:35:56 -04:00
Mustaqeem66 b80a888dc8 fix(catalog): enrich Alibaba Token Plan discovered model limits 2026-08-06 17:12:21 +00:00
can1357 43c1b245e7 chore: bump version to 17.2.10 2026-08-06 13:32:34 +02:00
can1357 9e738dc880 chore: reformat + rewrite changelogs 2026-08-06 13:30:08 +02:00
can1357 c5cce0f325 chore: normalized changelogs after merging 14 pull requests 2026-08-05 22:17:01 +02:00
can1357 a6918d3397 Merge PR #7669: fix(catalog): exposed low effort tier for deepseek-v4-flash (@roboomp) 2026-08-05 22:16:29 +02:00
can1357 0b0e7dd530 chore: normalized changelogs after merging 20 pull requests 2026-08-05 21:50:46 +02:00
can1357 e9888367d1 refactor: migrated packages to internal utility modules and removed external dependencies
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
2026-08-05 13:39:09 +02:00
roboomp e97d1fd21a test(catalog): removed tautological effort assertions 2026-08-05 02:34:04 +00:00
roboomp 736b496cc6 fix(catalog): exposed low effort tier for deepseek-v4-flash
DeepSeek's API accepts reasoning_effort low/high/max and only
deepseek-v4-flash supports all three tiers (V4 Pro is high/max). The
identity-derived effort fallback blanket-applied high/max to every
direct-DeepSeek reasoning model, hiding the low tier flash accepts.

Added isDeepseekV4FlashModelId and route flash to the low/high/max ladder
on every host; non-flash DeepSeek keeps high/max (high-only on OpenRouter).

Fixes #7668
2026-08-05 02:22:27 +00:00
can1357 f7f8e040ee chore: bump version to 17.2.9 2026-08-05 03:07:47 +02:00
Pete Samwel fce059930e fix(catalog): emit AWS GovCloud us-gov Bedrock Claude inference profiles
Bare anthropic.claude-* Bedrock rows already derived eu.* selectors; also
emit us-gov.* so GovCloud accounts can resolve system inference profiles
without requiring a full partition ARN.
2026-08-04 13:13:18 -05:00
can1357 003bb5548c chore: bump version to 17.2.8 2026-08-04 05:53:35 +02:00
can1357 a5090f1f81 chore: bump version to 17.2.7 2026-08-04 01:19:36 +02:00
can1357 60acbee44a docs(changelog): normalized and regenerated unreleased changelogs 2026-08-04 01:19:08 +02:00
roboomp 58ced35dc0 fix(catalog): honored disabled DeepSeek thinking
Moved the direct DeepSeek enabled toggle into the thinking-only compat variant and normalized stale cached compat metadata before request encoding.

Fixes #7559
2026-08-03 21:48:40 +00:00
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
can1357 01c1f91ff5 chore: bump version to 17.2.6 2026-08-03 16:44:19 +02:00
can1357 4ddc2d2cd4 chore: rewrite changelogs + fix stale tests 2026-08-03 15:32:19 +02:00
can1357 9fdb989c63 docs(changelog): restored released sections and normalized unreleased entries
- Union changelog merges interleaved stale pre-17.2.5 PR-branch entries into
  released sections; released bodies are restored byte-for-byte from the
  pre-merge main state.
- [Unreleased] now carries exactly the entries for PR #7080 and the nine
  merged fixes (#7495, #7460, #7466, #7468, #7473, #7481, #7477, #7368, #7453).
2026-08-03 14:51:01 +02:00
can1357 35ed5db9e1 Merge PR #7473: fix(catalog): use live Copilot default-tier prices (@roboomp) 2026-08-03 14:46:23 +02:00
can1357 11559395c2 Merge PR #7468: fix(catalog): add reasoning config for deepseek-v4 family in alibaba-token-plan (@21307369) 2026-08-03 14:46:23 +02:00
can1357 c48376d8f0 fix(catalog): made bedrock-mantle dynamic discovery authoritative
- Account-scoped bearer /v1/models responses now replace the static seed
  instead of merging, so models disabled for the account are not selectable.
- Extended the catalog regression to run a real online refresh and assert
  the static seeds are pruned to the fetched IDs.
2026-08-03 14:36:57 +02:00
can1357 e06ccbd907 Merge PR #7080: fix(ai): add authenticated Bedrock Mantle routing (@anatoli-tsinovoy)
# Conflicts:
#	packages/ai/src/registry/registry.ts
#	packages/catalog/scripts/generated-policies.ts
#	packages/catalog/src/models.json
2026-08-03 14:36:52 +02:00
roboomp 27c1d6e8d4 fix(catalog): used live copilot default-tier prices
Applied GitHub Copilot's discovered default token-price tier to base models while preserving the provider fallback for unreported cache-write costs. Added regression coverage for GPT-5.6 Luna base and long-context pricing.

Fixes #7471
2026-08-03 08:38:14 +00:00
lsmir2 7beee683b9 fix(catalog): add reasoning config for deepseek-v4 family in alibaba-token-plan
Dynamically discovered deepseek-v4* models (e.g. deepseek-v4-flash-0731)
now receive reasoning: true and [high, max] thinking efforts via prefix
matching in the discovery mapper.
2026-08-03 15:55:14 +08:00
can1357 c53b85aaf4 chore: bump version to 17.2.5 2026-08-03 05:53:12 +02:00
can1357 63b07f8ec0 chore: rewrite changelogs 2026-08-03 05:52:53 +02:00
can1357 a7f3bc2a17 chore: normalized changelog entries after merges 2026-08-02 21:01:11 +02:00
can1357 3177f6bdf7 test(catalog): cover DeepSeek policy without bundled models 2026-08-02 20:55:16 +02:00
can1357 46ca4c6095 Merge PR #7317: fix(catalog): downgrade forced tool choice for deepseek reasoning models (@roboomp) 2026-08-02 20:55:16 +02:00
can1357 990ec2b0e8 fix(catalog): exclude Token Plan speech recognition models 2026-08-02 20:53:35 +02:00
roboomp ac403dbb71 fix(catalog): excluded token plan embedding models
Filtered text-embedding model IDs from authoritative Alibaba Token Plan chat discovery.

Covered text-embedding-v4 alongside the existing media-only discovery fixtures.

Fixes #7391
2026-08-02 16:16:24 +00:00