Commit Graph

365 Commits

Author SHA1 Message Date
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
roboomp 9599689af3 fix(compaction): routed summary oneshots through provider cap
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).

Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.

Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.

Fixes #3749
2026-06-28 21:11:48 +00:00
roboomp edaaec398c fix(task): routed side requests through provider cap
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.

Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
2026-06-28 20:57:13 +00:00
ben 23dc53bdc9 fix(agent): preserve stream error detail 2026-06-28 17:31:06 +08:00
ben dc17aead5f fix(agent): narrow stream recovery review fixes 2026-06-28 16:18:46 +08:00
ben cb4e68888f fix(agent): recover completed tools after stream read errors 2026-06-28 16:18:46 +08:00
can1357 08e1ea7275 feat(agent): strengthened v2 stream integrity and replay logic
- Implement provider-native replay logic to enable reuse of remote compaction data across compatible models.
- Enhance compaction logic to re-expand and locally summarize remote history when provider-native replay is unavailable.
- Update OpenAI request setup to include session and routing identifiers for improved traceability.
- Refine token estimation for image content during truncation to ensure more accurate budget management.
- Update agent-session to resolve compaction model candidates before persistence, ensuring authentication availability.
2026-06-28 09:52:44 +02:00
can1357 102d6d54ad feat: implemented v2 streaming remote compaction for model history state
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
2026-06-28 07:27:02 +02:00
can1357 0ce330ab79 feat(session): implemented persistent session title tracking
- Added a fixed-width title slot system to serialize and persist session titles across physical files and backend storage.
- Integrated automated session title updates triggered by todo replan operations using conversation history context.
- Extended storage interfaces across memory, file-system, Redis, and SQL backends to support independent title metadata updates.
- Implemented title metadata parsing within session loaders and list utilities to ensure accurate retrieval and display.
2026-06-28 07:27:00 +02:00
roboomp ae87caeb2a fix(coding-agent): split per-file TTSR digests for multi-file edits
PR #3648 review (codex): the first iteration emitted every envelope
path alongside one combined digest, so a multi-file payload that added
`: any` to a README.md hunk and merely touched src/ok.ts would surface
a *.ts path in the TTSR match context and trip the bundled
tool:edit(*.ts) ts-no-any rule on text that belonged to the Markdown
hunk — aborting valid edits under interruptMode:always.

Add a per-file matcherEntries(args) hook on AgentTool / EditStreaming-
Strategy returning [{ path, digest }] entries, one per touched file
(same-path sections/hunks merged):
- replace / patch: one entry from the top-level path + matcherDigest
- hashline: regex-split by [path#TAG] section, body added-lines per
  entry (tolerant of streaming partial payloads)
- apply_patch: expandApplyPatchToPreviewEntries grouped by path

AgentSession.#checkTtsrStream / #checkTtsrAstStream now prefer
matcherEntries and iterate per-file with isolated filePaths + streamKey,
so each file's buffer and repeat-tracking are independent. Tools
without matcherEntries keep the existing combined matcherDigest +
matcherPaths path.
2026-06-27 11:22:32 +00:00
roboomp 33e425594e fix(coding-agent): route hashline + apply_patch edit paths into TTSR
AgentSession's TTSR match context only scanned top-level path/paths
arguments, so hashline and apply_patch edit streams (whose only target
path lives inside the wire payload — section headers or envelope
markers — not as a top-level argument) arrived without any filePaths
and silently skipped path-scoped rules like the bundled ts-no-any
(scope: tool:edit(*.ts)).

Add an optional AgentTool.matcherPaths(args) hook, companion to the
existing matcherDigest(args), so tools whose wire grammar embeds paths
can surface them. Implement on each edit streaming strategy:
- replace / patch: top-level path
- hashline: parse [path#TAG] (and tag-less [path]) section headers
  tolerant of streaming partial payloads
- apply_patch: parse *** Add/Update/Delete File: markers, also tolerant
  of pre-End-Patch buffers

AgentSession.#getTtsrToolMatchContext consults tool.matcherPaths first,
normalising its output through the existing path-candidate helper, and
falls back to the generic top-level argument scan for tools that don't
implement it.

Fixes #3646
2026-06-27 11:11:50 +00:00
can1357 cfa0cd84ba feat(agent): resolved partial json leakage by enforcing streaming cleanup
- Update scrubPartialJson to utilize clearStreamingPartialJson for consistent tool-call cleanup.
- Adjust execution order in streamProxy to ensure partial error messages are finalized before scrubbing.
- Remove redundant test expectation comment regarding partialJson leakage.
2026-06-27 12:26:11 +02:00
can1357 357c29224d feat: implemented symbol-based streaming state for isolated metadata
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
2026-06-27 12:07:08 +02:00
can1357 c3f7e849e5 refactor: centralized AI error handling into a dedicated module
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
2026-06-27 10:44:13 +02:00
can1357 cc45148cb1 refactor(ai): replace delete with stripVariant helper and use performance.now() for timing
`delete` on object properties degrades V8 hidden class optimization; the new `stripVariant` util sets the property to `undefined` instead, keeping the object shape stable.

`performance.now()` is used in place of `Date.now()` for duration and TTFT measurements to get a monotonic, high-resolution clock that is unaffected by system clock adjustments.
2026-06-27 09:08:37 +02:00
can1357 053da98ddc feat: removed pi dialect and its associated infrastructure
- Removed the pi dialect implementation and associated source files.
- Updated dialect resolution, factory registration, and type definitions to exclude pi.
- Cleaned up settings schema and user options to remove pi-related configurations.
- Deleted corresponding test suites covering pi dialect functionality, in-band tools, and examples.
2026-06-27 08:37:25 +02:00
can1357 0b8206a46d fix(compaction): preserve provider defaults for remote compaction 2026-06-27 01:39:34 +02:00
can1357 14cc9cba0f Merge PR #3106: fix(compaction): enable custom provider remote compaction (@roboomp) 2026-06-27 01:39:34 +02:00
can1357 3e5360ca4c Merge PR #3060: feat(provider): add GitLab Duo Agent provider (@jiwangyihao) 2026-06-27 01:39:31 +02:00
can1357 eb4a02433c refactor: consolidated json parsing and stream utilities
- Centralized JSON parsing and stream processing logic by moving utilities from `packages/ai` to the shared `@oh-my-pi/pi-utils` package.
- Standardized import paths for JSON parsing and streaming across the agent, ai, and coding-agent packages.
- Refactored SSE stream handling to use consolidated `parseStreamingJson` logic and introduced robust error recovery for malformed container-shaped tail events.
- Cleaned up legacy bundled registry references and updated related module exports and tests to reflect the new utility structure.
2026-06-27 00:11:42 +02:00
can1357 1ed343d9a4 Merge PR #3595: fix(agent): stop replaying provider refusals (@roboomp) 2026-06-26 23:27:40 +02:00
can1357 2899ce1bc2 Merge PR #3328: fix(proxy): scrub transient partialJson from final tool calls (@roboomp) 2026-06-26 23:27:38 +02:00
roboomp 22f43b26ae fix(agent): preserved refusals in summaries
Scoped provider-refusal filtering to live replay so compaction and snapcompact summaries retain the refused turn while outbound provider context still drops the refusal.

Fixes #3592
2026-06-26 18:04:34 +00:00
roboomp d68980fe58 fix(agent): stopped replaying provider refusals
API-level refusals now stay visible as terminal errors without being sent back as assistant dialogue on the next provider request. Added core and coding-agent conversion coverage for Anthropic refusal metadata.

Fixes #3592
2026-06-26 17:56:07 +00:00
Alexander Kirilin ab513757f8 fix(agent): preserve tool turn before mid-run compaction 2026-06-26 09:35:24 -04:00
Alexander Kirilin e99b8c4128 docs(agent): clarify mid-turn continuation hook 2026-06-26 09:17:07 -04:00
Alexander Kirilin bbe3d98375 fix(agent): compact long tool loops mid-turn
Co-authored-by: coderred <coderredlab@gmail.com>
2026-06-26 09:05:11 -04:00
can1357 415099c420 fix: stripped stale snapcompact archive state during compaction strategy migration
- Fixed stale `preserveData.snapcompact` frames leaking into context-full compaction after switching from `snapcompact` to `context-full` strategy, which inflated context usage and made sessions appear to compact prematurely.
- Added secret redaction for migrated snapcompact archive plaintext (`text`/`textHead`/`textTail`) during the snapcompact->context-full transition, while preserving opaque provider-replay state byte-identical.
- Added `archiveSourceText()` and `stripPreservedArchive()` utilities to snapcompact module for archive extraction and cleanup.
2026-06-26 14:31:50 +02:00
can1357 93b80e5b1b refactor: consolidated duplicate
- Consolidated duplicate `stripSnapcompactPreserveData` functions into `snapcompact.stripPreservedArchive`.
- Added unit tests to verify archive removal and empty state collapse behavior.
2026-06-26 14:28:49 +02:00
OpenAI GPT-5.5 d7fdeb8e2b fix(agent): use static snapcompact archive prompt
Co-Authored-By: OpenAI GPT-5.5 <noreply@openai.com>
2026-06-26 08:21:11 +00:00
OpenAI GPT-5.5 5c4e269900 fix(agent): migrate snapcompact archive on context compaction
Co-Authored-By: OpenAI GPT-5.5 <noreply@openai.com>
2026-06-26 08:15:17 +00:00
jiwangyihao af32c22ffe fix(agent): /move 后按会话实时 cwd 重新作用域 Duo 发现
机器人指出 agent.ts 的 #cwd 在构造时固定,/move 更新 SessionManager 与
进程 cwd 后不会重建 Agent,导致 GitLab Duo Agent 的 namespace/project 发现
持续读取旧仓库的 git remote。

按既有 resolver 模式(getReasoning/getServiceTier)修复:

- Agent 新增可选 cwdResolver;构造时存入 #cwdResolver。
- AgentLoopConfig 新增 getCwd 每调用解析器,config 同时携带静态 cwd 与
  getCwd。
- agent-loop 在 streamFunction 调用点计算 effectiveCwd = getCwd?.() ?? cwd,
  每次 LLM 调用读取一次,因此运行中途的 /move 也能被工作区级 provider 发现
  感知。
- sdk.ts 主 Agent 传入 cwdResolver: () => sessionManager.getCwd(),该值在
  /move 时由 SessionManager.#cwd 更新。

新增针对可观测契约的回归测试(mock streamFn 记录 options.cwd):resolver
覆盖静态 cwd、resolver 返回 undefined 时回退静态 cwd、以及运行中途变更可被
逐次调用读取(模拟 /move)。
2026-06-26 16:13:24 +08:00
jiwangyihao d6194030e5 feat(agent): thread cwd through to local tool execution 2026-06-26 16:13:20 +08:00
roboomp 2b68785e82 fix(agent): include provider payloads in append-only digests
Responses-style providers serialize providerPayload history items instead of the
visible message blocks when replaying native history. Include providerPayload in
the append-only per-message digest so payload-only history rewrites stop the
stable-prefix walk and re-sync the changed message before any later divergent
tail.

Add a regression where an assistant message keeps identical visible content and
id but changes its openaiResponsesHistory providerPayload while a later message
also diverges; syncMessages must preserve the prefix before the assistant and
refresh the assistant payload.

Fixes #3406
2026-06-24 23:11:12 +00:00
roboomp 5b9e009a8b fix(agent): bound append-only stable prefix by log length
Direct callers can clear AppendOnlyContextManager.log without resetting the
private sync cursor. The advisor reset path does this when recycling its helper
agent, leaving lastSyncCount and messageDigests describing the old transcript
while the physical log is empty.

Clamp the stable-prefix reuse count to the current log length before truncating
and appending. A direct log clear now forces the next sync to replay from index 0
instead of starting from a stale private cursor and dropping prefix messages from
the provider context.

Add a regression that clears the public log after syncing two messages, then
resyncs a context with the same first message and a rewritten second; both
messages must be present in the rebuilt append-only log.

Fixes #3406
2026-06-24 23:07:11 +00:00
roboomp cce627ee4b fix(agent): include tool result metadata in append-only digests
Track internal tool-result metadata in append-only per-message digests so
metadata-only rewrites of toolCallId, toolName, or isError stop the stable-prefix
walk and re-sync the changed tool result before any later divergent tail.

This prevents stale tool-result pairing or error state from being preserved when
the text content stays unchanged but provider-serialized metadata changes.

Fixes #3406
2026-06-24 22:42:51 +00:00
roboomp 65a974aa90 fix(agent): preserve append-only log prefix on in-place message rewrites
`AppendOnlyContextManager.syncMessages` hashed a single rolling digest
over the entire synced prefix, so any in-place rewrite of an already-
synced message — per-turn `pruneSupersededToolResults` / `pruneToolOutputs`
collapsing a tool result, image stripping, or a `transformContext` re-render
— triggered `log.clear()` and re-appended the full conversation from
the current (mutated) view. The provider's cached bytes still matched
the prefix, but every position past the divergence had to be re-prefilled.
On llama.cpp / Ollama / LM Studio this re-prefilled tens of thousands
of tokens every few turns (`n_past \u2248 end-of-system-prompt` collapse,
~40k-token full re-prefill, GPU pinned >400W).

Replace the rolling digest with per-message digests in `#messageDigests`,
walk the new sync against them to find the longest byte-stable prefix,
truncate the log down to that prefix via a new `AppendOnlyLog.truncate(count)`,
and only re-append the diverged tail. Genuine compaction (`length <
lastSyncCount`) still clears the log.

- Tail-only rewrite: prefix stays byte-stable; only the trailing message
  re-syncs.
- Deep rewrite: prefix up to the divergence stays byte-stable; the
  provider re-prefills from the divergent message onward (architectural
  minimum).
- True compaction: unchanged, full replay.

Replace the now-misleading `detects in-place rewrite of already-synced
messages` / `detects in-place rewrite via digest mismatch` tests with
`preserves the byte-stable prefix when a deep message is rewritten (#3406)`,
`preserves the prefix when the tail is rewritten (#3406)`, `appended
new messages keep the prefix stable even when the prior tail also
diverged (#3406)`, and `rewriting the first message still re-syncs
from scratch` so each invariant is asserted directly.

Fixes #3406
2026-06-24 22:33:22 +00:00
can1357 1052af5fa6 refactor(agent): restructured yield logic to support deterministic testing
- Extracted yield logic into a configurable YieldGate class to avoid process-global state.
- Injected time and sleep dependencies to support deterministic testing.
- Handled potential negative time progression by forcing a re-anchor instead of gating indefinitely.
- Maintained existing behavior for the public yieldIfDue export via a shared instance.
2026-06-24 15:30:29 +02:00
oldschoola c210bdd391 fix(proxy): scrub partialJson in catch-block error path on stream disconnect
Address review feedback: when the SSE stream disconnects after a
toolcall_delta but before toolcall_end/done/error, the catch-block
at lines 183-192 pushes the partial message as the error result
without calling scrubPartialJson. This leaked the internal partialJson
field into the final error message.

Added scrubPartialJson(partial) call in the catch block, before
pushing the error event.

Added test verifying partialJson does not leak when server disconnects
mid-tool-call (toolcall_start + partial toolcall_delta, no terminal event).
2026-06-23 10:16:06 -07:00
oldschoola 2a2aa9fe59 fix(proxy): preserve partialJson on content during streaming, scrub at terminal events
Address review feedback: downstream renderers (event-controller.ts:535)
read content.partialJson during toolcall_delta to pace streaming
previews (bash env assignments, write/edit smooth streaming).

Revised approach:
- toolcall_start: initialize partialJson on content via typed
  ToolCall & { partialJson: string } intersection (not as any)
- toolcall_delta: accumulate in side-channel Map, write onto content
  via typed intersection cast
- toolcall_end: delete partialJson from content + side-channel map
- done/error: scrubPartialJson() cleans any remaining blocks that
  never got toolcall_end (the original leak bug, now fixed for all
  terminal paths)

Added test verifying partialJson IS present during streaming and
IS absent after completion.
2026-06-23 09:40:32 -07:00
oldschoola 9c795f886e fix(proxy): use side-channel Map for partialJson, eliminating as-any casts
streamProxy stored internal partialJson streaming state directly on typed
ToolCall objects via 4 'as any' casts. If toolcall_end was skipped (stream
error, early done), the field leaked into the final AssistantMessage content,
corrupting downstream serialization.

Replace with a side-channel Map<number, string> keyed by contentIndex:
- toolcall_start initializes the map entry
- toolcall_delta accumulates into it
- toolcall_end cleans it up

The typed ToolCall object never carries non-spec fields. All 4 'as any'
casts are eliminated.

Added 4 contract tests covering argument parsing, partialJson isolation on
normal completion, partialJson isolation when toolcall_end is missing, and
multiple concurrent tool calls with interleaved deltas.
2026-06-23 08:17:29 -07:00
can1357 5c21b28786 feat: optimized handoff generation and harden request safety
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
2026-06-22 20:05:40 +02:00
can1357 0b5e2d8276 fix(ai-providers): normalized Anthropic tool call IDs
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
2026-06-21 06:04:49 +02:00
can1357 29f7b57fde feat: migrated core rendering and context transformation to async
- Updated `render`, `renderMany`, and native snapcompact methods to return promises, ensuring scalable async execution.
- Refactored `transformProviderContext` and `buildSideRequestContext` to support asynchronous operations in agent loops.
- Integrated `Promise.all` for improved concurrency when processing frame rendering and rendering batch operations.
- Updated all internal call sites, SDK hooks, and test suites to accommodate the asynchronous API signatures.
2026-06-21 00:32:13 +02:00
DarkPhilosophy 8c9ae9ef67 fix(compaction): exclude encrypted reasoning from the compaction floor
CI caught that flooring by the raw local estimate falsely triggers compaction on
thinking-heavy turns: estimateTokens counts the opaque thinkingSignature /
redactedThinking payloads (providers bill them on replay, #2275), but their local
byte size diverges wildly from what the provider actually charges — so a turn
with a large encrypted-reasoning blob but small provider usage would trip the
floor (broke agent-session-handoff 'provider-anchored usage' test).

estimateTokens now takes { excludeEncryptedReasoning } and the compaction floor
(#estimateStoredContextTokens) uses it: the floor counts only reliably-countable,
on-wire-compressible content (text, tool results, tool calls), while the provider
usage arm of compactionContextTokens still accounts for encrypted reasoning. This
keeps the encrypted-reasoning case provider-anchored while still flooring upward
when a before_provider_request hook compresses tool results.
2026-06-20 23:21:36 +02:00
DarkPhilosophy 7e8450192c fix(compaction): floor context tokens by local estimate so payload compression can't suppress auto-compaction
A before_provider_request extension (a context-compression proxy like Headroom,
an obfuscator, or inline snapcompact) can shrink the outgoing request below the
real stored conversation. The provider then reports deflated prompt tokens, so
the auto-compaction threshold never fires and the stored history grows unbounded
until it overflows the context window and can no longer be compacted at all.

Add compactionContextTokens(provider, storedEstimate) = max(provider, estimate)
and apply it to both the pre-prompt and post-response compaction decisions,
flooring the provider-reported tokens by the agent's own estimate of the stored
conversation. Display and cost accounting still use exact provider usage; only
the compaction trigger takes the floor.
2026-06-20 23:21:36 +02:00
can1357 b5437c2cad feat(coding-agent): improved prompt caching for ephemeral side-channel requests
- Added `buildSideRequestContext` to the `Agent` class to generate prompt-cache-friendly provider contexts.
- Updated ephemeral side-channel turns to forward the full tool catalog to maintain prompt cache hit rates.
- Injected a `developer` role reminder into ephemeral turns to instruct the model to suppress tool calls.
- Implemented automatic post-processing to strip any tool calls from ephemeral turn responses.
- Exported message and dialect helper functions in `agent-loop.ts` to support context construction.
2026-06-20 23:19:10 +02:00
roboomp d1066ae4ad fix(compaction): used Azure request shape for remote compaction
Azure Responses remote compaction was enabled by shouldUseOpenAiRemoteCompaction
but requestOpenAiRemoteCompaction still built OpenAI-style requests:
Authorization: Bearer plus a URL without api-version. Azure Responses uses an
api-key header and api-version query parameter, so normal Azure configs failed
and fell back to local summarization.

Derive Azure compact URLs from the Azure Responses base URL/resource-name
configuration, append api-version, and send api-key headers while preserving
custom headers. Existing OpenAI and Codex request shapes are unchanged.

Added a regression test that opts into azure-openai-responses compaction and
asserts the compact URL, api-key auth, absence of Authorization, custom header
preservation, and configured compaction model payload.

Fixes #3104
2026-06-20 07:48:33 +00:00
roboomp b7aefe0689 fix(compaction): enabled custom remote compaction
- Added provider/model remoteCompaction metadata and models.yml propagation.\n- Routed configured OpenAI-compatible compaction endpoints for custom providers.\n- Added compactionModel as a summary-only model selector that leaves the active session model unchanged.\n\nFixes #3104
2026-06-20 07:27:09 +00:00
can1357 1afa6ba68a feat(catalog): supported fireworks fast serving path
- Added support for "Fast" serving-path variants for select Fireworks models.
- Updated compatibility logic to route `-fast` suffixes to the appropriate router wire format.
- Extended the model generation catalog to include these Fast variants with their respective pricing.
- Updated AI types to allow the `priority` service tier for Fireworks providers.
2026-06-20 09:21:01 +02:00