505 KiB
505 KiB
Changelog
[Unreleased]
Added
- Added model metadata fields (
context_length,max_output_tokens,input_modalities, etc.) to auth gateway model listing responses
Fixed
- Fixed tool-argument repair applying lossy transformations (such as stringifying objects or stripping unrecognized keys) when validating union schemas (
anyOf/oneOf), preventing corrupted tool call and subagent payloads - Fixed 400 errors when communicating with local OpenAI-compatible inference servers that reject
chat_template_kwargs.reasoning_effortby improving reasoning effort parameter fallback and compatibility handling
[17.3.8] - 2026-08-19
Changed
- Fixed Gemini thought summaries occasionally leaking a raw
```thinking/``````thinkingfence delimiter into the reasoning block, so it no longer shows up as fence spam in the thinking display or persisted transcripts (#8719). - Fixed the OpenCode Go login prompting for an "OpenCode Zen API key": the shared login flow now names the provider you selected, so connecting OpenCode Go asks for an OpenCode Go key (the
opencode.ai/authconsole is still shared, as documented upstream) (#8738). - Fixed Anthropic-compatible endpoints with strict prompt validation (e.g. Z.AI GLM
api.z.ai/api/anthropic, which rejects the whole request with400 code 1213 "The prompt parameter was not received normally") failing sessions once a tool returned empty output on a vision-capable model: empty successfultool_resultblocks now encode ascontent: ""instead ofcontent: [], which both the official API and strict compatible endpoints accept. - Fixed
retry.usageReservePct(Reserve Margin) ignoring Claude Fable/Mythos weekly tier usage until it hit 100%, so a Fable model kept serving turns past the configured reserve; reserve health now honors the mapped tier row while credential-wide hard blocks still require confirmed exhaustion (#8773). - Fixed
cursor-agentstreams stalling with "Provider stream stalled while waiting for the next event" when Cursor asked the client to approve a hosted WebFetch / web search (reproduced oncursor-grok-4.6-xhighafter "I'll fetch the page…"). Thoseinteraction_queryframes — including the newer WebFetch field 9 this proto did not name — were dropped, so the server waited forever and the idle watchdog aborted a live connection. Permission queries are now answered; hosted search/fetch is approved, unnamed permission fields get anapprovedreply on the same field number, and prompts this client cannot serve are rejected so the turn can continue.
Fixed
- Fixed thinking effort selections being ignored for local Qwen 3.8+ models on llama.cpp and vLLM: the Qwen chat-completions dialects only toggled
enable_thinking, so the chat template always reasoned at itsxhighdefault no matter which level was selected. The encoder now routes the requested effort onto the template'sreasoning_effortkwarg (chat_template_kwargsfor both Qwen dialects, plus the top-level field newer llama.cpp builds map natively). - Fixed OpenAI Completions, Amazon Bedrock, and Cursor providers ignoring
onPayloadreplacement payloads. The hook now transforms the actual request body sent upstream on these providers, matching the Anthropic/Gemini/OpenAI Responses replacement contract.devin-agentstill does not fire the hook (its payload is a protobuf object). - Fixed Codex requests failing outright when the signed-in ChatGPT account is not entitled to the requested model; the exact model denial is now classified as an account-policy error so credential rotation can reach an entitled sibling account
- Fixed Perplexity email-OTP login after its verification response renamed the encrypted session token from
tokentochallenge_token. - Cloud Code Assist Gemini 3.6/3.7 Flash requests at
minimalnow sendthinkingLevel: LOWon the aliased-lowSKU instead ofMINIMAL, which the API rejects with HTTP 400. - Answer Cursor
interaction_querypermission gates (hosted web search, Exa, unnamed field-9 WebFetch) so the Run RPC continues instead of sitting silent until the 300s idle watchdog. - Fixed provider tool calls arriving with flattened array argument paths (e.g. Gemini's
questions[0].id) being stripped and rejected by argument validation; well-formed flattened paths are now rebuilt into the nested arrays the tool schema expects (#8886). - Fixed opencode-go (Console Go) rejecting Responses turns with
400 No tool output found for tool call …(naming a random call of the batch on each retry) when a model streamed a trailing text/thinking block after its tool calls:buildResponsesInputemitted that block as an assistantmessageitem wedged between thefunction_callbatch and itsfunction_call_outputitems. Such interleaved messages are now hoisted ahead of their call batch (canonicalmessage(s) → calls → outputs), which the strict gateway validator accepts; content is unchanged (#8789). - Fixed the OpenAI-wire transport sleeping on a LiteLLM concurrency-admission 429 (
rate_limit_type: max_parallel_requests,Retry-After: 60) and retrying it up to 6 times (~300s) before session recovery saw the error. Because a 60s hint equals the transport'smaxDelayMscap,fetchWithRetrykept sleeping and retrying; the request now surfaces on the first attempt soTurnRecovery's concurrency backoff/model fallback runs promptly. Genuine RPM/quota 429s (no such marker) still honorRetry-After(#8854). - Fixed OAuth login (Codex
localhost:1455, and anylocalhostcallback flow) failing on hosts with IPv6 disabled at the kernel (ipv6.disable=1). The::1companion listener added in #8081 fails there with Bun's generic "Is port X in use?" message (oven-sh/bun#7187), which the in-use check misread as a real collision — tearing down the healthy IPv4 listener and surfacing a bogus "port 1455 is in use" error. The dual-bind path now detects the missing IPv6 loopback up front and serves IPv4 alone (#8814).
[17.3.7] - 2026-08-17
Changed
- Send the
omp/<version>User-Agent on xAI chat (xaiandxai-oauth) unless the request already set its own.
[17.3.5] - 2026-08-16
Added
- Added retryable oneshot completion support (
retryTransientCompletion) so non-agent LLM calls correctly retry on transient provider failures (Anthropic overload/rate-limit errors, HTTP 429/500/502/503/529), honoring provider-supplied retry-after timing before giving up.
Fixed
- Fixed xAI availability detection so paid-key-only setups correctly default to
xai/grok-4.5instead of the free SuperGrok catalog; explicitxai-oauth/…selectors still work as before. - Fixed xAI Responses requests sending unsupported parameters (reasoning summary, presence/frequency penalties) that some models rejected.
- Fixed Umans usage reporting incorrectly marking quota as exhausted based on raw request counts instead of actual weighted usage, and improved the usage display to show both a soft-cap warning and a hard exhaustion limit with an accurate countdown to reset.
- Fixed
omp usage invalidateto fully clear stale usage data and force a fresh refresh, so upgraded subscriptions no longer show outdated quota information. - Improved session recovery to correctly treat certain Cursor HTTP/2 connection errors as transient instead of ending the session.
- Fixed OpenAI-compatible streams (e.g. DeepSeek) that are cut off mid-generation being silently treated as a completed response instead of being retried.
- Fixed DeepSeek resource-exhaustion interruptions not being automatically retried.
- Fixed tool-call IDs being lost during same-model replay, which could break correlation with custom gateways.
- Fixed Kimi Code multi-account routing to prefer accounts with more available quota, respect usage-limit cooldowns, and keep consistent usage history across token refreshes.
- Fixed Anthropic custom signing-proxy conversations losing tool-search results and thinking content during replay.
- Fixed rare runaway response loops across model providers so they now fail gracefully instead of repeating indefinitely.
- Fixed xAI rejecting entire turns due to certain MCP tool schema shapes, restoring compatibility while isolating any remaining incompatible tools rather than failing the whole request.
- Fixed Alibaba DashScope/Bailian transient per-minute rate limits being misclassified as full quota exhaustion, causing unnecessary long backoffs instead of quick retries.
- Fixed Anthropic-compatible streams dropping thinking content, which broke replay of prior reasoning.
- Updated the Alibaba Coding Plan China login flow to point to the current Bailian API-key management console.
[17.3.4] - 2026-08-14
Fixed
- Fixed
omp usage invalidateto discard stale OAuth and API-key usage snapshots, then force a cache-bypassing, per-provider serialized refresh with a broker request budget sized for the full unfiltered account batch, so upgraded subscriptions do not silently retain pre-change quota data. - Fixed quota reporting and Cookie capture guidance for China (Beijing) Alibaba Token Plan credentials (#8509).
[17.3.3] - 2026-08-14
Fixed
- Distinguished Gemini thought-only
STOPresponses from empty transports, avoiding repeated identical reasoning requests and duplicate Antigravity endpoint streams while surfacing the missing final output for session-level recovery.
[17.3.2] - 2026-08-13
Fixed
- Dropped unsigned thinking blocks from Antigravity Claude requests instead of sending them without a signature, preventing HTTP 400 responses when resuming sessions or switching models.
- Classified Antigravity HTTP 429 responses from structured
google.rpc.ErrorInforeasons (QUOTA_EXHAUSTED,RATE_LIMIT_EXCEEDED, andINSUFFICIENT_G1_CREDITS_BALANCE), using retry delays of five minutes or longer to distinguish rotatable quota windows from transient throttling instead of relying only on message regexes.
Removed
- Removed the Antigravity identity-prompt injection (
ANTIGRAVITY_SYSTEM_INSTRUCTIONandshouldInjectAntigravitySystemInstruction): Cloud Code Assist accepts arbitrary system instructions on gemini-3.x and Claude routes (verified live), and the injected stub never matched the real client's system prompt anyway. User system prompts are now sent unmodified (still taggedrole: "user"). - Fixed Antigravity
automode not failing over to the sandbox endpoint when the daily endpoint returned a thinking-onlySTOP, which caused Advisor turns to be falsely recorded as empty-response failures (#8480).
[17.3.0] - 2026-08-13
Breaking Changes
- Renamed
withGeminiThinkingLoopGuardtowithThinkingLoopGuard; the guard applies to Gemini, DeepSeek, and Grok model-id families.
Changed
- Updated OpenCode Go integration to use the official usage endpoint, removing hardcoded caps, enabling real-time credential validation, and routing multi-key pools based on rolling and weekly headroom.
- Optimized Anthropic prompt caching with rolling 5-minute breakpoints and idle refreshes to keep the prompt prefix warm.
Fixed
- Fixed Ollama chat adapter to correctly forward sampling parameters like temperature and topP to the provider.
- Fixed OpenAI agent turns ending prematurely after a web search with no visible answer, ensuring the agent continues processing the search results.
- Fixed a resource leak where completed model streams retained provider concurrency permits longer than necessary.
- Fixed image input support for qwen3.8-max and newer models when using DashScope compatible-mode.
- Fixed xAI usage reporting falling back to a stale cache when a new weekly cycle starts with 0% consumed credits.
- Fixed Together AI login validation failures by querying the authenticated models list instead of a hardcoded model.
- Fixed credential-health probes and usage fetches failing when using reference-stored API keys (such as environment variables or commands) by ensuring secrets are correctly resolved.
- Fixed Perplexity email-OTP login by preserving the session cookies required for verification.
- Fixed thinking configuration for OpenAI and Daybreak models to correctly send reasoning.effort: "none" when thinking is disabled.
- Fixed Grok runaway thinking streams bypassing the thinking-loop guard.
Removed
- Removed legacy local request-cost estimation machinery and database schemas previously used for OpenCode Go estimates.
[17.2.15] - 2026-08-12
Fixed
- Fixed an issue where AWS_BEDROCK_SKIP_AUTH failed to expose Amazon Bedrock models when AWS credential files were unavailable.
- Fixed an issue where forceReasoningOff was ignored by Anthropic and Google transports, which allowed native thinking alongside a caller-supplied external scratchpad.
[17.2.14] - 2026-08-11
Added
- Added
forceReasoningOffanddisableReasoningoptions to disable reasoning in OpenAI and Azure OpenAI models
[17.2.13] - 2026-08-11
Changed
- Standardized first-party outbound User-Agent headers on
omp/<version>via the sharedUSER_AGENTutility.
Fixed
- Fixed the Amazon Bedrock and Cursor transports ignoring
StreamOptions.headers; both built their request headers from scratch, so caller-supplied tracing or attribution headers were silently dropped while working on every other provider (#8107 by @svperfecta). - Fixed Antigravity Flash turns hanging after successful response headers when the endpoint never emitted an SSE event; the provider now cancels the stalled body and fails over after 60 seconds while retaining the longer allowance for Pro reasoning starts.
- Fixed Cursor exec-bridge bash/grep calls failing ArkType validation when the server omitted optional frame fields: synthesized and executed tool args now drop
undefinedkeys (cwd,case,skip,timeout) instead of writingoptional: value || undefined. - Fixed Cursor sessions double-executing settled tools when
tools.formatis an owned dialect (e.g.gemini):wrapInbandToolStreamrebuilt toolCall blocks without copyingkCursorExecResolved, so agent-loop re-ran bash/grep/todo and appended a second result for the same call id. - Fixed Codex Responses Lite requests for opaque model codenames such as Daybreak omitting the required
reasoning.context: "all_turns"value and failing with HTTP 400. - Fixed Cursor personal usage reporting for current Pro / Pro+ / Ultra
/api/usage-summarypayloads that exposeindividualUsage.plan(and optionalonDemand) instead of the olderindividualUsage.overallbucket (#7998 by @dnth). - Allowed passive Google callers to accept empty or thinking-only
STOPresponses as successful silence instead of exhausting the provider's empty-response retry budget. (#8223) - Fixed the AWS credential resolver ignoring
role_arnprofiles: shared-config role chaining (source_profilerecursion,web_identity_token_file,credential_source) now resolves via STSAssumeRole/AssumeRoleWithWebIdentity, honoringrole_session_name/duration_seconds/external_id, so Bedrock is detected on EKS/IRSA and multi-account setups instead of reporting "No models available" (#8209). - Fixed Bedrock availability being under-detected on Nitro/EKS hosts: the EC2 metadata probe now recognizes Nitro DMI markers (
board_asset_taginstance ids,Amazon EC2vendor fields) in addition to the Xenec2UUID prefix (#8209). - Fixed DeepSeek Responses targets (opencode-go) rejecting a thinking-mode continuation with
400 The reasoning_text in the thinking mode must be passed back to the APIafter a prewalk hand-off plus mid-run compaction: the Responses input builder re-encoded replayed assistant turns without a reasoning item, so the request enabled reasoning but shipped noreasoning_text. The encoder now synthesizes areasoning_textreasoning item for every replayed assistant turn when the target requires reasoning replay in thinking mode (requiresReasoningContentForAllAssistantTurns/requiresReasoningContentForToolCalls), mirroring the chat-completionsreasoning_contentsafety net (#8248).
[17.2.12] - 2026-08-08
Fixed
- Fixed account-scoped Codex cyber-policy denials bypassing sibling credential rotation; replay-safe requests now try every configured account before surfacing the error.
[17.2.11] - 2026-08-07
Breaking Changes
- Fixed handling of GitHub Copilot's model_not_available_for_integrator error to prevent unnecessary retries, preserving the actionable available models list.
Added
- Added support for reporting Cursor personal monthly USD quotas and remaining balances, labeled by verified profile email accounts.
Fixed
- Fixed an issue where ANTHROPIC_BASE_URL was ignored for Anthropic chat requests, ensuring requests are routed to the configured host and forwarding ANTHROPIC_CUSTOM_HEADERS to non-official gateways.
- Fixed an issue where a legacy pre-organization login credential could persist and cause a permanent error row in omp usage even after a successful organization-scoped re-login.
- Fixed an issue where lazy provider streams (including Amazon Bedrock, Google, Cursor, Devin, and Ollama) ignored model-specific idle timeouts, which previously caused healthy but slow reasoning turns to prematurely time out.
- Improved error classification for Simplified Chinese quota-exhaustion and rate-limit messages, ensuring affected credentials are correctly rotated or backed off instead of being treated as unknown errors.
- Classified subscription and plan-cap 429 responses as rotatable usage limits rather than transient rate-limit throttles, enabling smoother credential rotation.
[17.2.10] - 2026-08-06
Breaking Changes
- Removed the
zoddependency andz/ZodTypere-exports. Tool schemas now useomptypetype()schemas, with Zod-style authoring still available via@oh-my-pi/omptype/zod.
[17.2.9] - 2026-08-05
Fixed
- Fixed GitHub Copilot requests failing with a raw
HTTP 400 model_not_available_for_integratoron roughly half of all turns for recently rolled-out models. Copilot's fleet is not uniform — part of it rejects models that/modelsadvertises on the same host — and the transient classifier matched only the oldermodel_not_supportedcode at a fixed envelope depth, so these rejections surfaced as terminal errors instead of entering the existing retry path. Model-availability 400s are now recognized at any envelope depth and rerolled on a flat delay with a dedicated 8-attempt budget on the OpenAI transports; every other retryable failure keeps its previous backoff and attempt count. - Fixed Cursor reads with inline OMP range selectors reporting the returned slice length as the source file's
totalLines, which made sequential reads of an unchanged file appear inconsistent (#7590). - Made model-scoped usage health ignore Codex accounts that cannot use the requested plan-gated model while retaining conservative unknown-state handling and independent usage-window resets.
- Fixed OpenAI Codex usage telemetry blocking explicitly allowed ChatGPT Team credentials when a weekly
used_percentrounded to 100, which could route multi-account sessions to an actually exhausted sibling instead (#7617). - Fixed OpenAI Codex GPT-5.x requests sending optional
reasoning.summary,reasoning.context, andtext.verbositycontrols by default, reducing Codexserver_errordisconnects from unsupported request shapes. (#4949) - Classified concurrent-request caps separately from quota exhaustion so they use a short retry backoff without burning a credential, and rotate credentials for account-scoped 403 caps such as Devin's overall message limit.
[17.2.7] - 2026-08-03
Changed
- Replaced
arktypewith@oh-my-pi/omptypefor schema validation, delivering up to 100x faster schema construction and 60-100x faster validation while maintaining full compatibility with existingtype/Typeexports and theisArkSchemacontract.
Fixed
- Fixed OpenAI-Codex (ChatGPT OAuth) requests failing with an
Unsupported service_tier: autoerror on default or legacy sessions by omitting the implicitautoservice tier on the wire. - Fixed an issue where Cursor
kimi-k3sessions would break permanently when a same-model assistant turn was persisted without thinking blocks, replacing hard errors with graceful warnings.
[17.2.6] - 2026-08-03
Added
- Added profile-aware Bedrock Mantle region selection, authenticated model discovery, bearer-token or SigV4 authentication, and credential refresh handling for OpenAI Responses models.
Fixed
- Fixed an issue where Ollama requests without a user-role message would fail to generate output or silently fail with a misleading error.
[17.2.5] - 2026-08-03
Changed
- Standardized tool-call examples in
renderToolExamplesandrenderToolInventoryto use Python keyword-argument syntax (name(key="value")) across all models, removing the model-specific dialect parameter and theDialectRenderOptions.exampleflag. - Updated
renderToolInventoryto render the tool catalog as a unified OpenAI-Harmony-style## functionsblock using TypeScript type declarations and comments, replacing the previous per-tool Markdown sections. - Added a
style: "harmony"option tojsonSchemaToTypeScriptfor generating compact, comma-delimited TypeScript definitions.
Fixed
- Fixed a session-blocking issue where unescaped Harmony control tokens in replayed assistant responses and tool inputs caused subsequent requests to be rejected with
invalid_prompterrors. - Fixed an issue where Codex Responses dropped native image-generation results from assistant content and replays due to stale
generatingstatuses. - Fixed Anthropic stream truncation handling where unexpected connection closures were incorrectly treated as clean stops, causing the agent loop to halt silently mid-sentence.
- Optimized Anthropic prompt caching to prevent unnecessary cache invalidation of the entire system prefix when volatile project footer details (such as current working directory, date, or workspace tree) change.
[17.2.4] - 2026-08-01
Fixed
- Fixed Codex WebSocket tool-result turns replaying full history when the preceding tool-call ID required Responses API normalization (#7279).
- Fixed direct Anthropic provider streams ignoring
model.compat.streamIdleTimeoutMs. Requests dispatched throughstreamAnthropiccan now widen the inter-event idle watchdog or set it to0to disable that watchdog; caller options and environment overrides retain precedence. Setting the compat value to0disables only the inter-event watchdog and leaves the first-event watchdog enabled; wider idle values continue to floor the first-event budget under the existing timeout contract. - Fixed OpenRouter DeepSeek models failing structured subagents when the upstream returns an opaque HTTP 400 for a strict yield schema, retrying once without strict tools and remembering the fallback for the provider session (#7264).
- Fixed provider-native Codex compaction streams bypassing WebSocket-first transport selection and SSE transport fallback (#7198).
- Fixed
SqliteAuthCredentialStore.open()running theauth_credential_refresh_leasesDDL (CREATE TABLE/CREATE INDEX) with Bun's defaultbusy_timeout=0, before the constructor's#initializeSchema()installed the busy handler. Under a concurrent write lock (e.g. WAL recovery on parallel omp startups) the lock-taking DDL failed immediately and, since the error wasn't BUSY-classified, bypassedopen()'s bounded retry loop. The busy handler is now installed on the connection immediately after it opens, before any lock-taking statement, honoring the issue-#2421 invariant on every entry path. (#7298) - Fixed a corrupt credential store (
agent.db) silently disabling every persisted rate-limit block.AuthStoragecaught unrecoverable SQLite errors (SQLITE_CORRUPTfamily /SQLITE_NOTADB) from the persisted block read/write paths atdebuglevel with no latch, so the broken store was re-queried on every credential evaluation while blocks quietly stopped applying. The first unrecoverable error is now reported once aterrorlevel with the store location and repair guidance, and every later persisted-block read/write short-circuits for the process lifetime; in-memory backoff still preserves availability (#7296).
[17.2.3] - 2026-08-01
Added
- Added the ai& (
aiand) provider registry entry with API-key paste login validated againsthttps://api.aiand.com/v1/models.
Fixed
- Fixed Anthropic OAuth (Claude Pro/Max subscription) requests hard-429ing (
Usage credits are required for long context requests) on every beta-gated 1M model — e.g.claude-sonnet-4-6, which the defaulttask/smol/scoutsubagent roles resolve to — regardless of prompt size, breaking all subagents. The 17.2.1 cowork request profile reintroduced thecontext-1m-2025-08-07beta for any model with a 1M catalog window, but subscription credentials have no long-context credit balance so Anthropic rejects the request outright. The beta is no longer advertised on OAuth requests; subscription accounts transparently get the standard 200k window. (#7238) - Fixed OpenAI Codex Responses ignoring disabled cache retention when deriving
prompt_cache_key, while preserving transport session identity (#7219).
[17.2.2] - 2026-07-31
Added
- Added support for the
gmi-cloudprovider registry, including API-key paste login validation and integration with@oh-my-pi/pi-catalog.
Changed
- Updated
AuthStorage.redeemResetCreditto prioritize spending the soonest-expiring available saved reset credit, and improved error handling to distinguish between transport failures (credit_list_failed) and a genuine lack of credits. - Exported
SENSITIVE_TOKEN_REfromproviders/transform-messagesto allow hosts to route credential shapes through reversible obfuscation instead of irreversible redaction.
Fixed
- Fixed an issue where Cursor conversation checkpoints were incorrectly recorded as billable output tokens, ensuring accurate usage totals.
- Fixed an issue in
AuthStorage.refreshStoredOAuthCredentialwhere expired OAuth credentials were returned without being refreshed when a credential mismatch occurred, which previously resulted in misleading "No API key found" errors. - Fixed Cursor history replay issues by preserving structured message order for assistant tool calls/results, retaining Kimi K3 thinking blocks, and preventing unsafe mid-session switches to K3.
[17.2.1] - 2026-07-30
Added
- Added exact OAuth credential-row resolution by durable credential id. The targeted path refreshes only that row and never ranks, rotates, or falls back to sibling accounts.
Changed
- Anthropic OAuth requests now reproduce Cowork's current
claude-desktoprequest profile, including client/runtime metadata, beta selection, system and billing attestation, the 64K output cap, and stable HTTP/1.1 header ordering.
[17.2.0] - 2026-07-30
Added
- Added first-class parentTurnId support for nested Codex requests, allowing stream options and metadata helpers to accept and safely propagate the initiating turn's ID.
- Added preservation of the Codex
encrypted_function_argsplaintext-collaboration marker on replayed function calls, keeping server-marked plaintext tool arguments from being reinterpreted as encrypted on subsequent turns. - Added interactive Exa API-key login through
/login exa, opening the official API-key dashboard and saving pasted keys to the credential store (#1798). - Cursor's modern exec wire protocol is now handled end to end.
agent.protomodels the frames current Cursor CLI builds emit — the seven Pi tools (ExecServerMessage45-51), hooks, subagents, allowlist prechecks, MCP state, smart-mode classification, canvas diagnostics, conversation search, agent-store conflicts and git diff — and every one of them gets a typed answer. The Pi frames run their local equivalents (read/bash/edit/write/grep/glob); the rest answer with the error, not-found or empty-but-valid variant that is actually true of this client. Frames this build cannot name at all now raiseExecClientControlMessage.throwwithunknown_exec_variant, and recognised frames with no truthful answer (git_diff_request, whoseGetDiffResponsehas no error variant) raiseexec_variant_unsupported, instead of a silent ack that leaves the server waiting. lspis advertised in the MCP tool catalog again. It was filtered out as a Cursor-native tool, but the nativediagnosticsframe covers one of roughly ten LSP actions, so the other nine were unreachable.- Added
pinSessionOAuthAccountsupport for backdating the sticky's last-use timestamp (options.lastUsedAtMs), so pins restored from persisted sessions keep the provider's warm-window semantics: resumes inside the prompt-cache TTL reuse the account, stale resumes still re-rank.
Changed
- Codex turn metadata now reserves the codex-rs
code_mode_tool_nameskey, preventing caller-supplied client metadata extras from colliding with the core-owned field. - Codex SSE requests to the official endpoint now use zstd-compressed bodies by default to match the official client, which can be disabled with PI_CODEX_ZSTD=0.
- API-key validation now preserves provider HTTP status and retry headers, allowing authentication, rate-limit, and server failures to retain their original error classifications.
- The Cursor Pi arg translation (
piReadPath,piJoinPath,piLsPath,piEscapeRegexLiteral,piLimit) moved toproviders/cursor-pi-args, re-exported fromproviders/cursor/exec-modernso existing imports are unaffected. The legacy pi shim shares these helpers and is compiled into the bundled virtual module registry, where a nestedproviders/<dir>/<mod>specifier is unresolvable under bunfs — and importing them from the exec module would drag the whole protobuf graph in for two string functions.
Fixed
- Fixed Novita login rejecting valid API keys belonging to Developer and Basic team members by validating against the chat completions endpoint instead of the billing balance endpoint.
- Fixed Cursor resource_exhausted errors being incorrectly classified as QUOTA_EXHAUSTED (which caused 30-minute credential blocks), mapping them to MODEL_CAPACITY_EXHAUSTED with a shorter backoff instead.
- Fixed a crash in Amazon Bedrock and Devin providers when Context.systemPrompt is passed as a bare string.
- Fixed aborted usage-limit recovery incorrectly blocking credentials or waiting on local usage fetches after the session had already changed.
- Fixed Codex WebSocket sessions echoing stale or missing turn states by capturing x-codex-turn-state refreshes from response metadata event headers.
- Fixed Harmony-dialect models (e.g., gpt-5.x, openai-codex) failing with invalid_prompt or "Request blocked" errors by escaping reserved control tokens in untrusted user and tool-result text.
- Fixed named forced tool_choice not being enforced on string-only OpenAI-compatible hosts (such as llama.cpp and LM Studio) by narrowing the advertised tools to the forced tool.
- Fixed direct Anthropic Claude Opus requests failing with HTTP 400 when the endpoint rejects strict tool fields.
- Fixed usage-based credential ranking for Anthropic accounts where a missing long-window (7-day) metric was incorrectly treated as a short-window metric.
- Fixed legacy Codex usage blocks continuing to gate all models after per-meter backoff was introduced, splitting the old shared scope into independent chat and spark blocks while maintaining backward compatibility with older clients and database schemas.
- Fixed Anthropic retry loops ignoring
maxRetryDelayMsfor long serverretry-afterhints, so over-budget delays surface immediately without losing response details or abort cleanup (#7003). - Added interactive xAI API-key login with key validation through the xAI models endpoint.
- Fixed Google Gemini and Vertex tool declarations carrying numeric, boolean, object-valued, or mixed
enumarrays that the Google Schema wire type cannot represent. Unsupported enums are omitted while valid string enums remain constrained. - Umans usage provider: fetches
GET /v1/usageand surfaces the rolling 5h request window + concurrency limits in/usage,omp usage, and the TUI status bar. - Fixed ranged legacy Cursor reads reporting the returned window byte length as the full file size.
- Updated the Cursor client build advertisement to activate the modern exec-frame protocol handled by this provider.
- Fixed a windowed Cursor
readreporting the window's line count as the file's.total_linesandfile_sizewere derived from the payload, which is the whole file only for an unranged read — a 20-line page of a 100-line file answeredtotal_lines: 20, which a paginating server reads as the end of the file. The count now comes from the read's own record of the file (details.meta.truncation.totalLines), falling back to counting the payload when the read returned the file whole. - Fixed a
pi_grepthat hit the native backend's internal match ceiling answering as an unqualified success.GrepToolfolds that cap into the flatdetails.truncatedalone, setting neitherdetails.truncationnorperFileLimitReached— the two fields the Pi result was built from — so the one truncation a caller can neither detect nor page around was the one it was never told about. The flat flag is now translated into aPiTruncation, and only when no specific cap already reported itself. - Fixed a
pi_grepframe'scontextandlimitvanishing from the transcript. The bridge honors both by building a scopedgrep, but neither is expressible in the model-facing schema, so the synthesized block recorded a plain pattern/path search — replaying a context-widened or capped search as an ordinary grep sitting beside output no ordinary grep produces. Both are now recorded on the block. - Fixed a Cursor MCP resource listing shrinking to a count in the transcript. The full URI/name/mime catalog goes out on the wire, but the paired local result recorded
Listed N MCP resource(s)— and rebuilt history is serialized from that result, so one reload later the model knew it had seen N resources and could name none of them. The paired result now lists what the answer carried. - Fixed the
pi_readrange translation padding the slice it asks for.piReadPathcomposed a plain:N+Kselector, which the localreadtool expands by one leading and three trailing context line — so a frame naming offset 5/limit 20 received lines 4-27. Ranged Pi reads now compose:raw:N+K; the wire result is an opaque output string, so the line-number gutterrawalso drops carries nothing the contract needs. - Fixed four Cursor exec frames answering with a result whose oneof was never set. In proto3 that is not an empty result — the server reads it as "the tool ran and produced nothing", indistinguishable from real success.
listMcpResourcesExecResult,readMcpResourceExecResult,recordScreenResultandcomputerUseResultnow sendListMcpResourcesSuccess{resources: []},ReadMcpResourceNotFound{uri},RecordScreenFailureandComputerUseErrorrespectively. - The MCP resource frames now answer from the host instead of a fixed verdict.
CursorExecHandlersgainedlistMcpResources/readMcpResource, so a host holding live MCP connections advertises them; the empty catalog andnot_foundabove remain the answer when no handler is supplied. A handler that throws surfaces asListMcpResourcesError/ReadMcpResourceErrorrather than collapsing into "none exist", which the model cannot retry. A read carryingdownload_pathforwards it and answers withReadMcpResourceSuccess.download_pathand no content, which is what that mode means. - Fixed Cursor
connect_scmcalls losing their repository and settling on a fabricated verdict. The target rides in theConnectScmArgs.targetoneof, so reading a flatgithubproperty always sawundefined; and the authoritativesuccess/error/rejectedresult only arrives on the completion frame, so answering at the announcement persisted a fixed failure for every call — including the ones the server went on to accept. The block now opens on the start frame and settles from the completion's decoded result. - Fixed interleaved Cursor tool calls corrupting each other. The stream decoder tracked a single "current" block and settled it on any
toolCallCompleted, ignoring the envelope'scall_id: a completion for one call closed whichever block happened to be open and paired it with the wrong result, andstart A, start Borphaned A entirely so its own completion settled B while A was never paired — which strips the whole interaction from every rebuilt transcript. Open blocks are now retained per envelopecall_id, and end-of-stream closes all of them rather than only the last. - Fixed a Cursor
search_conversationscall leaving no transcript block. The frame is answered from a fixed verdict, so nothing downstream pairs a result for it, and an unpaired call takes its whole interaction out of every rebuilt transcript. - Fixed a Cursor
read_mcp_resourcecall leaving no transcript block. The frame runs locally — and in download mode writes a workspace file — but synthesized no tool call and paired no result, so the read was invisible in the UI and absent from every rebuilt history; a resource download could mutate the workspace with nothing on record. The frame now synthesizes aread_mcp_resourceblock (notread: it is a remote MCP operation, and the name drives rendering and prune semantics) and pairs a result on success, not-found and error alike. Frames answered without a handler still synthesize nothing, since nothing ran. - Fixed a Cursor
list_mcp_resourcescall leaving no transcript block. The model consumed the catalog, but the frame synthesized no tool call and paired no result — its streamedListMcpResourcesToolCallannouncement was equally unrecognized — so the listing was invisible in the UI and absent from every rebuilt history. Frames a handler answered now synthesize alist_mcp_resourcesblock and pair a result derived from the same answer that went on the wire; frames answered from the fixed no-handler catalog still synthesize nothing, since nothing ran. - Fixed an unavailable
pi_edit/pi_writeanswering with the error variant. Both results model refusal and failure as separate oneof cases, and a denial reported aserrorreads as "the tool ran and broke" — inviting a retry of an operation that was never permitted. A frame whose tool is not granted, or whose handler produced nothing, now answers withPiEditExecRejected/PiWriteExecRejected; execution failures keep the error variant. - Fixed a Cursor MCP approval probe actually running the tool. A modern
mcpArgsframe carryingsmart_mode_approval_onlyasks only whether a call would be permitted, not for the call itself. The decoder dropped the flag, so the frame ran a side-effecting MCP tool the user had not been asked about, then ran it again when the real call followed. The flag is now carried through and the probe is answered from the host's policy without executing: approved only for a definite allow, refused for a deny, for a mode that demands a prompt the frame cannot raise, and for a tool the session does not have. No transcript block is synthesized either, since nothing ran. - Fixed the Cursor stream's end-of-transport cleanup erasing the arguments of every block still open. Blocks whose args arrive whole (todo, connect-SCM, MCP) never feed the streamed partial-JSON buffer, and reparsing an absent buffer yields
{}, so a truncated or disconnected turn rebuilt those calls with no arguments at all. Only blocks that actually streamed their args are reparsed now. - Fixed a Cursor stream dying mid-turn stranding the call it left open.
connect_scmand native todo blocks are stamped resolved the moment they open, so the agent loop synthesizes no placeholder and only their completion frame pairs a result — a transport that closed first left the card animating and the call unpaired, which takes the whole interaction out of every rebuilt transcript. The terminal-error path now closes open blocks and pairs those server-owned calls with an interrupted result; the flush ran only on clean completion before, which is not the path a dying stream takes. Exec-settled MCP blocks are left alone, since the dispatch that ran them owns their result. - Fixed the Pi exec frames displaying a different operation than the one they run. The provider synthesized its transcript block from a second, hand-rolled translation of the frame args, so
pi_read'soffset/limitwere shown as a whole-file read,pi_grep'sliteralpattern as an unescaped regex, andpi_find's path/glob join differed from the executed one. Both sides now share a single translation. - Fixed the streamed
pi_*_tool_callannouncements that modern builds send alongside each exec frame being unrecognized. The exec channel already synthesizes those blocks when it runs the tool; the duplicate was avoided only because the decoder recognized none of the variants, which would have started double-rendering as soon as any one was added. - Fixed
pi_bashresults reaching Cursor clipped with no truncation notice. Two truncation records exist locally:read/grepsetdetails.truncation, which carries an explicittruncatedflag, whilebashsetsdetails.meta.truncation, whose record has no such flag — its presence is the signal.piTruncationread only the first shape and required the flag, so every real Bash truncation was dropped and the server was told the clipped output was complete. Both shapes now translate, and an explicittruncated: falsestill suppresses the field.
[17.1.8] - 2026-07-28
Fixed
- Fixed an HTTP 400 error when resuming or replaying OpenAI history after an interrupted native Computer Use turn.
- Fixed connection 404 errors when using Google Vertex AI in multi-region locations (eu and us) by correctly resolving regional endpoint (REP) hosts.
- Fixed a resource leak in SqliteAuthCredentialStore.close() where unclosed prepared statements kept the SQLite connection alive, preventing database file cleanup (especially on Windows where files remained locked).
[17.1.7] - 2026-07-27
Changed
- Upstream
403 Forbiddenresponses (e.g. Anthropicpermission_errorplan/model denials, Copilot model-policy rejections) now rotate through sibling credentials like usage limits do, instead of failing the session on the first denied account. The denied credential is soft-blocked for 60s and re-validated — never removed — and the original 403 surfaces only once every sibling has been tried. - Usage report filtering in the auth-broker remote store is memoized per (reports, snapshot) with a precomputed per-provider OAuth credential map, replacing an O(reports × credentials) scan on every credential-selection and status refresh
- Cursor and Devin Connect-frame readers no longer copy every stream chunk through
Buffer.concatwhen the pending buffer is empty
[17.1.6] - 2026-07-27
Added
- Added
getProxyForUrl()for transports that need provider-specific and standard proxy environment resolution withNO_PROXYsupport (#6770). - Added SiliconFlow and SiliconFlow (China) to the built-in API-key login provider catalog so
omp login siliconflow/omp login siliconflow-cnstores a reusable credential validated against each region's/v1/modelsendpoint.
[17.1.5] - 2026-07-27
Fixed
- Fixed OpenAI Responses replay treating a tool output as paired with a matching call that appeared later in the input, or a tool call as paired with an earlier output. Pair repair now respects wire order before preserving or synthesizing each side.
- Fixed adaptive-thinking Anthropic models omitting the interleaved-thinking beta on signature-enforcing proxies, which caused persisted interleaved assistant turns to fail on replay (#6717).
- Kimi Code now sends its session-stable prompt cache key on both supported transports:
prompt_cache_keyfor OpenAI-compatible requests andmetadata.user_idfor Anthropic-compatible requests. Explicit keys survive side-channel session IDs, whilecacheRetention: "none"still disables automatic affinity (#6049). - Fresh encrypted auth-broker snapshot caches are revalidated within a short startup budget, so one-shot clients see newly imported or revoked credentials immediately when the broker is reachable while retaining cache fallback for transport and server failures.
- Fixed custom
anthropic-messagesendpoints dropping native web-search call/result blocks in the leaked-thinking wrapper, preserving signed continuation history in source order without carrying a preceding text signature onto later unsigned blocks (#6703).
[17.1.4] - 2026-07-26
Added
- MiniMax Token Plan accounts now report quota in
omp usage.GET /v1/token_plan/remainsreturns one bucket per plan quota, each carrying a rolling interval window and a weekly window, sominimax-codesurfaces real remaining percentages instead of an empty report. A model the plan does not include comes back looking like an untouched quota; those buckets are dropped from the report and named in its metadata. The mainland idminimax-code-cnis untouched. - OAuth logins now stamp
authorizedAt(epoch ms of the interactive login) on the stored credential, and every refresh-persist path preserves it. Anthropic expires the whole OAuth grant family ~30 days after authorization regardless of refresh-token rotation (observed asinvalid_grant: "Refresh token expired"on the latest rotated token, exactly 30 days after login, across four production accounts), so the login anchor is what makes re-login deadlines computable. ExportedANTHROPIC_OAUTH_GRANT_TTL_MSalongside the anthropic OAuth flow. - Added
GET /v1/credentials/disabledto the auth broker andAuthBrokerClient.listDisabledCredentials: disabled-credential tombstones (DisabledCredentialSummary— identity, verbatim disable cause, disable timestamp; never token material) so auto-disabled accounts stay visible to clients instead of silently vanishing from the snapshot.AuthStorage.listDisabledCredentialsserves the same data locally from SQLite; clients of brokers predating the endpoint get an empty list (404 mapped, no error). - Added
AuthStorage.revalidateCredentials()and the optionalAuthCredentialStore.refreshSnapshothook: remote broker stores re-fetchGET /v1/snapshoton demand so callers pairing live per-credential data with stored identities (omp usage) never render against the up-to-an-hour-stale disk-cached snapshot; local SQLite stores are always current and only reload. - Added an optional per-request
codexSseMaxAttemptsstream option to bound Codex SSE pre-response retries while preserving the six-attempt default when omitted. - Fixed Cursor requests failing with
Connect error internal: Unable to parse image: ...whenever the session history contained an image:rootPromptMessagesJsonimage parts now embed adata:<mime>;base64,URI instead of bare base64, matching the convention used by the OpenAI-completions provider (#6564).
Fixed
- Fixed OpenAI Responses native history replay sending output-only
statusfields back as input, preventinginput[N].statusfailures in long-running sessions. (#6513 by @Ant39140) - Cursor no longer discards a local tool result when the transport fails mid-execution. The provider waits for in-flight exec dispatches before pushing
done, but the error path skipped that wait, so a handler decoded from the last chunk landed its result after the Agent had already finalized the call from the terminal error and cleared its buffer — losing the real outcome of a tool that may already have run side effects. Both exits now drain the same barrier. - Cursor exec handlers returning the bare-result form no longer record a failed call as successful. When an SDK handler returns only a protocol result (no paired
toolResult), the synthesized transcript entry was always"Tool produced no transcript result"withisError: false, even for arejectedorerrorresult — so Cursor saw a failure while the rebuilt transcript showed success. The synthesized entry now derives its state and message from the result's own oneof variant — including MCP, where an application-level tool failure rides inside thesuccessvariant asis_errorrather than as a separate variant. - Fixed Cursor models silently failing to maintain the todo list. Cursor resolves its native
update_todos/read_todostools server-side, but the bridge looked for them under flattenedupdateTodosToolCall/readTodosToolCallproperties, which a decodedagent.v1.ToolCallnever has — the variant only arrives through thetooloneof — so no native todo call was ever recognized. The synthesizedtodotool call was also emitted as locally runnable with a{todos}payload the local tool's schema rejects, so any update that did surface ended as a validation error and local todo state never followed Cursor's. Todo calls are now read from the oneof, both native todo blocks are marked as already-resolved, and local state is mirrored from the server's confirmed success snapshot (leaving state untouched onUpdateTodosError).TODO_STATUS_CANCELLEDnow maps toabandonedinstead of reverting the task topending. - Hardened Cursor todo mirroring against partial
read_todosresponses: a read narrowed bystatus_filter/id_filter, or one returning fewer rows than the server's owntotal_count, is a subset rather than the list, and is no longer treated as authoritative. Previously such a response would have deleted every task it omitted. - Fixed an empty
update_todosresponse whosetotal_countis nonzero being mirrored as an authoritative clear, deleting every local task at once. The count-mismatch guard skipped empty responses entirely; only a matching zero count is a genuine clear now. An emptyread_todosstays refused outright, since proto3 decodes an unsettotal_countas0and it cannot be told apart from a filtered read that matched nothing. - Fixed a Cursor todo call being left unpaired when the completion frame carried no
tool_callat all.ToolCallCompletedUpdate.tool_callis optional, but the block was already marked as server-resolved by the started frame, so nothing emitted a placeholder for it and every transcript rebuild stripped the interaction. It now settles as "nothing to mirror", the same as a refused snapshot. - Fixed local Cursor exec calls (
read/write/grep/delete/bash/lsp/MCP) vanishing from rebuilt transcripts when the tool produced no result. The assistant block is synthesized and marked server-resolved before the handler runs, so the three result-less paths — no handler installed, a handler returning nothing, and a thrown handler — left the call unpaired. Each now pairs a result carrying the same text the server receives. - Fixed Cursor MCP tool calls being unrecognized on the wire.
ToolCall.toolis a protobuf oneof, so a decoded message exposes the variant as{ case, value }and never as a flatmcpToolCallproperty — the same trap that made native todo calls invisible while hand-shaped fixtures kept passing. Both the streamed start and the completion arg merge now go through a shared selector. - Fixed a streamed Cursor MCP block being named from
namewhile its paired result usedtoolName, so the two disagreed whenever the server sent different values. Both now prefertoolName. - Fixed the Cursor stream emitting
donewhile a tool handler decoded from the final chunk was still running. Server messages are dispatched fire-and-forget so the socket keeps draining, but nothing waited for them: when an exec request,turnEndedand the stream close arrived in one chunk, the turn finished before the handler produced its result, and the result missed the buffer drain that pairs it with its call. In-flight dispatches are now awaited after the transport completes. - Fixed a server-resolved Cursor todo call leaving its transcript block stuck pending: the synthetic completion was emitted under a freshly generated id instead of the streamed call id the interactive transcript filed the block under, so the card animated indefinitely. The settled call id is now handed to the sync handler.
- Fixed server-resolved Cursor todo blocks disappearing from rebuilt transcripts: nothing produced a
toolResultfor them, andbuildSessionContextstrips anytoolCallleft unpaired, so the interaction vanished on reload, branch switch, or transcript rebuild. The result the host builds is now persisted verbatim — it carries thedetails.phasesthe todo renderer rebuilds the list from, which a summary-only result would have replayed as0 tasks. - Fixed a refused or failed Cursor todo call leaving its card animating forever. Only a successful snapshot settled the block, so a
read_todosnarrowed by a filter and a serverUpdateTodosErrorboth went unanswered — notool_execution_end, and notoolResultto keep the block from being stripped on rebuild. Every completed native todo call now settles. A server error is carried through as a failed result rather than collapsed into the benign "nothing to mirror" case, which would have replayed the failure as a success. - Hardened Cursor todo mirroring against snapshots whose rows collide on content. Cursor's wire model identifies todos by
idand can represent two rows sharing the same text; the local list is keyed by content alone and thetodotool rejects a duplicate outright, so importing such a pair would leave every task-targeteddone/drop/rmresolving to the first row and the second unreachable. The snapshot is now refused like any other that cannot be represented locally — local state is left untouched and the call still settles as a no-op. - Hardened Cursor todo mirroring against ambiguous empty
read_todosresponses.total_countis a proto3 scalar, so an unset field decodes as0and is indistinguishable from a genuinely empty list; acceptingtodos=[]+total_count=0would clear every local task. Empty and mismatched reads are now refused —update_todosremains the authoritative clear path. - Fixed refused Cursor todo results claiming
"No todo changes". A server-acceptedupdate_todoscan still be declined locally (content collision, etc.), so the persisted fallback now reads"Todo snapshot not mirrored"instead of implying the remote call changed nothing. - Hardened Cursor todo mirroring against snapshots carrying unresolved
TodoItem.dependencies. The wire model blocks a row behind other rows byid; the local list has no ids and no edges, so an imported dependent row files as plainpendingandnextActionableTaskthen offers work the server considers blocked. Snapshots with an edge pointing at a row that is not yetcompleted/abandonedare now refused like any other that cannot be represented locally. Edges whose blockers already finished constrain nothing and still mirror. - Extended the Cursor todo
total_countmismatch guard toupdate_todos. A partial or size-limited merge response is as incomplete as a filtered read, but the check only applied to reads, so an update returning fewer rows than its own count was mirrored as the full list and deleted every task it omitted. An empty update still syncs — it remains the authoritative clear path, unlike an ambiguous empty read. - Hardened Cursor todo mirroring against rows with empty
content.contentis a proto3 string, so a missing or default value arrives as""; the local list is keyed by content and rejects a falsy one before lookup, leaving the imported row unreachable to every task-targeteddone/drop/rm. Such snapshots are now refused like any other that cannot be represented locally. - Fixed a deterministic circular-import TDZ that crashed
packages/catalog's test process withReferenceError: Cannot access 'claudeCodeVersion' before initialization:registry/oauth/anthropic.tsimportedclaudeCodeVersionfromproviders/anthropic.ts, which transitively pulls the registry back in (providers/anthropic→stream→registry→registry/oauth/anthropic), so the module-levelclaude-code/${claudeCodeVersion}bootstrap user-agent const read the binding whileproviders/anthropic.tswas still mid-initialization.claudeCodeVersionnow lives in a zero-import leaf module (providers/claude-code-fingerprint.ts) thatproviders/anthropic.ts,registry/oauth/anthropic.ts, andusage/claude.tsall import from, removing the cycle at the source rather than deferring the read. - Fixed a circular initialization between the Anthropic provider and OAuth registry that could throw before
claudeCodeVersionwas initialized when package tests or consumers loaded modules in parallel (#6628 by @anatoli-tsinovoy). - Stopped the account-level Codex
rate_limit.limit_reachedflag from being applied to individual chat windows. Codex reports one shared flag for the whole account, so a window with real headroom was markedexhaustedbecause a different window (or a separate metered feature) was at its limit, which over-blocked sibling accounts during credential selection. Each window's status now reflects only its own usage - Scoped Codex reactive backoff per meter: a
usage_limit_reachedfrom a Spark request no longer persists a block that ordinary chat requests honour, and the reverse. Blocks written before scoping used a shared scope meaning "block everything", so requests still honour it and reconciliation still heals it - Implemented
scopeLimitsfor the Codex ranking strategy so a request gates only on the windows it actually consumes:-sparkmodels spend the Spark meter and every other model spends the 5h/weekly chat windows, instead of OR-ing every window and meter into one provider-wide block - Fixed native Anthropic adaptive-only models (Opus 4.6+, Sonnet 4.6+, Fable/Mythos 5) keeping thinking ON when reasoning was meant to be off.
mapOptionsForApinever consulteddisableReasoningon the Anthropic branch, so a caller-side disable left adaptive thinking at full effort; anddisableThinkingIfToolChoiceForceddeletedoutput_config.effortalongsidethinking, which for adaptive-only models silently re-enabled adaptive thinking (a bare omission defaults to adaptive-ON). Both paths now omitthinkingand pin the lowest adaptive effort, sodisableReasoningand forcedtool_choiceturns (e.g. the delivery reviewer'sreport_delivery) actually suppress reasoning instead of returning a thinking block withend_turn(#6589). - Fixed Bedrock Converse dropping captured Claude thinking signatures when replaying application-inference-profile ARN models, restoring adaptive-thinking multi-turn conversations (#6610).
- Fixed the
alibaba-token-planlogin only supporting the international Singapore endpoint, which rejected China (Beijing) Token Plansk-sp-keys with401 invalid_api_key. Login now selects a region (International / China (Beijing) / Custom), validates the key against that region's/modelsendpoint, and stores the chosen base URL in the credential so inference and discovery both target it (#6682). - Fixed statusless provider capacity errors such as
no_capacityand high-demand responses being treated as terminal instead of retryable. (#6503) - Fixed QwenCloud Token Plan quota reporting to call the current console usage RPC and document how to capture its optional Cookie during login.
- Fixed Cursor exec-channel MCP calls such as
web_searchomittingtoolCallblocks when no interaction block arrives, which rendered their tool cards below the final assistant answer or dropped them on transcript replay. (#6501) - Fixed Claude scoped weekly limits (e.g.
Claude 7 Day (Fable)) withis_active: falsebeing dropped by the/usageparser, rendering asnot reportedinomp usagedespite carrying real utilization. Live payloads mark only the currently binding limit active — an account pinned at a 100% Fable cap reports its 77% shared weekly row as inactive too — sois_activesignals severity ranking, not bucket existence, and is now ignored. Exhaustion gating is unchanged: tier rows still hard-block only at confirmed 100% with a future reset. - Fixed a TDZ crash (
Cannot access 'claudeCodeVersion' before initialization) whenproviders/anthropicwas the first module loaded:providers/anthropic→stream→registry→registry/oauth/anthropiccircled back into the still-initializing provider module. The Claude Code fingerprint constants now live in the leaf moduleproviders/claude-code-fingerprint(star re-exported fromproviders/anthropic, so import paths are unchanged).
[17.1.3] - 2026-07-24
Fixed
- Fixed Cursor sessions exposing
ast_edit(and other staged-previewxd://devices) without a reachable resolver: the built-inwritetool — which carries thexd://resolve/xd://rejecttransport that finalizes a staged preview — was filtered out of Cursor's forwarded catalog, so previews could never be resolved and the session aborted after three forcedwriteturns.writeis now re-included in the forwarded catalog whenever pi-agent devices are advertised (#6536). - Fixed OpenAI Responses and chat-completions streams honoring per-model first-event watchdog policy, allowing local llama.cpp-style backends to process arbitrarily large prompts without a premature client cancellation (#6524).
[17.1.2] - 2026-07-24
Added
- Added
GET /v1/usage/historyto the auth broker (recorded usage-limit snapshots withsinceMs/providerfilters) andAuthBrokerClient.fetchUsageHistory— in broker deployments the broker host performs every upstream usage fetch, so its durable history is the only complete utilization record - Added per-client burn tracking to the auth broker: clients batch observed request usage per (provider, model) and flush it to
POST /v1/usage/observedevery 10 seconds (install id as client key, hostname as display name); the broker persists 5-minute buckets inclient_usage/clientsand serves aggregates fromGET /v1/usage/clients. Brokers without the endpoint disable reporting for the process lifetime
Changed
- Renamed the Z.AI feature quota row to
ZAI Zread Quota(tier/id slugzread), replacing the 74-charZAI Web Search / Reader / Zread Quota (web-search-reader-zread)title that wrappedomp usagerows
Fixed
- Fixed every Claude (
anthropic-messages) model on theopencode-zenprovider failing with401 Missing API key: the gateway requiresx-api-key, soopencode-zennow uses X-Api-Key auth likeopencode-go/umansinstead of bearer-only, and no longer sends thecontext_managementfield its Anthropic proxy rejects on thinking requests (#6510). - Fixed Anthropic native server-tool blocks being dropped from persisted assistant turns, preserving signed web-search continuations in their original response order (#6495)
[17.1.1] - 2026-07-24
Added
- Added
setCodexAttestationProviderAPI for injectingx-oai-attestationheaders in ChatGPT-OAuth Codex requests - Added OAuth account session pinning and active status tracking in storage
- Added OpenAI Responses native computer-use transport, including batched actions and exact
computer_call/computer_call_outputreplay with pending/acknowledged safety checks andimage_url/file_idoutput references. Models without native support receive the same action surface as a regular function tool; provider-specific tool-choice forcing is used where supported. - Added
PI_CODEX_RESPONSES_LITEto override the catalog-selected Codex Responses transport for diagnostics (1/trueforces Lite;0/falseforces the standard body). - Added caller-owned
cachedContentongoogle-generative-aiandgoogle-vertexGenerateContent options: pass an opaque cache resource name through the shared builder (blank values rejected); no create/refresh/delete lifecycle and no guessed model/project/location validation; existingcachedContentTokenCount→Usage.cacheReadnormalization is unchanged. - Added Anthropic extra-usage reporting across
omp usage, interactive/usage, and ACP/usage: the OAuth usage endpoint's authoritativespendpayload (or legacyextra_usagefallback when absent) is normalized into aClaude Extra UsageUSD row; capped accounts show limit/remaining/fractions and status, while uncapped spend exposes only its absolute used amount—rendered as$… usedin CLI/TUI and123.45 usd usedin ACP—without a fabricated cap, percentage, or status. (#5575) - Added process-scoped OAuth account pools for trusted auth-broker clients via
OMP_AUTH_BROKER_ACCOUNT_POOL_FILE, consistently filtering snapshots, streaming updates, refreshes, and usage reports to selected OAuth identities while leaving API-key credentials and the shared encrypted snapshot cache unrestricted. - Added opt-in Vercel AI Gateway automatic prompt caching for OpenAI Chat Completions while preserving
onlyandorderrouting preferences. - Added Vercel AI Gateway Responses cache anchors and cache lifetimes, emitted only with automatic caching.
- Added opt-in OpenAI GPT-5.6 explicit prompt-cache controls for Responses and Chat Completions. Existing requests remain implicit; the policy marks at most one existing stable-history block and is rejected locally on unsupported explicit routes.
- Forwarded
statefulResponsesthroughstreamSimple, so diagnostic callers can explicitly disable OpenAI Responsesprevious_response_idchaining. - Added native QwenCloud Token Plan API-key login, model discovery, and an optional interactive console-Cookie prompt for 5-hour and 7-day quota reporting (#6151).
- Added model-scoped usage health and same-provider reselection for native coding-plan credential pools, preserving OAuth/login-pool precedence, scoped broker blocks, sibling rotation state, and conservative unknown-account handling while excluding ordinary configured API keys (#5018).
Fixed
- Fixed stateful OpenAI Responses explicit cache breakpoints being restored onto edited historical messages, ensuring full replays recompute the latest stable cache boundary.
- Fixed ChatGPT Codex standard and Lite transports rejecting or hiding native computer-use payloads by unrolling the tool definition, forced choice,
computer_call, andcomputer_call_outputinto ordinary function-tool forms.
[17.1.0] - 2026-07-24
Added
- Added support for caller-owned
cachedContenton Google Generative AI and Google Vertex AIGenerateContentoptions, allowing passing of opaque cache resource names. - Added Anthropic extra-usage reporting across CLI, interactive, and ACP usage endpoints, normalizing the authoritative
spendpayload into a 'Claude Extra Usage' USD row with accurate limit, remaining, and status details. - Added process-scoped OAuth account pools for trusted auth-broker clients via
OMP_AUTH_BROKER_ACCOUNT_POOL_FILEto filter snapshots, updates, refreshes, and usage reports to selected OAuth identities. - Added opt-in Vercel AI Gateway automatic prompt caching for OpenAI Chat Completions, including support for cache anchors and cache lifetimes.
- Added opt-in explicit prompt-cache controls for OpenAI GPT-5.6+ Responses and Chat Completions, supporting stable boundary selection, stateful Responses markers, and future GPT-5.x/6.x models.
- Added support for forwarding
statefulResponsesthroughstreamSimpleto allow diagnostic callers to explicitly disable OpenAI Responsesprevious_response_idchaining. - Added native QwenCloud Token Plan support, including API-key login, model discovery, and an optional interactive console-Cookie prompt for quota reporting.
- Added interactive Meta Model API key login and support for
MODEL_API_KEYandMETA_API_KEYenvironment variables. - Added model-scoped usage health tracking and same-provider reselection for native coding-plan credential pools.
Fixed
- Fixed AWS Bedrock cache checkpoints to use resolved model compatibility, falling back to the provider-default 5-minute cache for unsupported 1-hour retentions, emitting AWS-recommended explicit checkpoints for Nova models (Lite, Micro, Pro, Premier, Nova 2 Lite), and honoring configured checkpoint maxima.
- Fixed Pi-native and compatibility-wrapper requests dropping cache controls required by
omp bench --cacheto preserve explicit prompt-cache affinity and allow disabling OpenAI Responses chaining. - Fixed outbound credential-pattern redaction running unconditionally; it is now opt-in via
configureCredentialRedactionand disabled by default. - Fixed SuperGrok (
xai-oauth)/usagereporting for unified-billing accounts by falling back to the default monthly limit and usage payload when credit usage percentages are absent. - Fixed sessions wedging with a
400 Invalid signature in thinking blockerror when switching Anthropic-compatible providers by stripping signatures whose issuing provider differs from the target. - Fixed OAuth callback servers aborting login on premature invalid callbacks, and restricted
localhostcallback listeners to the IPv4 loopback interface. - Fixed Google Gemini CLI and Antigravity OAuth login hanging indefinitely during Cloud Code Assist project provisioning by introducing request timeouts, cancellation checks, and bounded polling.
[17.0.9] - 2026-07-23
Added
- Added Synthetic (synthetic.new) usage provider:
/usagenow reports the rolling 5-hour request limit and weekly credit quota viaGET /v2/quotas, including per-tick regeneration rates in the window labels. - Added optional
UsageWindow.resetLabelso rolling windows can render their countdown with an accurate verb (e.g. "tick in 12m" / "regen in 51m" instead of "resets in") — both quota windows on Synthetic regenerate incrementally rather than hard-resetting.
Fixed
- Fixed GitHub Copilot OpenAI-compatible requests being rejected when the session's native OpenAI service tier was set to
priority(#5160 by @audreyt). - Fixed OpenAI Responses token-cap truncations suppressing fully streamed function and custom tool calls whose inputs are complete.
- Added SuperGrok (
xai-oauth) usage tracking for weekly credits, product limits, and positive on-demand caps.
[17.0.8] - 2026-07-22
Fixed
- Fixed Gemini Flash Cloud Code Assist empty-response retries when responses contain only intercepted planning-leak JSON.
- Fixed Antigravity auto-routing to correctly fail over to the sandbox endpoint when the daily endpoint exhausts its retries.
- Fixed OpenAI-compatible providers configured with auth: none incorrectly sending an Authorization: Bearer N/A header, which broke custom endpoints using alternative authentication headers.
- Fixed auth-gateway model listings exposing duplicate or ambiguous model IDs by ensuring only provider-qualified routing IDs are advertised.
- Improved connection error handling by classifying generic connection failures as transient, allowing them to be retried, while keeping explicit authentication rejections non-retryable.
- Fixed custom Anthropic base URLs losing native thinking signatures during continuation requests.
- Fixed Alibaba Coding Plan Custom login rejecting valid API keys on endpoints that do not serve the default validation model by validating against the model catalog instead.
[17.0.6] - 2026-07-20
Fixed
- Fixed OpenAI Codex credentials limited to one ChatGPT workspace per email: a personal Plus/Pro plan and a Team/Enterprise seat under the same email now coexist in the auth store — with separate rotation and usage pools — instead of the second login silently replacing the first. The workspace (
chatgpt_account_id) is captured as the credential's org at login with the plan type as its display label, and two members of one workspace keep separate rows (#2966). - Fixed Devin total-token usage omitting cache reads and cache writes.
- Fixed model switches to Devin rejecting foreign provider response IDs, reasoning signatures, and empty interrupted turns as invalid Cascade history.
- Classified zero-output Devin
invalid_argumenttrailers as context overflow when the serialized message history is already large, routing cumulative tool-output payload failures through context maintenance—including artifact-backed shake rescue—instead of retrying the same rejected history.
[17.0.5] - 2026-07-18
Changed
- Changed Anthropic API-key requests to default to a 1-hour prompt-cache retention (using the extended-cache-ttl-2025-04-11 beta) to prevent cold-misses during idle sessions, with support for PI_CACHE_RETENTION values "short" and "none" to override this behavior.
Fixed
- Fixed transient OpenAI stream truncations by retrying once before output becomes replay-unsafe, preventing recoverable transport errors from failing the turn.
- Fixed native Kimi Code K3 thinking being disabled during named function selection by utilizing generic required tool choice.
- Fixed /login moonshot validating China-platform API keys against the international host instead of honoring MOONSHOT_BASE_URL.
- Fixed Anthropic session stickiness suppressing usage-based re-ranking indefinitely by gating stickiness on a 1-hour cache warmth window (configurable via ANTHROPIC_SESSION_STICKY_CACHE_WARM_MS) to restore proactive multi-account load balancing after long idle periods.
- Fixed credential ranking where clockless Anthropic usage windows incorrectly outranked clocked sibling credentials.
- Fixed tool request failures (HTTP 400) on local grammar-constrained OpenAI-compatible backends (such as llama.cpp, LM Studio, and vLLM) by widening bare boolean subschemas into a value-accepting primitive union.
- Fixed custom OAuth Anthropic-compatible endpoints receiving generated Claude Code fingerprint headers even when explicit header overrides were provided.
- Fixed active sessions for plan-gated OpenAI Codex models (Sol/Luna) silently re-routing to sibling OAuth accounts when usage headroom changed, ensuring session stickiness is preserved as long as the preferred credential remains usable and eligible.
[17.0.4] - 2026-07-18
Fixed
- Fixed Kimi Code usage reports dropping the 5h window reset time (
omp usageshowed no "resets in …" for the 5h limit): the API returnsresetTimeon the limitdetail, not onwindow, so the parsed row-level reset is now carried onto the window when the window itself has none. - Made Kimi device-id persistence best-effort: a missing or unwritable
~/.omp/agentdirectory no longer throws during Kimi header construction, which silently nulled everykimi-codeusage probe on fresh installs. - Coerced boolean tool-schema subschemas to MFJS object forms for native Moonshot/Kimi endpoints, preventing the task tool's
outputSchemafield from causing HTTP 400 responses (#5952).
[17.0.3] - 2026-07-17
Fixed
- Replaced the opaque
h2 is not supportedfailure on the Cursor run transport with an actionable error naming the ALPN-stripping proxy as the cause and pointing at theproviders.cursor.baseUrlHTTP/2 bridge workaround. The run RPC is HTTP/2-only, so behind a TLS-intercepting proxy that strips ALPN (e.g. Zscaler) bun cannot negotiateh2and the completion cannot proceed (#5828). - Restored the
createAssistantMessageEventStream()root export used by legacy provider extensions (#5879). - Fixed parallel Responses tool-result images interleaving synthetic user messages before all pending outputs, preventing strict OpenRouter/Moonshot backends from rejecting follow-up requests. (#5850)
- Fixed Kimi Code K3 requests to send native named efforts (
low,high,max) and use adaptive effort rather than generic token budgets on explicit Anthropic transport overrides (#5893). - Automatically invalidate and rotate OAuth credentials when an "invalidated oauth token" error occurs
- Fixed Anthropic usage reports treating the organization response header as the account identity, which caused the 5h/7d status-line segment to disappear for OAuth credentials without stored organization metadata. (#5698)
[17.0.2] - 2026-07-17
Fixed
- Automatically invalidate and rotate OAuth credentials when an "invalidated oauth token" error occurs.
- Fixed auth-broker snapshot validation rejecting API keys stored via the
/loginflow, restoring support for gateway/broker setups serving login-sourced keys on custom hosts. - Fixed an issue where literal reasoning tags (e.g.,
<think>) inside Markdown code blocks or inline code were incorrectly treated as reasoning boundaries, which corrupted the rendered Markdown. - Classified HTTP 402 and "balance exhausted" quota responses as persistent usage limits, enabling automatic rotation of multi-account requests to a sibling credential.
- Fixed
kimi-codeAnthropic-format requests ignoring custom provider base URLs. - Fixed an issue where GPT-5.6 Codex Responses-Lite requests failed with an HTTP 400 error due to invalid
tool_choiceparameters after tools were rewritten, by automatically downgrading forced hosted choices totool_choice: "auto"while preserving explicit tool-use constraints. - Fixed Cursor streams prematurely reporting success before late CONNECT or gRPC terminal failures were observed, and resolved issues rejecting transport ends without a
turnEndedsignal.
[17.0.1] - 2026-07-16
Fixed
- Fixed OpenRouter cost reporting to use the provider's authoritative account charge instead of catalog token-price estimates on both Responses and Chat Completions streams.
- Fixed OpenAI Responses and Chat Completions requests forwarding unsupported sampling parameters such as
temperatureto o-series and GPT-5+ models, preventing 400 errors for mnemopi memory calls through GitHub Copilot GPT-5.6 Luna. (#5606) - Fixed boolean JSON Schema subschemas (
true/false) in MCP tool inputs triggering400 INVALID_ARGUMENTon the Google/Cloud Code Assist (Antigravity) transport by coercing them to their object equivalents (true→{},false→{ not: {} }) before sending (#5604). - Fixed thinking-enabled Claude requests routed to
google-vertexsending theeffort-2025-11-24beta as ananthropic-betaHTTP header, which Vertex rawPredict rejects with a 400. The effort beta and theoutput_config.effortfield are now gated off the Vertex path the same waycontext-management-2025-06-27already is (#5614). - Fixed custom and Foundry-routed Anthropic endpoints receiving first-party eager/legacy tool-streaming controls (#5572).
- Parsed Ollama NDJSON response bytes directly instead of decoding and buffering every network chunk as text. (#5542)
- Fixed Amazon Bedrock stream error handling for non-
Errorvalues thatJSON.stringifycannot serialize (#5539). - Fixed concurrent provider OAuth refreshes by serializing rotating-token updates across processes, fencing stale writes, and preventing background usage probes from disabling otherwise usable credentials (#5396).
- Fixed OpenAI Codex WebSocket connections ignoring
PI_PROXY, provider-specific proxy settings, and standard HTTPS/ALL proxy variables (#5384). - Fixed Anthropic account quota exhaustion (
This request would exceed your account's monthly spend limit) hanging until the local deadline instead of surfacing the error: therate_limit_error"spend limit" wording is now classified as a persistent usage limit, so it fails fast and rotates to a sibling credential rather than looping in the provider retry backoff. (#4787) - Fixed OpenRouter daily free-model allowance errors (
free-models-per-day) being treated as transient rate limits, so requests rotate from an exhausted API key to a healthy sibling credential. (#4832)
[17.0.0] - 2026-07-15
Changed
- Improved Ollama streaming performance by parsing NDJSON response bytes directly instead of decoding and buffering network chunks as text.
Fixed
- Fixed Cursor TLS connection resets causing process-fatal uncaught exceptions, allowing the active turn to fail or retry gracefully without terminating the session.
- Fixed Amazon Bedrock stream error handling to correctly handle non-Error values that cannot be serialized by JSON.stringify.
[16.5.2] - 2026-07-14
Added
- Added OpenAI Codex rate-limit response-header ingestion to proactively refresh account usage snapshots and rotate credentials before hitting 429 errors.
Changed
- Optimized multi-account credential ranking to maximize quota utilization and prevent mid-session blocks by prioritizing expiring quota and demoting heavily used accounts.
- Improved responsiveness of credential blocking by bypassing the usage-ingestion throttle immediately when an account is detected as exhausted.
Fixed
- Fixed empty provider responses (such as from Cloud Code Assist API) being treated as non-retryable, allowing session retries and model-fallback chains to engage.
- Fixed OpenAI Codex watchdog timeouts bypassing transport and session retries by ensuring each request attempt has an independent timeout signal.
[16.5.1] - 2026-07-14
Added
- Added Cursor OAuth and access-token usage reporting to
omp usagevia Cursor's account usage endpoint.
Fixed
- Fixed OpenAI Responses
content_filterterminal events being auto-retried as provider finish errors, ensuring content-filtered turns remain hard failures without triggering a retry loop. - Improved credential rotation on usage and account-quota failures to cycle through all eligible credentials instead of stopping early, while maintaining rate-limit backoffs and safety guards.
- Fixed GLM tool call parsing to correctly handle and recover from missing or mistyped argument closers, preventing subsequent arguments from being swallowed.
- Fixed Anthropic credential management and usage routing for users with multiple organizations under a single email. Credentials, OAuth refreshes, usage reports, and active sessions are now correctly partitioned and isolated by organization, preventing subscriptions from overwriting or merging with each other.
- Fixed OpenAI and Codex response finalization to preserve streamed text when receiving empty content on completion. ([#5146])
- Fixed OpenAI Chat Completions request parsing to correctly accept assistant tool-call replay messages with null content. ([#5121])
- Fixed session-sticky OAuth credential mappings remaining active after credential changes, ensuring sessions correctly reselect accounts after login or logout. ([#4982])
- Fixed concurrent reasoning summaries to ignore legacy streaming events under cutoff contracts.
- Fixed Codex saved-reset redemption to apply to the selected OpenAI account in multi-account configurations. ([#5054])
- Updated the OAuth completion page to instruct users to close the tab manually when the browser blocks automatic window closing. ([#4855])
- Fixed Cursor
max_moderequests to correctly send max-mode metadata on both model payload fields. ([#4797]) - Fixed configuration discovery to support both nested and flat YAML formats for
auth.broker.urlandauth.broker.tokenkeys. ([#4734])
[16.5.0] - 2026-07-13
Added
- Added diagnostic response headers to auth-gateway inference endpoints, including request IDs (x-request-id/request-id), LiteLLM model metadata (x-litellm-model-id/x-litellm-model-api-base), and performance/cost metrics (x-litellm-response-cost, x-litellm-response-duration-ms, openai-processing-ms) on non-streaming responses.
Changed
- Updated Google and Google Vertex providers to always use streamGenerateContent requests.
Fixed
- Fixed empty provider responses (such as from Cloud Code Assist API) being classified as non-retryable, allowing session retries and model-fallback chains to engage instead of failing the turn.
Removed
- Removed automatic /interactions chaining for follow-up turns in Google provider calls, along with the useInteractionsApi, storeInteraction, and previousInteractionId stream options.
[16.4.6] - 2026-07-12
Added
- Added asynchronous
invalidateUsageCachemethod to clear cached usage reports - Added support for cross-service usage cache invalidation between AuthStorage and AuthBroker
Fixed
- Fixed OAuth credential resolution returning "No API key found" when every plan-eligible OpenAI Codex account was rate-limit blocked and the only unblocked account failed the model's plan gate: resolution now runs a last-resort ladder that first yields a plan-fitting account regardless of usage blocks (so callers get real usage-limit retry semantics), then tries every account with the plan filter dropped before reporting no credential
[16.4.5] - 2026-07-11
Fixed
- Fixed an issue in GLM tool calling where missing or malformed argument closers (such as
<arg_value>mistyped as</arg_key>) caused subsequent arguments to be swallowed or merged into a single field, affecting both in-band and native tool calling.
[16.4.3] - 2026-07-11
Fixed
- Fixed auth database upgrades from schema v5 by creating the OAuth credential refresh-lease table before lease statements are prepared.
- Fixed an issue in the Responses API where empty tool results were incorrectly serialized with a "(see attached image)" placeholder, causing models to look for non-existent attachments.
- Fixed OpenAI Responses server non-streaming envelopes to always include the required "incomplete_details" field, using null for completed responses.
- Preserved Cloud Code Assist tool schemas when mixed-type unions carry branch-local validation descriptions.
[16.4.2] - 2026-07-10
Fixed
- Fixed compatibility with xAI by automatically downgrading OpenAI-specific tool calls and image detail settings during message history replays.
- Fixed a race condition in shared SQLite OAuth token refreshes by implementing durable credential ownership and compare-and-set persistence to prevent stale refresh failures.
- Fixed OpenAI Codex requests to include the required version header for newly gated models.
[16.4.1] - 2026-07-10
Changed
- Enforced
all_turnsreasoning context for all Responses Lite requests
[16.4.0] - 2026-07-10
Added
- Added "max" as a first-class reasoning effort option across providers (including Anthropic, Google, Bedrock, and OpenAI), supporting a maximum reasoning budget of 32,768 tokens.
- Added and standardized the "Responses Lite" wire contract and transport, enabling automatic activation via model-level catalog flags, moving tools and instructions into developer input items, disabling parallel tool calls, and stripping image detail instead of falling back to the full transport.
- Added support for concurrent reasoning summaries on Codex Responses using the sequential-cutoff streaming contract.
- Added Novita API-key login with authenticated key validation and automatic NOVITA_API_KEY environment variable discovery.
Changed
- Recognized Pro Lite as a paid plan tier for OpenAI Codex models.
Fixed
- Fixed xAI SuperGrok multi-account rotation to correctly treat HTTP 403 credit exhaustion and spending limit errors as usage limits, triggering a credential rotation to a sibling account.
- Fixed error classification for AWS credential-resolution failures (AwsCredentialsError) to correctly map them as authentication failures.
- Fixed OpenAI-compatible chat-completions streams to preserve vLLM-style trailing cached-token usage chunks, ensuring accurate cacheRead and billable input session statistics.
- Fixed xai-oauth/grok-4.5 Responses requests to omit the unsupported reasoning.summary field while preserving the reasoning.effort payload.
- Fixed Codex OAuth credential selection to re-check blocked accounts during ranking and clear stale usage-limit blocks once live usage indicates recovery.
- Fixed sequential-cutoff reasoning summaries duplicating section headers across Codex reasoning items by tracking the cumulative summary response-globally, so replayed sections and replay-only items no longer re-emit text earlier thinking blocks already streamed.
[16.3.15] - 2026-07-09
Breaking Changes
- Renamed
OpenAIResponsesCacheOptions,normalizeOpenAIResponsesPromptCacheKey, andgetOpenAIResponsesPromptCacheKeyto the endpoint-neutralOpenAICacheOptions,normalizeOpenAIPromptCacheKey, andgetOpenAIPromptCacheKey.
Added
- Added automatic prompt-cache affinity header injection for OpenAI-family chat completions
- Added support for explicit prompt-cache affinity headers in OpenAI-family chat completions
- Added OpenAI pro reasoning mode support: models carrying the catalog
reasoningMode: "pro"marker (GPT-5.6 Pro aliases) sendreasoning: { mode: "pro" }on OpenAI Responses and Codex Responses requests, alongside the configured effort. The Codex request body now honorsrequestModelIdso catalog aliases request the base upstream model id.
Changed
- Updated xAI OAuth to use a dedicated device-code flow instead of redirect/loopback server
Fixed
- Improved account routing for GPT-5.6 models to better respect paid tier requirements
- Refined account selection logic to correctly identify plan types from account metadata
- Fixed OpenAI Codex multi-account routing for GPT-5.6: Sol and Luna requests now prefer Plus-or-higher accounts while Terra remains available to Free/Go accounts; local pro-mode aliases inherit their base model's Codex plan eligibility.
- Fixed xAI Grok OAuth login to use xAI's device authorization flow:
/loginnow opens the verification URL, displays the device code, and polls for approval instead of asking for a pasted redirect or linking to Hermes Agent documentation.
[16.3.14] - 2026-07-09
Changed
- Updated Codex reasoning effort mapping to support shifted wire tiers for newer models
Fixed
- Fixed the Codex Responses request transformer bypassing catalog/compat reasoning effort maps: the clamped user effort is now remapped to the provider wire tier (GPT-5.6's shifted five-tier scale sends
maxfor userxhighandxhighforhigh), failing loudly if a map produces a value outside the Codex wire vocabulary.
[16.3.13] - 2026-07-09
Changed
- Changed the xAI Grok OAuth (
xai-oauth) provider to use manual code-paste login by default./loginnow accepts a pasted authorization code or fullhttp://127.0.0.1:56121/callback?code=...redirect URL without starting a local callback listener (#3277 by @Jaaneek). - Renamed the xAI Grok OAuth provider in login and credential prompts to "xAI Grok OAuth (SuperGrok or X Premium+)" (#3277 by @Jaaneek).
Fixed
- Fixed the generic lazy-stream idle watchdog aborting healthy
cursor-agentstreams with "Provider stream stalled while waiting for the next event" while a Cursor exec-channel local tool (shell/read/grep/write/MCP/…) legitimately ran longer than the idle budget. Provider streams now advertise consumer-side local work in flight and the watchdog slides its deadline instead of aborting; genuinely silent streams still time out. (#4593) - Fixed OpenAI Codex/Responses reasoning streams so streamed thinking content is preserved when the final
output_item.donereconstructs to an empty summary (#4918). - Fixed Anthropic streams hanging forever when generation wedges mid-stream (notably long
writetool calls on Opus 4.8 high/xhigh) while the server keeps sendingpingkeepalives: pings now extend the idle watchdog only within a bounded window (3x the idle timeout) since the last real stream event, so a stalled tool-call stream times out and recovers instead of hanging with no retry path (#4900).
[16.3.12] - 2026-07-08
Added
- Added
AssistantMessage.toolCallAbortMessagesfor per-tool placeholder labels on aborted assistant turns (#2783).
Fixed
- Fixed Anthropic replay 400s (
tool_use ids were found without tool_result blocks immediately after) when a persisted assistant turn carries content after a completed tool call — such as a mid-turnserver-side-fallbackhandoff (fallback block plus continued text/tool calls after the primary model'stool_use) or trailing text from cross-provider replays — by stable-partitioning assistant content so alltool_useblocks trail the non-tool_usechain. (#4781, #544) - Fixed access-token-only OAuth credentials attempting token refresh with an empty refresh token after expiry.
- Fixed gateway usage-limit retries falling through to cross-provider model fallback before trying a sibling credential from the same provider.
- Fixed Codex usage-limit rotation treating Plus and K-12 accounts as separate quota groups for shared 5-hour/7-day windows.
- Fixed OpenAI Responses streams that end with
response.donebeing misclassified as premature stream closures. - Fixed OpenCode Go
/logincredentials being shadowed by an existingOPENCODE_API_KEYenv fallback after switching accounts. (#4688) - Fixed OpenAI Codex WebSocket continuations to treat proxy stale-anchor codes such as
codex_previous_response_staleas an expiredprevious_response_idchain — same recovery class as the OpenAI-standardprevious_response_not_found— so the turn is retried with full context instead of surfacing the error to the user (#4624). - Fixed Azure Foundry Anthropic utility requests to omit the structured-output beta whenever strict tools are disabled, preventing
structured_outputs not supported in your workspacefailures for Sonnet 5 compaction (#4679). - Fixed OAuth
launchUrladvertisement for flows whose redirect never returns to the local callback server: custom-scheme redirects (e.g. GitLab Duo'svscode://URI, whichnew URLparses without complaint) and fixed non-loopback hosts no longer receive ahttp://localhost:<port>/launchcopy target that misrepresents the callback endpoint and resolves nowhere for remote users. - Codex load balancing: clear stale persisted and in-memory usage-limit blocks for an
openai-codexaccount when a fresh live usage report shows it is allowed and below all limits, including broker-backed gateway snapshots, so traffic returns to recovered accounts instead of funneling to one sibling.
[16.3.11] - 2026-07-06
Fixed
- Fixed
openai-codex-responsesfresh plan execution requests that contained only system/developer guidance by mirroring the final instruction as user input so Codex accepts the first turn. (#4714) - Fixed Codex WebSocket compact/resume delta diagnostics to record request shape and raw-vs-displayed usage buckets, so persistent server-reported uncached suffixes without
orchestration_*fields are visible in debug stats. (#4707)
[16.3.10] - 2026-07-06
Fixed
- Fixed Ollama/Ollama Cloud EOS-only completions to retry empty stops with a single output token before the agent loop can halt silently. (#4659)
- Fixed Claude Sonnet 5 failing every request on feature-gated gateways (Azure Foundry, OpenAI-compatible relays) that reject strict tools with "structured_outputs not supported" — the rejection is now classified as a strict-tool rejection, so the request retries without strict tools and the session remembers the downgrade.
[16.3.7] - 2026-07-05
Fixed
- Fixed formatting of demoted reasoning blocks to prevent accidental concatenation with prose
- Fixed terminal whitespace issues in assistant messages that caused rejections by Anthropic API
- Fixed Cursor provider handling of empty-pattern grep arguments to return a clear, actionable error instead of a generic error and a broken TUI rendering.
- Fixed Google Cloud Code Assist API (Antigravity) and Gemini CLI to immediately bubble up underlying API errors (such as safety or recitation blocks) instead of incorrectly retrying and hiding them behind a generic empty-response message.
- Fixed GitHub Copilot OpenAI Responses replay to prevent empty reasoning-only assistant turns from being persisted in history and poisoning subsequent requests.
- Fixed classification of provider gateway quota-insufficient errors so they are correctly identified as usage-limit errors rather than generic 403 failures.
- Fixed OpenAI-compatible Responses models (such as DeepSeek endpoints) to preserve user-configured tool strictness settings unless strict mode is explicitly unsupported.
- Fixed custom openai-codex-responses providers failing when no ChatGPT account ID claim is present by omitting the header when it cannot be derived.
- Fixed token accounting for OpenAI Responses and Codex providers to correctly include provider-side orchestration tokens in billing totals without misclassifying them as uncached prompt input.
- Fixed Google Gemini and Cloud Code Assist providers to preserve the requested reasoning tier when sending requests with hidden thinking summaries.
- Fixed parallel OpenAI-compatible tool-call streaming to prevent argument data from bleeding across concurrent commands when identifiers are missing.
- Fixed Anthropic Claude reasoning and thinking replay handling. Same-model replays now drop unsigned prior reasoning blocks to prevent reasoning-extraction refusals, while cross-model replays (including Bedrock cross-region profiles) correctly demote reasoning without emitting raw thinking tags or causing text-flattening formatting issues.
- Fixed custom OpenAI-compatible relays serving standard OpenAI model IDs to be correctly classified as OpenAI-family targets for fast mode.
[16.3.6] - 2026-07-04
Added
- Persisted credential rate-limit blocks across processes:
auth_credential_blocks(auth schema v5) stores per-credential blocks keyed by row id + provider key + block scope with MAX-upsert semantics,AuthStoragemerges persisted and in-memory blocks on read, and auth-broker snapshots/SSE carry per-entry blocks withPOST /v1/credential/:id/blockandDELETE /v1/credential/:id/blocksendpoints so gateway and sibling omp processes stop re-discovering exhausted accounts by burning a 429 each.
Fixed
- Fixed Anthropic credential selection sampling Fable/Mythos-exhausted accounts on every new session: a Fable/Mythos weekly cap now proactively hard-blocks the credential when confirmed exhausted (server
exhaustedstatus or used fraction >= 1) with a liveresetsAt, and a live Fable 429 extends the reactive block to the confirmed tier reset instead of the 60s default. Unconfirmed rows (missing/expired reset, below cap) remain ranking hints only, preserving the false-100% guard. - Fixed Ollama/Ollama Cloud tool requests failing with HTTP 400 by rewriting boolean subschemas (
true/false) into a value-wideninganyOfunion of primitive types, stripping booleanadditionalProperties/unevaluatedProperties, and flattening nullabletypearrays before serializing tool parameters, so unconstrained fields still advertise "any JSON value" to grammar-constrained samplers (llama.cpp) instead of collapsing to an empty object. (#4488)
[16.3.5] - 2026-07-04
Added
OAuthCallbackFlownow serves aGET /launchroute on its loopback callback server that 302-redirects to the pending authorization URL, and exposes that short URL asOAuthAuthInfo.launchUrl. UIs can advertise it as a truncation-safe copy target (~30 chars) instead of the full authorize URL, so terminals narrower than the composed row cannot silently drop OAuth query parameters likecode_challenge_method=S256(#4418).- Preserved explicit
tool.strict === falseon OpenAI-family function tool payloads (openai-responses, openai-codex-responses, openai-completions) so backends that distinguishstrict: falsefrom an omitted flag stop over-filling optional arguments (#4336).
Fixed
- Fixed tool-call validation to strip stray trailing line terminators on schema-matching enum values and on well-known identifier fields (
path,paths,file,file_path,url,uri,title,label) before dispatch, keeping ordinary trailing spaces and content-carrying fields (content,input,code,command, etc.) intact (#4461).
[16.3.4] - 2026-07-03
Added
- Added support for Baseten as an AI provider
Changed
- Improved Claude usage reliability by removing proactive hard-blocking for Fable and Mythos tiers
Fixed
- Fixed Anthropic OAuth account rotation to exclude unreliable model-scoped Fable/Mythos weekly caps from proactive hard-blocking, ensuring they act only as ranking priority hints while still allowing reactive 429-fallback to rotate and reach serviceable siblings.
[16.3.3] - 2026-07-02
Added
- Added comprehensive tracking and credential-ranking support for Anthropic per-tier and weekly usage limits, including Claude Fable weekly caps. This prevents a single exhausted model-scoped cap from blocking the entire OAuth credential and improves credential selection based on drain-rate pressure.
Changed
- Updated Claude Fable reasoning replay to use bare text instead of wrapped thinking tags
Fixed
- Improved robustness of single-argument tool calls by automatically remapping mislabeled string arguments.
- Fixed Anthropic OAuth usage reporting to stop retrying on 429 rate-limit errors.
- Fixed usage cache to correctly persist null values during cold-start failure backoff windows.
- Fixed cursor-agent persisted transcripts losing tool-call structure for native execution tools, ensuring replayed tool results are correctly paired with their corresponding calls.
- Fixed OpenAI-compatible streaming usage parsing to prefer non-zero nested cached token counts when the root cached_tokens value is zero.
- Added automatic detection and remediation for custom proxies returning signature errors on Anthropic thinking blocks, allowing the client to automatically retry with unsigned blocks and prompt the user to adjust their configuration.
- Fixed potential hangs in GitLab Duo Workflow setup by adding proper timeout and abort signal handling to REST fetches.
- Fixed Cursor proxy tunnel setup hanging indefinitely by adding abort and timeout handling.
- Fixed Devin Connect streaming reader vulnerability to corrupt frame lengths by capping payloads at 16 MiB and throwing an envelope error immediately.
[16.3.1] - 2026-07-02
Changed
- Removed automated injection of reasoning suppression prompts in OpenAI responses
[16.3.0] - 2026-07-02
Added
- Added opt-in support for Anthropic's server-side fallback beta (server-side-fallback-2026-06-01) on the anthropic-messages provider, including support for AnthropicOptions.fallbacks and automatic filtering of fallback blocks during cross-provider message transformations.
Changed
- Improved stream healing for official first-party endpoints (Anthropic, OpenAI, and OpenAI Codex) by skipping leaked-thinking healing, preventing misfires on legitimate code blocks while maintaining healing for third-party gateways and custom base URLs.
- Updated CoreWeave Serverless Inference login instructions to clarify persisting COREWEAVE_PROJECT in shell startup files.
Fixed
- Fixed an issue where same-model Anthropic message replays incorrectly demoted unsigned thinking into textual content during API calls
- Fixed a performance issue where broker usage fetch failures were not cached, causing redundant network requests when the broker is offline.
- Fixed Xiaomi MiMo API key validation to use the supported mimo-v2.5 model.
- Fixed certificate verification errors for custom gateways behind private CA bundles by ensuring NODE_EXTRA_CA_CERTS is respected across all provider fetches.
- Fixed Claude Fable demoted-thinking replay to use markdown-italic assistant prose instead of tags, preventing context issues after model switches.
- Fixed OpenAI Responses replay errors (400 Bad Request) caused by missing reasoning items during history replay.
[16.2.13] - 2026-07-01
Fixed
- Fixed pre-5.4 OpenAI Codex models (
gpt-5.1-codex,gpt-5.3-codex,gpt-5.3-codex-spark) rejecting requests withUnsupported parameter: 'reasoning.summary' is not supported with this modelby gatingreasoning.summarybehind the same gpt-5.4 wire floor asreasoning.context: "all_turns".
[16.2.12] - 2026-07-01
Changed
- Improved streaming performance for Cursor and Devin providers by optimizing mid-stream tool-call argument parsing to prevent UI stalls when handling large payloads.
Fixed
- Fixed issues with tool call streaming where tool call IDs, partial JSON payloads, or late-arriving IDs could be lost, filtered, or incorrectly initialized.
- Fixed an issue where stream healing for leaked thinking blocks could replace live tool-call blocks with empty-id placeholders, breaking streamed tool arguments on Anthropic-compatible streams.
- Fixed an issue where stalled auth-gateway SSE responses could hang indefinitely in pi-native streams by ensuring first-event and idle timeout watchdogs are properly honored.
- Fixed cross-turn tool-call loops going undetected by adding a guard for consecutive identical tool calls. (#3971)
[16.2.11] - 2026-07-01
Fixed
- Fixed streaming UI glitches and resolved an issue where invalid empty tool call IDs were persisted in the chat history.
[16.2.10] - 2026-06-30
Added
- Added streaming support for keyed parameter argument deltas in XML-family in-band tool call scanners (Anthropic, DeepSeek, XML, Minimax)
Changed
- Improved native tool-call passthrough in
wrapInbandToolStreamto accurately mirror live streaming IDs, arguments, and partial JSON states from the underlying provider
Fixed
- Fixed a bug where tool calls with empty or missing IDs were not detected as malformed, causing API validation failures (e.g., 400 errors with Anthropic) on subsequent requests
- Raised Gemini header runaway threshold to prevent premature interruption of complex reasoning loops
- Fixed leaked
```thinkingfences with nested language-tagged Markdown code blocks so inner fences remain inside structured thinking instead of leaking as visible reply text.
[16.2.9] - 2026-06-30
Added
- Added
OAuthCallbackFlowOptions.allowPortFallbackto allow disabling random-port fallback, enabling strict port enforcement and early configuration errors for OAuth flows with static redirect URIs.
Changed
- Improved
OAuthCallbackFlowport conflict error messages to include the busy port, configured redirect URI, and actionable remediation steps.
Fixed
- Fixed an issue where malformed tool-call JSON from local Ollama or llama.cpp models was incorrectly retried as generic 500 errors, now surfacing a clear recovery message.
- Fixed a race condition in OAuth callback flows where abort signals triggered before the callback listener was registered were ignored.
[16.2.7] - 2026-06-30
Added
- Added service tier support for Google Gemini and Vertex AI, including model-specific service tier configurations via ServiceTierByFamily.
- Added Google Vertex AI Interactions API support for Gemini 3+ models by default, with automatic fallback to :streamGenerateContent and a useInteractionsApi: false option to force standard generation.
- Added support for explicit Vertex bearer access tokens via GOOGLE_CLOUD_ACCESS_TOKEN or CLOUDSDK_AUTH_ACCESS_TOKEN environment variables.
Changed
- Updated service tier logic to use per-provider configurations instead of global scopes.
- Refactored priority request billing and accounting to better align with specific provider capabilities.
- Updated API key resolution precedence so explicit environment variables (e.g., GEMINI_API_KEY) override stored or broker-migrated static API keys, while deliberate OAuth logins still take highest precedence.
Fixed
- Improved Vertex AI reliability by automatically falling back to global endpoints on 404 errors.
- Fixed safety setting application for Google Vertex AI models.
- Fixed Kimi Code's Anthropic-compatible request path to keep thinking enabled and downgrade forced tool choice for Kimi K2.7 Code title generation.
- Fixed leaked reasoning fences (such as ```thinking or ) across all providers by splitting them into structured thinking blocks during streaming.
- Fixed Codex requests failing with unsupported all_turns errors on older models (gpt-5.1 and gpt-5.3) by gating the reasoning.context: "all_turns" default to gpt-5.4+ models.
[16.2.6] - 2026-06-29
Fixed
- Fixed Antigravity usage reporting to correctly infer daily and weekly quota windows from unlabeled reset-only rows, preventing Cloud Code Assist payloads from collapsing these counters into the default category.
[16.2.5] - 2026-06-28
Fixed
- Fixed Google and Cloud Code Assist streams that end without a finish reason (dropped connections or truncated responses) being treated as fatal; they are now classified as transient so the coding agent automatically retries.
[16.2.4] - 2026-06-28
Added
- Enabled freeform tool patch support for Azure OpenAI and Codex models
Fixed
- Fixed usage reporting for Antigravity and Z.AI to correctly surface and preserve distinct quota windows (daily, weekly, monthly) instead of collapsing or duplicating them
- Fixed an issue where
/usage showreturned "No usage data available" when using a custom proxy base URL for Codex - Fixed OpenAI stream read errors being incorrectly classified as non-transient, enabling the coding agent to automatically retry after recoverable stream failures
[16.2.3] - 2026-06-28
Changed
- Enabled automatic removal of leaked reasoning tags for all models
- Prevented reasoning text duplication when models emit both structured and inline thinking
- Defaulted reasoning context to all turns for all Codex requests.
Fixed
- Enabled freeform tool patch support for Azure OpenAI and Codex models.
- Fixed an issue where the
/usage showcommand returned "No usage data available" when using a custom proxy base URL for Codex.
[16.2.2] - 2026-06-27
Added
- Added a comprehensive, public-facing error module exported via the "./error" path, featuring structured error classification, provider-specific HTTP error classes (e.g., Anthropic, OpenAI, Gemini), OAuth/Auth-specific errors, rate-limit utilities, and retryability predicates.
Changed
- Updated OpenAI Codex defaults to increase default text verbosity to medium, enable detailed reasoning summaries by default, and include all turns in the reasoning context by default.
- Updated the OpenAI Codex WebSocket transport to resolve its configuration (via PI_CODEX_WEBSOCKET_* environment variables) once at startup rather than re-parsing on every request.
- Enhanced cross-model reasoning recovery and preservation to render demoted reasoning in the target model's canonical inline thinking dialect (such as Gemini's thinking fence or standard think tags) to prevent leaking inert context or control tokens into history.
- Broadened the leaked-thinking stream healer to recover reasoning emitted in any dialect's canonical idiom (including Gemini, Gemma, Harmony, and scratchpads) and route them to thinking events instead of raw markup.
- Implemented automatic retry logic for detected thinking-loop stalls to improve response reliability.
- Hardened stateful delta chaining to ignore transient streaming bookkeeping symbols during structural equality checks, preventing unnecessary full-transcript replays.
Fixed
- Fixed preservation of OpenAI Responses assistant message phase values across auth-gateway parsing, streaming, and history replay, ensuring GPT-5.4/GPT-5.5 intermediate updates and final answers retain their original phase labels.
Removed
- Removed Pi dialect support and related serialization/parsing logic.
[16.2.0] - 2026-06-27
Breaking Changes
- Removed the
@oh-my-pi/pi-ai/utils/json-parsemodule. The JSON repair and parsing helpers (repairJson,parseJsonWithRepair,parseStreamingJson,parseStreamingJsonThrottled) have been moved to@oh-my-pi/pi-utilsto be shared across utilities.
Added
- Added the GitLab Duo Agent provider (
gitlab-duo-agent) and built-in implementation, renaming the existing AI Gateway proxy provider to "GitLab Duo Non-Agentic" (gitlab-duo). - Added GitLab Duo Workflow provider support, featuring OAuth login via the official VS Code OAuth application, automatic project discovery, and automatic session-time namespace Duo settings enablement.
- Added runaway detection for Gemini models to interrupt streams stuck in excessive planning steps.
- Added a per-provider in-flight request limiter for LLM streams, shared across local OMP processes and configurable via
maxInFlightRequests. - Added a
creditsfield toUsageResetCreditsto display when banked rate-limit resets expire, with support for OpenAI Codex usage details.
Changed
- Optimized GitLab Duo Agent and Workflow providers to use an inline custom "ambient" flow with MCP-only agent privileges, registering MCP tools under their bare names.
- Improved GitLab Duo Agent context management and auto-compaction by lowering the soft overflow threshold to 1 MB and stripping redundant bytes (such as tool-call UUIDs and escaped JSON) from the goal transcript.
- Enhanced GitLab Duo Agent prompt engineering to render replayed tool calls as past-tense records, reducing model confusion and preventing the model from mimicking historical markers.
- Added caching for discovered GitLab Duo Agent root namespaces per account to avoid redundant discovery requests.
Fixed
- Fixed various GitLab Duo Agent and Workflow stability issues, including infinite tool-call loops, connection hangs on half-open WebSockets, and unhandled step-limit or generic server-side failures.
- Improved GitLab Duo Workflow routing, namespace resolution, and project-path handling, ensuring correct numeric ID resolution and support for self-managed GitLab relative install base paths.
- Fixed GitLab Duo Workflow checkpoint streaming to correctly map reasoning entries to thinking blocks, preserve tool boundaries, and accurately report token usage.
- Fixed
AuthStorage.loginto only synthesize manual-code paste prompts for paste-code providers, preventing terminal-blocking races on loopback OAuth flows. - Fixed llama.cpp compatibility by downgrading named forced
tool_choiceobjects to the string"required"in the chat-completions encoder. - Fixed
omp usageomitting Ollama and Ollama Cloud accounts by registering placeholder usage providers. - Fixed Gemini reasoning-runaway detection to expose a dedicated thought-summary header guard to interrupt streams stuck in planning loops.
Removed
- Removed legacy GitLab Duo Workflow
chatandsoftware_developmentflow paths and the non-MCP action bridge in favor of the inline customambientflow.
[16.1.23] - 2026-06-26
Added
- Added a third streaming thinking-loop detection heuristic to catch "progress-lexicon stalls" where models endlessly reshuffle motivational filler without introducing new vocabulary or concrete technical references
- Added branded wordmark and logo animation to authentication flow
- Added a third streaming thinking-loop detection shape — a progress-lexicon stall — alongside verbatim tail repetition and near-duplicate (trigram) segments. It catches reasoning-summarizer loops that reshuffle the same motivational filler ("just doing it, pushing ahead, maintaining momentum") into fresh word order every paragraph: word-trigrams never cluster, but a run of substantial segments that recycle the recent vocabulary and introduce no new concrete reference (path / identifier / code-span) trips the guard. Summarizer title/heading lines (
**Bold Title**,## Heading) are stripped before analysis so their ever-changing wording cannot mask the stall by inflating novelty. Calibrated against 537k real non-Gemini reasoning blocks (zero false positives at novelty floor 0.2 / run length 8; the real loop sustains runs of 10+). - Added CoreWeave Serverless Inference provider login support via
COREWEAVE_API_KEYandWANDB_API_KEYfallback.
Changed
- Redesigned the OAuth callback page (
oauth.html) to match the oh-my-pi web brand language: OKLCH purple-tinted dark neutrals, magenta→iris→cyan brand gradient on the wordmark, frosted-glass card over an ascii grid backdrop, and a colored status halo around the success/error icon. All assets are inlined; the__OAUTH_STATE__injection contract and success/error JS logic are unchanged.
Fixed
- Fixed local llama.cpp (and any local OpenAI-compatible server rendering the Qwen3.6+ chat template) re-processing the full prompt every new user message even with
replayReasoningContentenabled (#3541 follow-up to #3528). Sendingreasoning_contentalone wasn't enough: Qwen3's chat template strips<think>...</think>from any assistant turn whose index is<= last_query_index, so the moment a new user message (the user's next prompt, or the auto-learn capture-at-stop nudge) lands, every prior assistant turn becomes "older" and is re-rendered without the<think>block — diverging from the generation tokens still in the slot's KV cache. The chat-completions encoder now emitspreserve_thinking: truefor Qwen thinking dialects on local servers, route-split the same way the existingenable_thinkingemission is: theqwendialect rides the top-level field (llama.cpp's--jinjahook and Alibaba Cloud Model Studio's compatible-mode), theqwen-chat-templatedialect (NVIDIA NIM, vLLM/SGLang's chat-template-kwargs path) rides onlychat_template_kwargs.preserve_thinkingbecause NIM's request schema isadditionalProperties: falseand rejects unknown top-level fields (#2299). The emission is hoisted above thereasoning.enabledgate so it fires for THREE cases the original gating missed: (1) runtime-discovered local Qwen models that ship withreasoning: falsebecause the upstream/v1/modelsdoesn't advertise the capability (same gotcha #3532 fixed forreplayReasoningContent), (2) caller-disabled reasoning (/think off) — the kwarg is a history-rendering knob, not a per-turn thinking switch, and the slot still holds<think>tokens from earlier turns, and (3) forced-tool-choice / DeepSeek-style auto-disable. Qwen3.6+ then renders<think>...</think>for every assistant turn regardless of position, and the next-turn render matches the cached generation tokens. (#3541)
[16.1.22] - 2026-06-26
Fixed
- Fixed llama.cpp / LM Studio / vLLM (and any local OpenAI-compatible server on a loopback or RFC1918 baseUrl) re-processing the full prompt on every assistant continuation when the prior turn produced
reasoning_content: theopenai-completionsencoder dropped the preservedthinkingblock on re-serialization for compat profiles withoutrequiresReasoningContentForToolCalls/thinkingFormat: "zai", so the chat template re-rendered the assistant turn without<think>…</think>and the rendered tokens diverged from the slot's KV cache state. The auto-learn capture-at-stop nudge made it reproduce on every turn. The encoder now replays preserved thinking asreasoning_content(honoring the streamed signature when it identifies a recognized wire field —reasoning_content/reasoning/reasoning_text— and falling back to the configuredreasoningContentFieldfor opaque signatures) whenever the newcompat.replayReasoningContentflag is set, and the cross-APItransformMessagespredicate (openAICompletionsReplaysUnsignedThinking) honors the same flag ahead of themodel.reasoninggate so a switch into a discovered local target (where the spec carriesreasoning: falsebecause the upstream/modelsendpoints don't advertise the capability) still preserves the prior turn's thinking block as signature-stripped reasoning instead of demoting it to conversation text. The chat-template-rendered prefix stays byte-stable across turns and llama.cpp's prefix KV cache survives. (#3528)
[16.1.21] - 2026-06-26
Fixed
- Restored the
pollOAuthDeviceCodeFlowexport from@oh-my-pi/pi-ai/oauthso legacy provider extensions can reuse the host OAuth device-code poller. (#3508)
[16.1.20] - 2026-06-25
Fixed
- Fixed Ollama/Ollama Cloud native chat responses that finish with
done_reason: "length"and no assistant content surfacing as a normal empty stop; they now become a context-window error instead of entering empty-stop retry recovery. (#3464) - Fixed direct Anthropic Claude Sonnet/Haiku 4.5 requests serializing
output_config.effort. The catalog classification (packages/catalog/src/model-thinking.ts) drove theanthropic-budget-effortbranch inbuildParams, which Anthropic's first-party Messages API rejects on Sonnet/Haiku 4.5 with HTTP 400This model does not support the effort parameter.Sonnet/Haiku 4.5 now use plainthinking.budget_tokens; Opus 4.5 still emitsoutput_config.effortbecause Anthropic supports it there. (#3497)
[16.1.19] - 2026-06-25
Fixed
- Fixed Ollama/llama.cpp chat payloads serializing user-attributed mid-conversation developer messages (auto-learn capture nudge, advisor cards, file-mention companions) as
systemturns; they now serialize asuserso llama.cpp can reuse the warm prompt prefix instead of forcing full re-processing. Agent-owned developer reminders (attribution: "agent"— empty/unexpected-stop retries, checkpoint rewind warning, todo reminders) keep theirsystempriority. (#3456) - Fixed prior-turn reasoning being lost on cross-API provider switches: when a session moved from an Anthropic-compatible 3p endpoint to an OpenAI-compatible one (Z.AI Anthropic → Z.AI OpenAI, Kimi Anthropic → Kimi OpenAI, DeepSeek, OpenCode-hosted reasoning models, or any custom
models.yamlswitch that crosses API types), the cross-API path oftransformMessagestext-demoted every priorthinkingblock, so the next request shipped the reasoning chain as plain conversationcontentinstead of structuredreasoning_content— losing it as reasoning context and re-billing it.convertMessagesnow threads the request-time resolved compat intotransformMessages, which preserves the prior reasoning as a native, signature-strippedthinkingblock whenever that resolved target acceptsreasoning_contentas a continuation hint (requiresReasoningContentForToolCalls— including thewhenThinkingpolicy OpenCode reactivates for thinking-on requests, #1071/#1484 — orthinkingFormat: "zai"); theopenai-completionsencoder surfaces those blocks viareasoningContentField, with a new branch for Z.AI-format hosts (Z.AI, Zhipu, Moonshot Kimi, Xiaomi MiMo) that accept but don't require the field. Targets that can't replay unsigned reasoning (encrypted reasoning blobs, signed thought parts, non-reasoning models, thinking-disabled OpenCode) still text-demote so the reasoning survives as conversation context. (#3437, #3439 by @roboomp; #3433, #3434) - Fixed Bedrock cross-region inference profiles routing to
us-east-1regardless of their geo prefix: a profile such aseu.anthropic.claude-…(orapac./au./jp.) sent to the hardcodedus-east-1endpoint returned HTTP 400The provided model identifier is invalid.streamBedrocknow derives the runtime region from the profile's geo prefix — honoring an ambientAWS_REGION/AWS_DEFAULT_REGIONonly when it can serve that geo and falling back to the geo's default region otherwise — while explicit per-request and ARN-embedded regions still win and region-agnosticglobal.profiles stay unchanged. - Fixed malformed tool calls (empty
name) wedging entire sessions in HTTP 400 loops: when a model occasionally emits{ "name": "", "arguments": "{}" }(observed: GLM-5.2 + thinking on long turns), the agent rejected the call at execution time withTool not found, but the malformed block plus its errortoolResultstayed in conversation history and every subsequent request 400'd ontool_use.name/tool_calls[i].function.namevalidation until the user ran/clear.transformMessages— the canonical sanitize boundary every provider passes through — now dropstoolCallblocks with empty/whitespacename, pairs them with theirtoolResultmessages only inside the same assistant→tool-result window (per-id FIFO queue cleared at non-result boundaries, so stale malformed calls without a result cannot consume later valid duplicate-id outputs), and drops the assistant turn when it has no replayable content left. Defensive (provider-agnostic, fires regardless of model), idempotent (no-op on a clean history), and self-healing (one round-trip after the fix lands sanitizes an already-poisoned session). (#3458)
[16.1.18] - 2026-06-25
Added
- Added
listOAuthAccountsfor retrieving a read-only list of stored OAuth account identities - Added
getOAuthAccessAtto resolve an OAuth token exclusively for a specific account position
Changed
- Refactored OAuth token persistence and disable logic to use stable credential IDs instead of positional indices to prevent race conditions during concurrent updates
- Updated OAuth failure classification to treat 403 status codes, rate limits, and network errors as transient, preventing unnecessary credential invalidation
Fixed
- Fixed Codex Responses Lite staying enabled for image prompts, which caused GPT/Codex image turns to be rejected as
Invalid value: 'input_image'; image-bearing Codex requests now fall back to the full Responses transport. (#3421) - Fixed the auth-broker background refresher disabling OAuth credentials unconditionally (
disableCredentialById) on a definitive refresh failure, so a credential another process or a fresh login rotated mid-refresh could be torn down even though the stored row already held a valid token. The definitive-failure teardown now happens insideAuthStorage.refreshCredentialByIdvia the same compare-and-set the in-stream and usage-probe paths use — it disables only when the persisted row still matches the credential the refresh actually attempted, and reloads on a CAS loss; the refresher now only logs. - Fixed OAuth refresh persisting the rotated token by a positional index captured before the refresh
await. A concurrent disable could reorder or shrink a provider's credential array while the refresh was in flight, landing the new token on the wrong row (or silently dropping it) and leaving accounts with a stale refresh token that failed — and was then disabled — on the next cycle. Refresh persistence, selection-index resync, and CAS-disable now address the row by id acrossforceRefreshCredentialById, candidate preflight, and in-stream selection (#replaceCredentialById/#disableCredentialByIdIfMatches). - Fixed
isDefinitiveOAuthFailuretreating a bare HTTP 403 (and genericunauthorized/ access-token-expired wording) as a definitive credential failure, which permanently disabled healthy OAuth accounts on WAF, egress rate-limit, permission, and account-verification responses. Bare 403, rate limits (429), gateway/5xx, and more network errors (ECONNRESET,ETIMEDOUT,EAI_AGAIN, …) are now classified transient; only explicit dead-grant errors (invalid_grant,invalid_token,unauthorized_client, revoked,refresh token … expired) or a bare 401 tear the credential down.
[16.1.17] - 2026-06-24
Added
- Added provider-level
notes?: string[]field toUsageReportfor disclaimers that apply to every limit (e.g. "OMP-observed spend only"). The field is declared in both theusage.tsschema and the auth-broker wire schema copy so it survives the"+": "reject"deserialization gate. (#3268)
Fixed
- Moved the OpenCode Go "OMP-observed spend only" disclaimer from per-limit
notesto provider-levelnotes, so it renders once per provider instead of duplicating across every account × window. (#3268) - Fixed Anthropic rate-limit header usage cache entries retaining legacy missing account metadata after refresh.
- Fixed Anthropic-compatible budget-effort models dropping the selected effort before request serialization, so
output_config.effortis emitted alongsidethinking.budget_tokenswhen model metadata declaresmode: "anthropic-budget-effort". - Fixed
anthropic-messagessilently dropping caller-suppliedAuthorization/X-Api-Keyfrommodel.headersandANTHROPIC_CUSTOM_HEADERS, blocking custom proxy auth schemes. Non-OAuth requests now honor the caller's value (matchingopenai-responses); the lower-level client also suppresses itsX-Api-Keyadd when a customAuthorizationis supplied for a non-official endpoint so the proxy receives a single credential. OAuth bearer + Cloudflare AI Gateway keep their pre-existing enforced auth headers. (#3391) - Fixed Ollama Cloud
num_predictignoring the provider's 65536 output-token cap so stalemodels.dbrows (or custommodelOverridesre-enabling output caps) that carriedmaxTokens: 1048576from a pre-omitMaxOutputTokens catalog 400'd every request withmax_tokens (1048576) exceeds model's maximum output tokens (65536) for model deepseek-v4-pro. The Ollama provider now clampsnum_predictfor anyollama-cloudrequest at the documented 65536 cap before sending, independent of the cached spec'smaxTokensand on top of the existingomitMaxOutputTokenspolicy — so the request stays valid even when the load-time policy never normalized the spec. Self-hostedollamatraffic is unaffected. (#3392) - Fixed OpenRouter Anthropic models on the Responses path omitting
cache_control, so prompt caching engages without forcing Chat Completions. (#3397) - Fixed OpenRouter Anthropic Responses follow-up requests replaying prior reasoning items with stale signatures, which caused HTTP 400
Invalid signature in thinking blockerrors after a thinking turn. (#3399) - Fixed OpenRouter Anthropic models on the Responses path omitting
cache_control, so prompt caching engages without forcing Chat Completions.cacheRetention: "long"now upgrades the breakpoint tottl: "1h". (#3397)
[16.1.16] - 2026-06-23
Fixed
- Fixed Anthropic-compatible thinking requests sending replayed thinking blocks without
context_management.keep: "all", preserving multi-turn reasoning context for API-key providers. API-key requests now also advertise the requiredcontext-management-2025-06-27beta header so the field is honored instead of rejected. Injected SDK clients, GitHub Copilot's Anthropic proxy, and Vertex rawPredict are excluded because this code path cannot add the beta to caller-owned clients, Copilot strips Anthropic betas and demotes thinking blocks to text upstream, and Vertex expects betas in the JSON body rather than the Anthropic HTTP beta header. (#3288) - Fixed OpenRouter Responses native history replay leaking Gemini reasoning item
formatmetadata back into follow-up requests, which caused HTTP 400 rejections while preserving encrypted reasoning replay.
[16.1.15] - 2026-06-22
Fixed
- Fixed API-key
/loginproviders replacing sibling credentials instead of appending new keys for the same provider. (#3265) - Fixed OpenAI Codex OAuth account rotation for quota failures that surface as bare HTTP 429 or
insufficient_quota, so pre-content failures temporarily block only the exhausted credential and retry a healthy sibling. The 429 status-only fallback applies only to absent/opaque bodies; informative transient bodies (Too many requests,Service overloaded 529,Please retry in 5s, …) defer toparseRateLimitReasonand stay in the provider's own backoff layer instead of burning sibling credentials. (#3231)
[16.1.14] - 2026-06-22
Added
- Added proxy support for model providers via
PI_PROXYandPI_PROXY_<PROVIDER>variables - Added
NO_PROXYenvironment variable support for bypassing proxy configuration - Added support for Sakana AI provider
- Added Sakana AI login and request base URL support for
SAKANA_*/FUGU_*environment variables
Changed
- Consolidated API key authentication logic across registry providers
- Disabled parallel tool calls for Devin provider requests
Fixed
- Improved proxy bypass logic to correctly handle private IP ranges and local metadata services
- Enhanced memoization for proxy environment variable lookups to improve performance
[16.1.13] - 2026-06-22
Added
- Added support for Devin as a provider
Changed
- Updated tool call arguments to use
Record<string, unknown>andunknownfor tool results
Fixed
- Fixed OpenAI Responses native history replay dropping failed/incomplete image generation calls instead of resending their transient
ig_...item IDs, preventing follow-up requests from failing with404 Item with id ... not found. (#3225) - Fixed
/login fireworksrejecting validfw_…keys withFireworks API key validation failed (500): Error listing deployed models. The validator pinged/inference/v1/models, which Fireworks serves from the per-account deployment registry and 500s for accounts without active deployments. Login now hits the static control-planeList Modelscatalog (GET /v1/accounts/fireworks/models?filter=supports_serverless=true&pageSize=1) — the same endpoint discovery already uses — so authentication no longer depends on the caller's deployment state. (#3219)
[16.1.11] - 2026-06-21
Fixed
- Fixed OpenAI Responses native history replay leaking image generation provider-only fields into the next request, which made OpenAI-compatible proxies reject
pitool-calling sessions withUnknown parameter: input[1].action. (#3201) - Fixed a stream thought-leakage issue for
gemini-3.5-flashwhere the model's internal reasoning JSON could leak into the visible text stream. The stream parser now uses a brace-balanced counting algorithm to accurately slice and discard the leading thought JSON block, with a robust fallback for unescaped double quotes, dynamic tool-name derivation, and preservation of subsequent text deltas without triggering empty-response retries.
[16.1.10] - 2026-06-21
Changed
- Improved JSON robustness by replacing external dependency with a custom, high-performance parser
- Strengthened streaming JSON parsing to prevent non-finite numbers from surfacing as
undefined/NaN - Configured JSON parser to reject JS-specific
NaNandInfinityvalues for tool arguments - Replaced the JSON repair/parse helpers (
parseJsonWithRepair,parseStreamingJson) with a single from-scratch tolerant parser (RelaxedJson) that accepts single-quoted strings, unquoted object keys, trailing/stray commas,//and/* */comments, PythonTrue/False/None, raw control characters, invalid escapes, and unescaped apostrophes ('it's'). Final parsing still throws on truncated/garbage input (so a malformed tool call is skipped rather than executed with half-formed args) and rejects JS-onlyNaN/Infinity; streaming parsing stays non-throwing and rolls back incomplete trailing tokens instead of surfacingundefined/NaN. The Cursor provider's ad-hoc regex + JSON5 tool-argument parser now routes through the shared parser.
Fixed
- Fixed tool call ID normalization for Anthropic-compatible models
- Fixed Anthropic Messages replay sanitizing malformed tool-call IDs, including aborted native tool calls with empty IDs, so retries no longer send invalid
tool_use.id/tool_result.tool_use_idpairs. - Fixed the Codex Responses WebSocket transport attributing a prior turn's output to the current one on a reused connection: a trailing/duplicate frame from a cleanly-completed previous response that slipped past the queue drain could be consumed as this request's terminal (ending the turn with empty output) or as a stale tool call. Frames are now keyed by
response.id— a frame carrying the previous response's id is dropped, and one carrying a third id (or a regressedsequence_number) fails closed so the turn retries instead of mixing two responses' streams. Idless frames (deltas, the rate-limit/metadata preamble,response.created-less streams) still pass through, matching upstream codex-rs. - Fixed
transformMessagespulling an earlier, orphaned tool result onto a later tool call that reused the same id (left behind when compaction folded the originatingtool_useinto a summary). The pending-call flush now pairs each call with a result positioned after its assistant turn, so a reused id surfaces its own output rather than a prior turn's. - Fixed DashScope 429 rate-limit messages that mention authorization being classified as credential failures, preventing valid API keys from being invalidated after throttling. (#3172)
- Fixed OpenCode Go
401 Insufficient balancequota errors being treated as unknown failures instead of usage-limit errors, restoring credential rotation and fallback chains. (#3169)
Removed
- Removed the
partial-jsondependency; streaming JSON parsing now uses the in-houseRelaxedJsonparser.
[16.1.9] - 2026-06-21
Added
- Added
llama.cppto the interactive/loginprovider list, accepting an optional API key while defaulting to local no-auth mode.
Changed
- Optimized generated AI tool schemas by collapsing verbose
anyOfunions into standardenumtypes
Fixed
- Fixed tool-call argument validation dropping nested keys that were accidentally double-encoded
- Fixed the
moonshotprovider being locked to the international Kimi host (api.moonshot.ai): OpenAI-completions requests now honor aMOONSHOT_BASE_URLoverride so users can reach the Kimi China platform (api.moonshot.cn), which rejects keys issued for the international endpoint. (#2883) - Fixed tool-call argument validation dropping fields whose object keys were accidentally JSON-encoded a second time (e.g.
{ "\"op\"": "done" }), which surfaced as spurious missing-required errors. A schema-agnostic pre-validation pass now recursively unwraps such double-encoded keys — through arrays and nested objects, and again after a JSON-string container is parsed — before the unrecognized-key repair can delete them.
Removed
- Removed the
setNextRequestDebugPath,clearNextRequestDebugPath, andgetNextRequestDebugPathutility functions for request debugging, as request/response recording now relies exclusively on thePI_REQ_DEBUGenvironment variable. - Removed Wafer Pass (
wafer-pass) login support; Wafer Serverless remains available aswafer-serverless.
[16.1.8] - 2026-06-20
Changed
- Changed OpenAI Responses and Codex Responses custom grammar tool requests to leave
parallel_tool_callsunset instead of forcing serial tool calls; CodexresponsesLitestill disables parallel tool calls when tools are present.
Fixed
- Fixed Bedrock
/btwand other no-tool ephemeral turns failing after prior tool calls by sending the required sentineltoolConfigwhenever replayed history containstoolUse/toolResultblocks. (#3124) - Fixed Anthropic Messages pre-content TLS
bad record MACserver transport errors surfacing before the provider retry loop exhausts its budget. (#3134) - Fixed API-key login flows replacing existing stored keys for the same provider, so providers such as NVIDIA NIM can keep multiple active keys available for session-level rotation. (#2923)
- Fixed
openai-codex-responsesforwarding sampling controls (temperature,top_p,top_k,min_p,presence_penalty,repetition_penalty) into the Codex request body — the ChatGPT-subscription Codex backend rejects each of them with a 400{"detail":"Unsupported parameter: temperature"}, so any caller setting non-defaultStreamOptionssaw every turn fail. The provider now drops the full sampling set (matching codex-rs), and the auth-gateway's defensive strip on bothbuildStreamOptionsand the pi-native path was widened from{temperature, topP}to the same set plusstopSequences/frequencyPenalty. (#3117) - Fixed Anthropic Messages retry classification for transient TLS/server-error failures such as
tls: bad record MAC (type=server_error). These pre-content transport blips are now retried inside the provider loop before the session sees an error banner.
[16.1.4] - 2026-06-19
Added
- Added bounded auto-retry for empty assistant completions specifically to the OpenAI Responses provider
- Added bounded auto-retry for empty assistant completions across the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages providers. A benign terminal stop that streamed no content and billed no output tokens — the signature of a flaky OpenAI-/Anthropic-compatible gateway that intermittently 200s with an empty body — is now retried up to twice with exponential backoff (honoring
providerRetryWait) before being surfaced, instead of silently stalling the agent loop. Retries fire only before any content streams, so live streaming (including thinking) is never delayed, retried, or duplicated.
Fixed
- Fixed the Antigravity (
google-antigravity) request builder droppinglabels.model_enumwhen the wire profile does not declare one. Required for Claude 4.6 ids whoseAntigravityModelWireProfilecarries onlymaxOutputTokens(no capturedmodel_enum); the label is now emitted only when the catalog defines it. (#3067)
[16.1.3] - 2026-06-19
Added
- Added regression test pinning that
openai-completionsemits athinkingblock forreasoning_contentdeltas even whendelta.contentis explicitly JSONnull(the DeepSeek-format dual-key pattern used by custom GLM/Qwen reasoning providers). See #2996.
Changed
- Improved the thinking loop guard to treat assistant text loops as retryable errors
- Refined text normalization logic to reduce false positives in the thinking loop detector
Fixed
- Fixed Ollama chat requests sending image payloads to text-only models. Image blocks are now omitted and replaced with the standard non-vision placeholder for models without vision support, while vision-capable Ollama models continue to receive images. (#3009 by @serverinspector)
- Fixed
SqliteAuthCredentialStore.close()leaking one-off prepared statements created by inlinethis.#db.prepare()calls in#authCredentialsTableExists,#readAuthSchemaVersion,#inferAuthSchemaVersion,#migrateAuthSchemaV0ToV1,#backfillCredentialIdentityKeys, andupdateAuthCredential. Each statement is now wrapped intry/finallywithstmt.finalize(), and theclose()method finalizes#insertUsageCostStmtand#listUsageCostsStmtwhich were previously missed. This caused EBUSY on Windows when tests tried to delete temp dirs containing open SQLite handles.
[16.1.2] - 2026-06-19
Added
- Added improved JSON repair capabilities for Anthropic tool arguments
- Added authentication broker discovery to sync credentials between local SQLite and remote state
Fixed
- Improved error feedback and transparency for malformed Anthropic tool call arguments
- Added automatic fallback for unsupported OpenAI reasoning effort levels
- Improved reliability when handling invalid reasoning parameter errors across OpenAI-compatible APIs
- Fixed OpenAI-compatible Chat Completions, Responses, and Azure Responses requests to retry once with the nearest provider-supported reasoning effort when an endpoint rejects
xhigh/minimal-style effort values.
[16.1.0] - 2026-06-19
Added
- Added utility functions to strip schema descriptions for optimized LLM context usage
[16.0.10] - 2026-06-18
Added
- Replaced the old legacy XML-ish
piowned tool-calling dialect with the new sigil-delimited format (§call header with inlinekey=valuescalars,«…»verbatim body fence for the dominant string argument,¤reasoning,‡‡tool result) using single-token markers that never occur in source code. Verbatim fences escalate Markdown-style (««…»») so re-rendered history never collides with payload content, and the scanner gates a bare§on an exact known-tool name to avoid swallowing prose. Round-trips and streams through the existing scanner contract at ~46% fewer tokens than the legacy format on typical calls; selectable viatools.formatorPI_DIALECT=pi.
Changed
- Updated
pidialect formatting to use a token-frugal, sigil-delimited format (§,¤,‡‡) - Updated
pidialect body fences to automatically escalate when content contains fence markers - Changed
pidialect tool results response format to‡‡blocks
Fixed
- Fixed Bedrock application inference profile ARNs to route requests to the ARN's region instead of the default Bedrock runtime region. (#3004)
[16.0.9] - 2026-06-18
Fixed
- Fixed OAuth login replacing all other active accounts for the same provider, allowing multiple OAuth accounts to coexist concurrently.
- Fixed legacy
api_keycredentials not being replaced/disabled atomically upon upgrading to OAuth login. - Fixed a logic issue where AuthStorage lost session-to-credential stickiness upon CLI restarts, causing cold-starts for server-side prompt cache (KV cache) and wasting tokens.
- Fixed GitHub Copilot Responses requests rejecting image inputs that carry the
detail: "original"hint with an HTTP 400 by degrading the hint to"auto"for hosts that do not support it; other hosts still preserve native-resolution frames (snapcompact). (#2822)
[16.0.8] - 2026-06-18
Fixed
- Improved reliability of auth-broker snapshot loading by implementing a robust manual schema check
- Fixed MCP tool argument validation to drop optional empty-string parameters before schema validation, matching the existing optional null handling and avoiding pattern/type failures for omitted model-filled fields. (#2981)
- Fixed API-key credential replacement to hard-delete superseded disabled
api_keyrows soauth_credentialsdoes not grow indefinitely after key rotation. (#2941) - Fixed Cursor provider streaming to close text blocks before tool calls so post-tool text opens a new content block and TUI transcript cards render inline instead of grouped near the bottom. (#2924)
[16.0.7] - 2026-06-18
Changed
- Switched Google OAuth callback hostname from
localhostto127.0.0.1to prevent IPv6 loopback fallback delays and proxy routing interception.
Fixed
- Fixed OpenCode Go usage reporting to synthesize
/usagelimits from OMP-observed request costs for the 5h, weekly, and monthly provider caps. (#2942) - Fixed MiniMax Anthropic-compatible requests to serialize adaptive thinking without an invalid Anthropic
output_config.efforttier (#2928).
[16.0.6] - 2026-06-18
Added
- Added support for ArkType schemas as tool parameters alongside existing Zod schemas
- Added
getOpenRouterHeadersutility to export standard OpenRouter integration headers
Changed
- Expanded thinking loop detection guard to also cover DeepSeek models (family, provider, or id matches).
- Extended loop guard to monitor assistant response prose (via
text_deltaevents) in addition to thinking logs, customizable via request options. - Modified loop guard error reporting to emit a non-retryable partial content block containing the accumulated streamed text if a loop is detected after response prose has started streaming, preventing unsafe agent session rollbacks.
- Migrated internal wire-schema validation (auth-broker, Anthropic Messages request, OpenAI Chat/Responses requests, and /v1/usage shapes) from Zod to ArkType
- Replaced the dedicated
xai-responsesprovider with a unifiedopenai-responsespath that handles xAI-specific reasoning effort stripping dynamically - Updated OpenAI Responses stream handling to throw a clearer error message when a stream closes without a terminal response event
- Consolidated shared OpenAI-compatible routing and strict-tool fallback helpers across Chat Completions and Responses providers.
- Consolidated the OpenAI-family provider stack: merged
openai-responses-sharedintoopenai-sharedand removed the now-deadopenai-responses-sharedre-export shim; folded the three duplicatedservice_tierrequest blocks and the per-provider wire model-id transform into sharedapplyOpenAIServiceTier/applyWireModelIdTransformhelpers; moved residual provider-name wire-quirk checks (DeepSeek special-token strip, cumulative reasoning deltas, Ollama empty-length context error, OpenAI tool-call-id cap, Fireworks thinking drop, OpenRouter/OpenAI Responses request fields) into resolved compat fields; shared the Responses stream per-block accumulation helpers plus the terminal pending-tool-call finalization (finalizePendingResponsesToolCalls) and toolUse/pause stop-reason promotion (promoteResponsesToolUseStopReason) betweenprocessResponsesStreamand the Codex stream handler; and removed the redundantgetOpenAIResponsesCacheSessionIdalias in favor ofgetOpenAIResponsesPromptCacheKey. - Centralized OpenAI-family request-param policy into shared
resolveOpenAIOutputTokenParam(output-token field selection, OpenRouter default-cap omission,alwaysSendMaxTokensdefaulting, model/provider clamp),applyOpenAIGatewayRouting(OpenRouterprovider+ Vercel AI GatewayproviderOptions), andapplyOpenAIExtraBody(extra-body merge + Fireworks thinking drop) helpers used by both Chat Completions and ResponsesbuildParams, and moved the Chat Completions reasoning/thinking dialect dispatch (applyChatCompletionsReasoningParams+disableChatCompletionsReasoningForDialect) plus theOpenAICompletionsParamsrequest type intoopenai-sharedalongsideapplyResponsesReasoningParams. As a consistency consequence, directstreamOpenAIResponsescalls (bypassingstreamSimple) now emitmax_output_tokensforalwaysSendMaxTokens(Kimi-family) models even without a caller cap — matching Chat Completions and the valuestreamSimplealready supplied. - Centralized OpenAI-family reasoning compat resolution behind a shared
resolveOpenAICompatPolicyconsumed by both Chat Completions and Responses request builders. Shared policy now drives tool-choice reasoning suppression, dialect-specific disable encoding, reasoning-history replay filters, encrypted-reasoning inclusion, Mistral/OpenAI tool-call-id modes, stream healing/DeepSeek token stripping, and xAI/OpenRouter cache-affinity wiring instead of endpoint-local provider/model checks.
Fixed
- Fixed OpenAI Responses cost accounting to apply standard service-tier pricing multipliers (flex 0.5×, priority 2×) to the calculated cost based on the served (or requested) service tier for provider
"openai"models. - Fixed OpenAI Chat Completions to consume the dedicated
requiresReasoningContentForAllAssistantTurnscompatibility flag, preventing unnecessary reasoning replay on non-tool-call turns for OpenRouter DeepSeek and OpenCode models. - Fixed the Kimi Code and Synthetic dual-surface shim (
streamOpenAIAnthropicShim) to correctly forward caller-suppliedtoolChoice,serviceTier, anddisableReasoningoptions. - Fixed the OpenAI Responses tool-choice compatibility helper to drop
tool_choicewhensupportsToolChoiceis false, and downgrade forced choices to"auto"whensupportsForcedToolChoiceis false. - Fixed Azure Responses to avoid emitting
tool_choice: "none"whencontext.toolsis empty. - Fixed Kimi via OpenRouter forced-tool requests to omit the OpenRouter
reasoningobject instead of sendingreasoning: { enabled: false }, preserving the generic OpenRouter explicit-disable behavior while avoiding Kimi's forced-tool reasoning conflict. - Fixed Google Gemini CLI credential parsing schema to gracefully handle empty or unexpected non-string shapes without throwing unhandled exceptions
- Fixed Google Gemini CLI credential parsing to correctly prioritize
projectIdoverproject_ideven when empty, and drop non-string values gracefully - Fixed OpenRouter Responses requests to omit default max token fields unless an explicit caller cap is provided, preventing upstream filtering issues
- Fixed Chat Completions reasoning suppression (
disableReasoningOnToolChoice/disableReasoningOnForcedToolChoice) to turn thinking off symmetrically across every dialect via a shareddisableChatCompletionsReasoningForDialecthelper. Previously the conflict path only deletedreasoning_effort/reasoning(and set Z.AIthinking: { type: "disabled" }on the forced branch alone), leaving Qwenenable_thinking, Qwen chat-templatechat_template_kwargs.enable_thinking, and OpenRouter nestedreasoningenabled — so those hosts could keep thinking on under forced/required tool choice and re-trip the incompatibility the policy guards against. OpenRouter is now set to{ reasoning: { enabled: false } }(not deleted, which OpenRouter treats as default-on). - Fixed OpenRouter Responses requests to send
session_idfromsessionIdin the request body for sticky provider routing and observability grouping. - Fixed OpenRouter Responses request shaping to preserve provider routing, variant suffixes, caller header overrides, and strict-tool fallback behavior while omitting only unsafe default max-token caps.
- Fixed OpenAI Responses stateful chaining so a non-ZDR stale
previous_response_idretry keepsstore: true: the full-context retry stays chainable on the next turn and the consecutive stale-failure circuit breaker trips after the configured limit instead of alternating cold turns. Zero Data Retention rejections still disable chaining on the first strike. - Fixed Anthropic Messages tool schema normalization demoting root
anyOf/allOfand alloneOfconstraints into descriptions instead of forwarding provider-rejected keywords in MCP toolinput_schema. - Fixed Ollama Cloud GLM-5.2 reasoning efforts to map
xhighto native think"max"(#2911 by @serverinspector) - Fixed OpenRouter Responses requests tagging the streamed assistant message with a hardcoded
openai-responsesAPI instead of the runtimemodel.api, which silently disabled native-history replay (buildResponsesInput) and cross-model tool-call item-id stripping on subsequent OpenRouter turns. The message now carriesmodel.api(matching the Chat Completions path). - Fixed OpenAI-family streaming leaking a pre-retry
errorMessageonto a successful turn: the OpenRouter Anthropic compiled-grammar strict-tool fallback seterrorMessagebefore retrying with strict tools disabled and never cleared it on success, and the Chat Completions success path could carry anerrorMessagefrom an internally-retried attempt — both made a successful turn read as errored in agent state and telemetry. The Responses fallback no longer assignserrorMessage, and the Completions success path clears it before emitting the terminaldoneevent. - Fixed Codex stream-error
.coderesolution to use the same nested-first precedence (error.code→error.type→ top-levelcode) asisRetryableCodexFailureEventand the formatted message. Previously the error factory resolved top-level-first, so a failure event carrying both a top-level and a differing nested error code surfaced a.codethat could disagree with its ownretryableflag and message text.
[16.0.5] - 2026-06-17
Added
- Added
antigravityEndpointModestream option withauto,production, andsandboxvalues to control Antigravity endpoint routing - Added
seedApiKeyResolverfor reusing a pre-resolved request key while preserving resolver-driven auth retry and credential rotation - Added optional
contextSnapshotproperty toAssistantMessagewith token usage metadata via newContextSnapshotinterface (promptTokens,nonMessageTokens, and optionallastMessageTimestamp) - Added
LITELLM_BASE_URLguidance to the LiteLLM login prompt so non-default proxy endpoints are discoverable. (#2726) - Added a Gemini thinking-loop guard that watches streamed
thinkingdeltas for degenerate reasoning loops — verbatim tail repetition and near-duplicate paragraph cycling — and terminates the stream with a retryable, empty-contenterrormessage (worded as a transient stream stall) so the turn is discarded and re-sampled instead of committing a runaway transcript. Gated to Gemini models across every transport (OpenRouter, direct Google, Vertex) and disarmed once visible answer text or a tool call starts; disable withPI_NO_THINKING_LOOP_GUARD=1.
Changed
- Changed the Antigravity (
google-antigravity) request builder to mirror the capturedantigravity/hubclient: gemini-3.x sendthinkingConfig.thinkingBudgetper tier, a fixed per-modelmaxOutputTokens, a defaultfunctionCallingConfig.mode: "VALIDATED"tool mode (auto/unset tool choice only), arole: "user"system instruction, a structuredrequestId(agent/<id>/<ts>/<trajectoryId>/<step>), andlabels(model_enum,trajectory_id,last_step_index,last_execution_id,used_claude*) tracked across the conversation via provider session state.
Fixed
- Fixed Gemini usage-tier mapping so
gemini-3.5-flashis treated asFlashandgemini-3.1-proplusgemini-pro-agentare treated asProin usage accounting - Fixed Antigravity stream state handling so a request’s
last_execution_idis committed only after a successful completion and cleared between retry attempts - Fixed
streamSimple()Gemini streams to run through the thinking-loop guard for custom API and pi-native transports, so degeneratethinkingloops now abort with the same retryable empty-content error path as other Gemini stream paths - Fixed Antigravity model streaming and usage fetch paths to retry on transient
429/5xxerrors by failing over to the alternate endpoint before surfacing an error - Fixed Antigravity endpoint tracking to prefer a previously successful endpoint in
automode for subsequent requests - Fixed Antigravity and Gemini CLI model requests failing with an opaque error when Google requires account verification. Cloud Code Assist
403 VALIDATION_REQUIREDresponses now surface thevalidation_urland the signed-in account email when available, so users see an actionable account-verification message instead of the raw API error body. - Fixed MiniMax M3 in-band tool calls by adding a MiniMax dialect that parses
<minimax:tool_call>wrappers instead of falling back to generic XML. (#2759) - Fixed GitHub Copilot OAuth for Business seats by storing the login-discovered API endpoint and routing model enablement plus chat requests to that endpoint. (#2876)
[16.0.4] - 2026-06-17
Fixed
- Fixed tool argument coercion to parse double-encoded JSON strings, including quoted values like
"300", when schema expects a number - Fixed object-array coercion to parse JSON object and array strings into proper array arguments instead of wrapping raw strings
- Fixed handling of malformed JSON container strings for array schema fields so validation now surfaces a top-level
expected array, received stringerror rather than nested element errors - Fixed ChatGPT/Codex browser login missing connector OAuth scopes and rendering object-shaped token endpoint errors as
[object Object]. (#2825) - Fixed Zhipu/BigModel GLM-5.2 chat-completions requests so internal
xhigheffort serializes as provider-nativereasoning_effort: "max"and tool calls opt intotool_stream. (#2833) - Fixed Google Gemini CLI and Antigravity tool calls with
toolChoice: "auto"serializing an explicittoolConfigAUTO mode, which can cause Gemini-3 models to leak raw planning JSON instead of executing tools. (#2830)
[16.0.3] - 2026-06-16
Added
- Exported
renderDelimitedThinkingfrom the@oh-my-pi/pi-ai/dialectbarrel so consumers can reuse the dialect's<thinking>envelope unwrap-and-rewrap logic (the only./dialect/renderingprimitive re-exported; the rest stay dialect-internal).
Fixed
- Fixed OpenAI Responses/Codex tool schema normalization stripping provider-rejected regex lookaround patterns from MCP tool parameter schemas. (#2784)
- Fixed OpenAI Responses parallel tool-call routing so late keyed argument deltas for a closed call are dropped instead of being appended to another open call.
[16.0.2] - 2026-06-16
Added
- Added
UMANS_WEBSEARCH_PROVIDER=native|exasupport for routing Umans gateway-owned web search requests.
Fixed
- A single MCP tool whose input schema can't be emitted as a valid strict tool schema for the active provider no longer fails the whole turn with HTTP 400.
convertTools(openai-responses) now validates each tool's emitted parameter schema forenum/const-vs-typecontradictions that pass structural JSON-Schema validation but the provider rejects — e.g. a non-nullenumon atype: "null"node, or anenumon anarraynode — and quarantines just the offending tool with alogger.warnnaming the tool and schema path, keeping every other tool usable. AddsfindStrictToolSchemaViolationto@oh-my-pi/pi-ai/utils/schema(#2652) - Fixed OpenAI Responses-compatible streams from Ollama/local hosts dropping arguments for parallel tool calls whose deltas use
fc_<call_id>item ids, which left earlierast_grepcalls with{}and failed validation. (#2715) - Fixed dialect transcript rendering so literal thinking envelopes are unwrapped before adding the dialect's own thinking tags, preventing nested
<thinking>output in advisor raw dumps (#2700). - Fixed Anthropic-compatible Umans requests escaping client tool names and forwarding gateway web search headers so Kimi answers normally instead of returning raw gateway search results.
- Fixed Google Gemini tool calls with
toolChoice: "auto"serializing an explicittoolConfigAUTO mode, which can cause Gemini-3 models to leak raw planning JSON instead of executing tools. (#2776) - Fixed OpenAI-compatible Ollama completions that return empty
finish_reason:lengthafter fillingnum_ctxso they surface an actionable context-window error instead of an empty length stop. (#2774) - Fixed Codex browser login issuing credentials for the
opencodeOAuth originator while OMP requests identify aspi, which could make the first authenticated Codex request return 401 (#2696).
[16.0.1] - 2026-06-15
Added
- Added Umans AI Coding Plan API-key login support and
UMANS_AI_CODING_PLAN_API_KEYenvironment fallback (#2636 by @oldschoola).
Fixed
- Fixed OpenAI Responses, Azure OpenAI Responses, and Codex Responses providers ignoring async
onPayloadreplacement bodies. Provider payload hooks can now transform the actual request body sent upstream, matching the Anthropic/Gemini replacement contract. - Fixed OpenAI-compatible chat-completions streams that send object-shaped tool arguments in fragments by deep-merging nested objects and task arrays instead of replacing earlier chunks. (#2617)
- Fixed OpenAI Responses strict-mode tool schema normalization for nullable enum MCP parameters so enum constraints are distributed to matching
anyOfbranches instead of being copied onto thenullbranch. (#1835) - Fixed Cursor provider formatting tool errors with the same
[Tool Result]prefix as successful results, causing Composer models to misinterpret error messages (e.g. "Pattern must not be empty") as directives over long conversations. Errors now use a[Tool Error]prefix so the model can distinguish failures from successes in the prompt history. (#1853) - Fixed
validateToolArgumentssilently accepting JSON-encoded array strings (e.g.'["a","b"]') againstunion(string, array<string>)schemas — providers that double-serialize tool-call arguments (Z.AI / GLM) caused tools likesearchto receive the literal["a","b"]as a single path, producing zero matches (single element) or glob parse errors (multi-element). A new pre-validation pass parses JSON-array-shaped strings when the schema explicitly accepts both shapes. (#1788) - Fixed Anthropic thinking summaries that arrive wrapped in literal
<thinking>tags so advisor/raw transcript dumps do not render nested thinking tags (#2695).
[16.0.0] - 2026-06-15
Breaking Changes
- Renamed the public dialect entrypoint from
@oh-my-pi/pi-ai/grammarto@oh-my-pi/pi-ai/dialect. - Renamed grammar dialect identifiers from
ToolCallSyntaxtoDialect, renamed theGrammarinterface toDialectDefinition, and renamedGrammar.syntaxtoDialectDefinition.dialect. - Added
DialectDefinition.renderThinkingandDialectDefinition.renderTranscriptso dialect implementations serialize complete native chat transcripts, not just tool call/result blocks.
Added
- Added
renderTranscriptmethod to dialect definitions for serializing complete native chat transcripts - Added
renderThinkingmethod to dialect definitions for rendering thinking/reasoning blocks - Added support for 11 dialect implementations: Anthropic, DeepSeek, Gemini, Gemma, GLM, Harmony, Hermes, Kimi, Pi-native, Qwen3, and XML
- Added
createInbandScannerfactory function to instantiate dialect-specific scanners - Added
getDialectDefinitionfunction to retrieve dialect implementations by name - Added
renderToolCatalogandrenderInbandToolPromptfunctions for tool catalog rendering - Added
renderToolInventoryfunction to generate human-readable per-tool documentation with examples - Added
renderToolExamplesfunction to render tool usage examples in the model's native dialect - Added
encodeInbandToolHistoryfunction to encode tool call history in dialect-specific format - Added
wrapInbandToolStreamfunction to process streaming responses with in-band tool call parsing - Added
ThinkingInbandScannerfor parsing thinking/reasoning blocks across dialects - Added
OwnedStreamclass for managing dialect-aware streaming with tool call events - Added in-band thinking channels to every dialect that was missing one:
gemini(a```thinkingfence mirroring```tool_code),gemma(its native<|channel>thought…<channel|>reasoning channel),kimi(<think>…</think>), andpi(<thinking>…</thinking>). Each scanner now parses reasoning into thinking events instead of leaking chain-of-thought into the visible reply, and every dialect'srenderThinkingis a real channel that round-trips back through its scanner (no passthrough renderers).
Changed
- Moved public dialect entrypoint from
@oh-my-pi/pi-ai/grammarto@oh-my-pi/pi-ai/dialectin package exports - Updated internal imports in
stream-markup-healing.tsto use new dialect module path - Changed
renderToolInventoryto demote a tool description's own markdown headers by one level when it contains a top-level#header, so they nest under the wrapping# Tool: <name>heading instead of reading as sibling sections. Descriptions that already start at##and headers inside fenced code blocks are left untouched.
Fixed
- Fixed Gemini, Gemma, Kimi, and Pi in-band scanners to respect
parseThinking: false, leaving private reasoning markers in visible text when parsing is disabled - Fixed thinking-channel parsing for streaming Gemini, Gemma, Kimi, and Pi outputs so split or partial
<thinking>blocks no longer leak into visible replies - Fixed in-band thinking finalization and Kimi stream-healing interactions so leaked
<think>blocks are preserved when structured tool calls are present, not duplicated when explicit reasoning is present, and closed on stream flush.
Removed
- Removed
src/grammar/factory.ts(replaced bysrc/dialect/factory.ts) - Removed
src/grammar/rendering.ts(functionality moved tosrc/dialect/rendering.ts) - Removed
src/grammar/xml.ts(replaced bysrc/dialect/xml.ts)
[15.13.3] - 2026-06-15
Added
- Added the
geminiin-band tool-call syntax with Python-styletool_codeblocks anddefault_apiinvocations - Added the
gemmatoken-delimited in-band tool-call syntax using<|tool_call>and<|tool_response>blocks - Added
geminiandgemmato owned stream tool-result token detection so their tool responses are recognized - Fixed truncated Gemini and Gemma tool blocks from being emitted as plain text during streaming
- Added the Azure OpenAI provider definition (
azure) to the registry;AZURE_OPENAI_API_KEYresolves as its env-var API key via the catalog provider table.
Changed
- Gemini tool-call examples now render without the
default_api.namespace prefix, keeping<example>blocks concise. The live wire format still usesdefault_api.per the Gemini grammar.
Fixed
- Fixed duplicate tool call projections by deduplicating provider-native
toolCallevents against in-bandtool_codecalls and keeping only the first real channel - Dropped nameless native
toolCallevents so they no longer appear as surfaced tool calls in owned-mode streams - Fixed truncated Gemini and Gemma tool blocks from being emitted as plain text during streaming
- Fixed Gemini/Gemma in-band tool-call parsing around Python comments, raw/unicode string literals, and Gemma close-token text inside string values.
[15.13.2] - 2026-06-15
Added
- Added
jsonSchemaToTypeScriptto@oh-my-pi/pi-ai/utils/schemato render JSON Schema argument shapes as compact, human-readable TypeScript-style signatures - Added the generic
ToolExampletype (ToolCallExample/ToolCompareExample/ToolNoteExample, parameterized over a tool's argument shape) and anexamplesproperty on theToolinterface for defining tool-call examples once as data. - Added
renderToolExamples(via@oh-my-pi/pi-ai/grammar) to render a tool's examples into an<examples>block in the model's native tool-call syntax, with an optional_iintent-field placeholder injected when intent tracing is active. - Added per-grammar
renderToolCallrendering of a single tool-call invocation (the inner element only, without the parallel-call block envelope), distinct fromrenderAssistantToolCallswhich renders a complete block of one or more parallel calls. - Added a
GrammarRenderOptions.exampleflag torenderToolCall: when set, the invocation renders as the bare payload — Harmony emits just the JSON arguments, dropping the verbose<|start|>…<|message|>…<|call|>envelope — sorenderToolExampleskeeps<examples>blocks legible. - Added an
abortOnFabricationparameter towrapInbandToolStream(defaulttrue): whenfalse, a fabricated in-band tool-result continuation is discarded without aborting the provider request instead of cutting the turn short. - Added
@oh-my-pi/pi-ai/utils/harmony-leakexport with helpers to detect, audit, and recover GPT-5 Harmony tool-call header leaks - Added the
@oh-my-pi/pi-ai/grammarpublic entrypoint for grammar factories, prompt/call rendering, in-band scanning, history encoding, and related typed utilities - Added a unified in-band tool-call grammar engine with syntax-owned scanners, prompts, history rendering, tool-result rendering, and stream adaptation for GLM, Hermes/Qwen, Kimi, XML/Anthropic, DeepSeek, Harmony, and pi-native formats.
Changed
- Changed Harmony in-band tool-call rendering to omit the
<|constrain|>jsonmarker before the payload incommentarychannel calls - Changed tool inventory rendering to present each tool’s
Parameterssection as a simplified TypeScript-style signature derived from its wire schema - Added raw in-band tool-call block capture to parsed owned tool calls so debugging can inspect the exact model-emitted call syntax.
- Moved the canonical
ToolCallSyntaxunion to@oh-my-pi/pi-catalog/identityand re-exported it from@oh-my-pi/pi-ai/grammarso the catalog can own the syntax vocabulary without an@oh-my-pi/pi-airuntime import; all existing import paths are unchanged. - Made tool-call argument validation more lenient for schema-directed scalar coercions, including object/array stringification and 0/1 boolean coercion.
- Changed
renderToolInventory(the verbose system-prompt inventory and/dump) to render each tool as a# Tool: <name>markdown section instead of a<tool name="…">…</tool>wrapper.
Fixed
- Fixed Harmony leak handling support by adding
recoverHarmonyToolCallplus leak-detection workflows for contaminated assistant messages so recoverable tool-call arguments can be safely truncated and retried - Fixed false-positive gating in Harmony leak heuristics using signal-based checks so unrelated text containing
to=functions...is not treated as leaked tool-call markup - Routed Kimi, DeepSeek DSML, and plain thinking markup healing through the shared in-band scanners so provider leak repair and owned tool calling parse the same wire formats.
- Fixed Cursor provider (
cursor-agentAPI) streaming dropping large MCP tool-call arguments — most visibly the built-intasktool'stasksarray on multi-subagent dispatches, which failed downstream schema validation withtasks: Invalid input: expected array, received undefined. Two upstream behaviors were fighting the stream handler inpackages/ai/src/providers/cursor.ts: (1)args_text_deltacarries the cumulative args text so far peragent.proto, but the handler concatenated each snapshot onto the buffer, garbling the JSON; (2)tool_call_completedcarries anMcpArgsmap that omits oversized parameters entirely and downgrades unparsable values to their raw string fallback, but the handler unconditionally overwrote the streamed args with that map. The handler now strips the already-buffered prefix from eachargs_text_deltasnapshot (falling back to append when the snapshot doesn't extend the buffer) and merges the decodedMcpArgsmap into the streamed args — preserving streamed keys the completion frame omits and the structured value when the completion frame downgrades to a string. (#2615) - Fixed Codex Responses stream mis-routing interleaved
function_call_arguments.deltaevents when more than one tool call was open concurrently. The runtime tracked a singletoncurrentItem/currentBlock, so every delta — regardless ofitem_id— was appended to whichever item was most recently added, andoutput_item.donefor the earlier call then overwrote a sibling's stored arguments (visible astasks: Invalid input: expected array, received undefinedon thetasktool). Open items are now keyed byitem_idwithoutput_indexfallback; deltas/done events route to the matching block, late deltas whose item already closed are dropped instead of corrupting a sibling, andtoolcall_*stream events emit the rightcontentIndexper call (#2619).
[15.13.1] - 2026-06-15
Fixed
- Fixed the auth-broker (
OMP_AUTH_BROKER_URL) rejecting OAuth credentials that carry provider-specific extension fields (e.g. an MCP server'stokenUrl/clientId/clientSecret/resourceembedded for self-contained token refresh): the OAuth credential wire schema was.strict(), soPOST /v1/credentialfailed with400 unrecognized_keysand a broker-backed MCP reauth reported success while the reloaded credential lacked its refresh material and could no longer refresh. The OAuth wire schema now uses.loose()to preserve unknown fields — matching the field-preserving local SQLite store — so extra OAuth fields round-trip through broker set->get (envelope and API-key schemas stay strict).
[15.13.0] - 2026-06-14
Fixed
- Fixed OpenAI Responses/Realtime SSE stream handler crashing with "Error Code undefined: undefined" when parsing error events with nested error details by falling back to the nested error object fields.
- Fixed OpenAI-compatible providers that reject forced
tool_choiceon thinking-required models by downgrading unsupported forced choices toautowhile keeping tools available (#2546). - Fixed GitHub Copilot Anthropic transport (
api.githubcopilot.com/v1/messages) returning400 tools.0.custom.eager_input_streaming: Extra inputs are not permittedon every tool-bearing turn by stopping the emission of the per-tooleager_input_streamingflag and thefine-grained-tool-streaming-2025-05-14beta header on the Copilot transport — the proxy whitelists neither (#2558). - Disabled Bun's native ~300s pre-response
fetchtimeout in every streaming provider (OpenAI completions/responses, Azure responses, Anthropic, Codex SSE, Bedrock, Gemini CLI, Ollama). The configurable first-event/idle/SDK watchdogs (PI_STREAM_FIRST_EVENT_TIMEOUT_MS,PI_OPENAI_STREAM_IDLE_TIMEOUT_MS,compat.streamIdleTimeoutMs) were silently capped by Bun's hidden ceiling, so cold large-context streams (e.g. self-hosted vLLM at multi-hundred-K prompts) died at exactly 300s withTimeoutError: The operation timed out.Direct callers of./providers/{amazon-bedrock,google-gemini-cli,ollama,openai-codex-responses}(which bypassregister-builtins' iterator-level watchdog) now install a pre-responseAbortSignal.timeout(firstEventTimeoutMs)alongside the disable, so a stalled upstream still fails within the configured budget instead of hanging forever (#2422) - Fixed Gemini / Antigravity streams (Google Cloud Code Assist API) creating a trailing empty text block and emitting redundant
text_start/text_delta/text_endevents at the end of the turn when the final SSE chunk contains an empty text part (text: ""). The parser now ignores empty text parts, preserving the active transcript block state and ensuring proper nesting and rendering of subsequent background jobs or new turns. - Preserved terminal Google
thoughtSignatures by still extracting and applying the signature on the active block even when the text part is empty or undefined. - Stopped Gemini Antigravity sessions (
gemini-3*/ Claude under Cloud Code Assist) from leaking system rule reminders and personality preambles into the final response, by appending an explicit 'do not output rule checks' instruction to the injected system parts. - Fixed Gemini / Antigravity streams (Google Cloud Code Assist API) letting a
functionCallpart's ownthoughtSignatureclobber the preceding text or thinking block's signature onthink → toolandtext → toolturns. A signed function-call part hastext: undefined, so it fell into the terminal-signature branch while the prior block was still active; that branch now skips function-call parts, leaving the tool call's signature on the tool call where it belongs and preventing corrupted signatures on same-model replay. - Fixed MiniMax-M3 OpenAI-compatible streams rendering reasoning twice when the same chunk carried both
<think>…</think>content and structuredreasoning_content; structured reasoning now wins and cumulative MiniMax reasoning snapshots are collapsed to deltas using a per-signature snapshot tracker that survives the</think>-to-text block transition (so post-answer cumulative snapshots don't reinstate a duplicate thinking block). (#2433)
[15.12.6] - 2026-06-14
Changed
- Bumped Z.AI (GLM Coding Plan) API key validation probe to glm-5.2.
Fixed
- Fixed tool schema conversion for non-Cloud Code Assist Google Gemini models by normalizing parameters with
normalizeSchemaForGoogleto prevent un-normalized schema properties (such asadditionalProperties: falseor type arrays) from causing Gemini API errors. - Fixed OpenAI-family request builders dropping forced named
tool_choicedirectives when the named tool is absent from the serializedtoolsarray, preventing spec-strict providers from rejecting self-inconsistent requests. (#1701)
[15.12.4] - 2026-06-13
Added
- Added
GITLAB_CLIENT_IDandGITLAB_REDIRECT_URIenv-var overrides for the GitLab Duo OAuth login flow so users running with their own GitLab OAuth application can replace the bundled credentials when GitLab rejects the bundledclient_id's redirect URI. SettingGITLAB_REDIRECT_URIalso disables the random-port fallback (strict OAuth providers reject mismatched URIs anyway). (#2424) - Added
AuthStorage.listStoredCredentials()andAuthStorage.removeCredential()for per-account credential management.
Changed
- Replaced the OpenAI SDK client usage in
openai-completions,openai-responses,azure-openai-responses, andopenai-codex-responseswith the new internalpostOpenAIStreamOpenAI-wire JSON/SSE transport
Fixed
- Fixed streaming providers to cancel upstream model requests when the client closes the response body, so interrupted SSE sessions stop instead of continuing in the background
- Fixed: provider request builders treat unknown
model.maxTokens(null) as "no model cap" instead of coercing to0viaMath.min; Anthropic falls back to the 64k Claude-Code cap for its requiredmax_tokens. - Fixed transient stream failures on OpenAI-compatible providers by retrying HTTP 408/429/5xx responses and transient network errors with Retry-After/quota-hint aware backoff
- Fixed SSE stream handling for OpenAI-compatible responses by parsing wire-level JSON frames directly and honoring
[DONE]termination - Fixed stream error handling for OpenAI-compatible providers by preserving structured HTTP status/headers and response body details from failed requests for retry and strict-tool fallback logic
- Fixed OpenAI-compat streams ending with a bare
finish_reason: "error"(gateways like OpenRouter reporting upstream failures, e.g. GeminiMALFORMED_FUNCTION_CALL) surfacing as a non-retryableProvider finish_reason: error. The reason is now mapped toProvider returned error finish_reason, which the session retry classifier recognizes as transient, so the turn auto-retries instead of stopping with a pinned error banner. - Fixed
SqliteAuthCredentialStore.open()crashing withSQLITE_BUSY_RECOVERY(errno 261) when severalomp --sessionpanes restore concurrently after an unclean shutdown:PRAGMA busy_timeout = 5000now runs as a standalone statement BEFOREPRAGMA journal_mode=WAL(the first lock-taking statement during WAL recovery), andopen()retries the BUSY family —SQLITE_BUSY,SQLITE_BUSY_RECOVERY,SQLITE_BUSY_SNAPSHOT,SQLITE_BUSY_TIMEOUT— with bounded exponential backoff. The exhausted-retry error message includes the DB path. ExportedisSqliteBusyError(err)for callers that need the same classifier (#2421). - Fixed MiniMax-M3 OpenAI-compatible streams rendering reasoning twice when the same chunk carried both
<think>…</think>content and structuredreasoning_content; structured reasoning now wins and cumulative MiniMax reasoning snapshots are collapsed to deltas. (#2433) - Fixed Gemini turns silently halting the agent when the model returned
finishReason: STOPwith only an empty (or whitespace-only) text part and no tool call — the well-known "empty response" failure. All Google surfaces (public Generative LanguagestreamGoogle, VertexstreamGoogleVertex, and Cloud Code Assistgoogle-gemini-cli/google-antigravity) now classify such a turn as empty via the sharedhasMeaningfulGoogleContentcheck and retry it up toMAX_EMPTY_STREAM_RETRIEStimes before surfacing an error. The Cloud Code Assist path previously had an empty-stream retry that never fired for this case (itshasContentflag counted an empty-string text part as content), and the public/Vertex path had no retry at all; the retry now emits a singlestartevent so no duplicate partial message leaks downstream.
[15.12.1] - 2026-06-12
Added
- Added the optional
ToolResultMessage.uselessflag: tools can declare a finished result contextually useless (zero matches, elapsed wait) so compaction passes may elide it once consumed. Never serialized to provider wire formats and never set together withisError.
[15.12.0] - 2026-06-12
Fixed
- Fixed Anthropic requests bypassing lone-surrogate sanitization after payload hooks or Anthropic-origin tool-call replay: the model itself can emit unpaired surrogate escapes in its own tool-argument JSON (streamed out fine, then rejected with
400 The request body is not valid JSONon every subsequent request, bricking the session). The final Anthropic payload is now deep-sanitized withtoWellFormed()immediately before SDK serialization; the pass is identity-preserving, so well-formed arguments stay byte-identical and prompt-cache prefixes are unaffected.
[15.11.8] - 2026-06-12
Breaking Changes
- Removed the Codex SSE stateful transport path, so SSE turns no longer send
previous_response_idwith delta input and now always send the full transcript
Changed
- Scoped
x-codex-turn-statehandling to within-turn continuations so only tool-loop follow-ups include the turn-state header and new user turns start without it
Removed
- Removed the
statefulResponsesoption fromOpenAICodexResponsesOptions, and SSE stateful mode is no longer controlled by thePI_CODEX_STATEFUL-style flag
Fixed
- Fixed the platform OpenAI Responses and Codex websocket stale-chain classifiers missing the "Unsupported parameter: previous_response_id" rejection phrasing (FastAPI-style
detailbody with noerror.code), so a chained turn now falls back to a full-transcript replay instead of surfacing the 400 - Fixed the HTTP-400 raw-request dump for Codex SSE to record the body actually sent on the wire instead of the pre-transport request body, which made chained-request failures look like the rejected parameter was never sent
[15.11.7] - 2026-06-12
Added
- Added
requestModelIdandthinking.suppressoptions togoogle-gemini-cliso collapsed effort-tier variants serialize their per-effort upstream wire id, and thinking-off requests on models withthinking.suppressWhenOffsend an explicitthinkingConfig(includeThoughts: falsewiththinkingLevel: "MINIMAL"orthinkingBudget: 0) — Cloud Code Assist re-applies the per-id baked server default when the config is omitted, silently thinking and billing the tokens - Added mandatory-reasoning clamping: models baked with
thinking.requiresEffortfloor omitted or disabled reasoning to the lowest supported effort in every api mapping, anddisableReasoningno longer emits OpenRouterreasoning: { enabled: false }for them — fixesomp benchand utility requests 400ing with "Reasoning is mandatory for this endpoint and cannot be disabled" on OpenRouter Gemini 3.x
Changed
- Changed
google-gemini-clirequest mapping to route per-request wire ids viaresolveWireModelId: the session effort picks the backing variant id (collapsedgemini-3.5-flashat high →gemini-3.5-flash-low; claude pairs route off → bare id, efforts →-thinking) whileAssistantMessage.modeland usage attribution stay on the logical id. A thinking budget clamped to zero now falls through to the thinking-off path (off routing plus suppression) instead of only disabling thinking - Changed
openai-completionsandanthropic-messagesto serialize per-request wire ids viaresolveWireModelId, so collapsedX/X-thinkingpairs on aggregators and custom providers switch to the thinking SKU when reasoning is enabled (previously onlygoogle-gemini-clirouted effort-tier variants)
Fixed
- Fixed
google-gemini-cliignoringModel.requestModelIdwhen serializing the request model id
[15.11.5] - 2026-06-12
Added
- Added
AuthStorage.listUsageHistoryto retrieve historical usage snapshots with optionalproviderandsinceMsfiltering - Added durable usage-history persistence in the sqlite auth store so successful usage reports are recorded as time-series snapshots of limit utilization for later trend inspection
- Added
AuthStorage.redeemResetCreditto redeem stored OpenAI Codex saved rate-limit reset credits for a target account bycredentialId,accountId, oremail - Added
listCodexResetCreditsandconsumeCodexResetCreditexports for OpenAI Codex saved reset-credit listing and redemption - Added
resetCreditswithavailableCounttoUsageReportso OpenAI Codex usage data now exposes redeemable rate-limit resets - Added
openai-codex-resetexports via package barrel for out-of-band tooling usage - Added a one-shot request-debug target that writes the next provider HTTP request JSON to an explicit path.
Changed
- Changed
AuthStorage.redeemResetCreditto invalidate cached usage data after a successful redemption so the next usage report reflects the reset immediately
Fixed
- Fixed temporary credential block state so redeemed reset credits immediately make the affected account selectable again after
redeemResetCreditsucceeds - Fixed one-shot request-debug path handling so an explicit request log target is consumed after the next request and no longer affects subsequent calls
- Fixed explicit request-debug path mode to create missing parent directories before writing request logs
- Fixed explicit request-debug mode to overwrite existing
.res.logfiles for the requested path instead of failing when they already exist - Fixed OpenAI Responses
previous_response_idchaining on Zero Data Retention orgs: the in-provider retry classifier missed the ZDR-specific 400 ("Previous response cannot be used for this organization due to Zero Data Retention"), so chained turns kept failing every other request after a brief recovery — the chain was reset but not disabled, so the next successful full-replay turn re-armed it. The ZDR phrasing is now classified categorically: one strike disables chaining for the session (skipping the three-strike circuit breaker) and the in-call retry dropsstore: true/previous_response_idand replays the full transcript instead (#2341).
[15.11.4] - 2026-06-12
Added
- Codex/Responses providers now map
end_turn: falseon the terminal stream event (Codex backend signal for "response ended, turn didn't" — commentary-only progress updates) tostopDetails: { type: "pause_turn" }with stopReason"stop", so the agent loop can re-sample instead of ending the turn. Wired inopenai-codex-responsesandprocessResponsesStream(openai-responses/azure-openai-responses); inert for backends that never send the field. - Added Codex upstream protocol features to
openai-codex-responses(tracking codex-rs as of June 2026):onModerationMetadatacallback surfacingresponse.metadata→openai_chatgpt_moderation_metadataon both transports;reasoningContextoption emittingreasoning.context(auto/current_turn/all_turns);clientMetadataoption emittingclient_metadatain the request body (canonicalx-codex-turn-metadataenvelope) without breaking the websocket append fast-path; and an opt-inresponsesLitemode mirroring codex-rs — lite header on HTTP requests and the websocket upgrade,ws_request_header_*marker inresponse.createclient metadata, lite-keyed socket pooling, image-detail stripping, forced serial tool calls, andreasoning.context: all_turnsdefault. Dormant until OpenAI flipsuse_responses_litein the model catalog. - Added
withOAuthAccess— thewithAuthcounterpart for OAuth-access consumers: runs an operation through the central a/b/c auth-retry policy (resolve → force-refresh same account → rotate to a sibling) while handing the attempt the fullOAuthAccess(bearer plusaccountId/projectId/enterpriseUrlidentity metadata). Use it instead of hand-rolledgetOAuthAccess+ fetch flows so 401s and usage-limits rotate credentials instead of failing the call. - Added
ProviderHttpError— a typed HTTP error carryingstatus,headers, andcode— replacing the ad-hocas Error & { status?... }/Object.assignhacks at provider throw sites, with per-provider subclassesCodexApiError,AuthGatewayError,GoogleApiError,GeminiCliApiError,OllamaApiError, andBedrockApiError;AnthropicApiErrornow extends it. Google, Gemini CLI, Ollama, and Bedrock HTTP errors now also carry response headers, so server-suggestedretry-afterdelays are visible to retry classification on those paths. The internalwithHttpStatushelper was removed. - Added stateful SSE turn chaining for OpenAI Codex (on by default; disable with
PI_CODEX_STATEFUL=0orstatefulResponses: false): SSE requests now reuseprevious_response_idwith delta-only input instead of replaying the full transcript, mirroring the websocket fast-path via a shared transport-aware builder. Any history mutation or option change falls back to a full replay; a server-sideprevious_response_not_found(HTTP or in-stream) resets the chain and retries the turn with full context, and three consecutive stale failures disable chaining for the session. - Added stateful
previous_response_idchaining to the platform OpenAI Responses provider (openai-responses): on by default against the official api.openai.com endpoint (forcesstore: true, which chaining requires), off for other Responses endpoints; override withstatefulResponsesorPI_OPENAI_STATEFUL. Chain detection compares the wire form of the conversation arguments alone — per-turn trailing scaffolding such as the GPT-5 "Juice: 0" developer item is excluded from the append-baseline prefix check and re-appended to the delta — and a rejected/stale previous response falls back to a one-shot full replay with the same circuit breaker. - Added
AuthStorage.getOAuthAccountIdentity()and theOAuthAccountIdentitytype — a read-only lookup returning theaccountId/email/projectIdof the OAuth credential a session is currently routed to, for display and metadata paths.
Changed
- The GPT-5 "Juice: 0" no-reasoning developer item in
applyResponsesReasoningParamsis now gated on the resolvedcompat.requiresJuiceZeroHackflag (auto-detected from GPT-5-family model names by@oh-my-pi/pi-catalog, overridable per model) instead of an inline model-name check.
Fixed
- Fixed websocket append fast-path to remain usable when only
client_metadatachanges between turns - Fixed
onModerationMetadatahandling so exceptions thrown by callback observers no longer terminate the response stream - Fixed local SQLite OAuth credential caches returning a stale Anthropic access token after another
ompprocess refreshed and persisted the same row.AuthStoragenow syncs the selected row from storage before returning or force-refreshing OAuth credentials, so concurrent sessions pick up peer-rotated tokens instead of surfacing a one-turn401 Invalid authentication credentials. - Fixed forced OAuth preflight refresh failures being swallowed silently in credential selection; they now emit a debug log (
OAuth preflight refresh failed) so stale-refresh-token replays from concurrent sessions are diagnosable.
[15.11.3] - 2026-06-11
Fixed
- Fixed GitHub Copilot long-context model requests to use the upstream
requestModelIdwhen calling Anthropic, OpenAI Responses, and OpenAI Completions APIs - Fixed GitHub Copilot model enablement to deduplicate catalog variants by upstream model ID when enabling all models
[15.11.2] - 2026-06-11
Fixed
- Fixed Anthropic encoding of error tool results with whitespace-only content so requests no longer 400 with
tool_result: content cannot be empty if is_error is true
[15.11.1] - 2026-06-11
Changed
- Exported
resolveAnthropicMetadataUserIdso non-streaming Anthropic Messages consumers (e.g. the coding-agent web search provider) can produce the same Claude-Code-shapedmetadata.user_idas the main streaming path.
Fixed
- Preserved Anthropic
stop_detailson assistant messages so refusal and sensitive classifier stops remain structurally visible to callers. (#2290) - Fixed OpenAI Responses, Azure OpenAI Responses, and OpenAI Completions streams hanging until the 120s idle watchdog errored the turn when a provider delivers the terminal frame but never sends
[DONE]nor closes the connection.processResponsesStreamnow breaks out of the event loop onresponse.completed/response.incomplete(mirroring the Codex websocket/SSE terminal break), and the completions consumer breaks oncefinish_reasonplus a usage payload arrived — or, for hosts that never send usage, ends the stream cleanly via a short post-finish grace window (iterateWithTerminalGrace) that aborts the transport to release the socket.
[15.11.0] - 2026-06-10
Added
- Added optional
ImageContent.detail("auto" | "low" | "high" | "original"): an OpenAI resolution hint forwarded by theopenai-responsesserializers (default staysauto) and byopenai-completionsfor the values Chat Completions supports."original"preserves native resolution — required for snapcompact frames, whose pixel-font glyphs do not survive the default downscale. Providers without a detail knob ignore the field.
Fixed
- Fixed OpenRouter DeepSeek V4 strict tool schemas nesting
anyOfinside the nullable wrapper for optional unions, which produced a branch withouttypeand triggered OpenRouter'sInvalid tool parameters schema : field anyOf: missing field type400. (#2270) - Hardened strict tool-schema handling beyond the optional-union case:
enforceStrictSchemanow splices natively nested pure unions into the parentanyOf(only when the inner node carries no constraining siblings, since sibling keywords are conjunctive withanyOf), so source schemas with nested unions no longer produce type-lessanyOfbranches that strict upstream validators reject. (#2270) - Made the openai-completions non-strict retry reachable for
"mixed"strict mode (previously gated toall_strict, i.e. Cerebras only) and taught it to recognize upstream tool-schema validation 400s (Invalid tool parameters schema …,Invalid schema for function …). A matching rejection now retries the request with base (non-strict) schemas and persistsstrictToolsDisabledon the provider session, so later requests skip the doomed strict attempt instead of paying a 400 + retry round-trip each turn. (#2270) - Cross-model
anthropic-messages → anthropic-messagescontinuations now preserve prior assistant turns' reasoning chains end-to-end: every priorthinking/redactedThinkingblock survives (not just the latest surviving assistant), and third-party ↔ third-party replays keep their signatures intact so the reasoning chain stays signed for the next turn. Signatures are stripped (and anyredacted_thinkingsibling without a native landing spot is dropped) only when an official Anthropic endpoint is on either end of the replay — official Anthropic cryptographically binds reasoning signatures to its key+session+model, while compatible reasoning endpoints (Z.AI, DeepSeek, custom anthropic-messages providers configured viamodels.yaml) treat them as opaque continuation hints. Source-side official detection uses the canonical catalog provider id"anthropic"(assistant messages carry nobaseUrl); target-side detection reuses the bakedcompat.officialEndpointflag. Latest-turn byte-for-byte behavior (Anthropic's "thinking blocks in the latest assistant message cannot be modified" rule) and existing aborted/errored last-block sanitization are unchanged. (#2257, #2265)
[15.10.12] - 2026-06-10
Added
- Added
antigravityRankingStrategyand registered it forgoogle-antigravityinDEFAULT_RANKING_STRATEGIES, so new sessions are routed to OAuth credentials with quota headroom for the requested model backend (lowest relevantremainingFractioncounter as the sole ranked window, 24hwindowDefaultsmatchingdaily-cloudcode-pa.googleapis.comresets). Without it, the existingantigravityUsageProviderdata never reached credential selection. (#2198)
Changed
- Updated MiniMax and MiniMax Token Plan defaults to
MiniMax-M3and refreshed Token Plan login copy/links (#1725).
Fixed
- Fixed OpenAI Responses and Azure OpenAI Responses streams silently surfacing incomplete output as successful when a custom/proxy provider drops the connection without sending a terminal
response.completed/response.incompleteevent. Both providers now detect premature stream closure and throw withstopReason: "error"(#2184) - Fixed
isUsageLimitErrormissing Antigravity / Cloud Code Assist'sIndividual quota reached429 phrasing. TheUSAGE_LIMIT_PATTERNonly knewquota.?exceeded/limit_reached, soauth-retryandAuthStorage.markUsageLimitReachedtreated the response as a terminal provider error and pinned sessions to the exhausted OAuth account instead of rotating to a sibling credential. The pattern now also matchesquota.?reached. (#2198) - Scoped Antigravity usage blocking and ranking by model family (
gemini-*/gemma-*→ Google,claude-*→ Anthropic,gpt-*/openai/*→ OpenAI), so an exhausted Gemini counter no longer makes a healthy Claude/OpenAI Antigravity credential unavailable until reset. (#2198) - Fixed no-model Antigravity credential lookups (e.g. image-provider discovery) inheriting provider-wide exhaustion:
scopeLimitsnow returns no limits without a concrete backend counter, andblockScopealways returns a counter scope so missing model context can never fall through to AuthStorage's provider-wide block bucket. (#2198)
[15.10.11] - 2026-06-10
Breaking Changes
- The model catalog moved to the new
@oh-my-pi/pi-catalogpackage. Deep subpath exports@oh-my-pi/pi-ai/models.json,/models,/model-cache,/model-manager,/model-thinking,/effort,/provider-models*,/utils/discovery*,/providers/openai-codex/constants,/providers/google-gemini-headers, and/providers/openai-completions-compatare gone — import the@oh-my-pi/pi-catalogequivalents (/models.json,/models,/model-cache,/model-manager,/model-thinking,/effort,/provider-models*,/discovery*,/wire/codex,/wire/gemini-headers,/compat/openai). The pi-ai root barrel re-exports only the model/effort types its own signatures use (Model,Api,ThinkingConfig,Effort,Usage, compat interfaces) — catalog values (getBundledModel(s),calculateCost,modelsAreEqual,clampThinkingLevelForModel,DEFAULT_MODEL_PER_PROVIDER, …) must be imported from@oh-my-pi/pi-catalog. ProviderDefinitionis now auth-only:defaultModel,createModelManagerOptions,catalogDiscovery,dynamicModelsAuthoritative,allowUnauthenticated, andspecialModelManagermoved to pi-catalog'sCATALOG_PROVIDERStable, andKnownProviderIdwas replaced by pi-catalog'sKnownProvider(registry completeness is enforced by a compile-time check against that union). The pure GitHub Copilot key/endpoint helpers moved fromregistry/oauth/github-copilotto@oh-my-pi/pi-catalog/wire/github-copilot.
Added
- Exported
wrapFetchForCchso non-streaming OAuth callers (e.g. the web-search provider) can patch the Claude Code billing-headercchattestation into their request bodies instead of shipping thecch=00000placeholder.
Changed
- Reduced idle-watchdog churn on the token hot path: the abort promise/listener is created once per stream instead of per yielded item, the deadline uses a persistent re-armed timer instead of a
setTimeoutcreate/destroy pair per delta, and the persistent race promises are re-minted every 1024 items so per-race reaction records cannot accumulate for the stream's whole life. - Memoized Anthropic many-image downscaling by content-block identity, so long sessions with stable message objects no longer re-decode and re-encode every oversized image on each request and retry.
- Tool-argument validation errors now truncate embedded argument strings at 256 chars per field — a failed
write-class call no longer echoes hundreds of KB of payload back to the model as the error message. - Auth storage no longer issues per-boot no-op writes: the schema-version row is only rewritten when the recorded version actually changes, and the credential identity-key backfill skips rows whose derived identity is null — reopening a current-schema database now performs zero write transactions
- Plain provider env-var names moved to the catalog table: registry defs dropped their 48
envKeysliterals (including the pure$pickenvpickers forhuggingface/qwen-portal/xai-oauth),getEnvApiKeynow derives those fallbacks fromCATALOG_PROVIDERS[].envVars, andenvKeysremains only for computed resolvers (Anthropic Foundry, Vertex ADC, Bedrock credential chains) and non-catalog providers (kagi,tavily,parallel,perplexity) - Protocol handlers are now pure
model.compatreaders — the per-requestresolve*Compat/detect*Compatcalls (anthropic ×11, responses ×3, completions wrappers), inlinestrictResponsesPairinghost detection, the OpenCodereasoning_contentmutation block, and allresolvedBaseUrlthreading are gone. Compat is materialized once at model build time (@oh-my-pi/pi-catalogbuildModel); the OpenCode thinking-mode quirk is a precomputedcompat.whenThinkingpointer swap, and request-time base-URL overrides only feed the HTTP client. Behavior is unchanged (the AnthropicsupportsLongCacheRetentionofficial-endpoint gate is folded into detection). - Providers now read baked thinking/wire metadata instead of re-parsing model ids per request: the Anthropic handler gates sampling params on
model.compat.supportsSamplingParamsand adaptivedisplayonmodel.thinking.supportsDisplay(Bedrock too), adaptive effort tiers come from the bakedthinking.effortMap, the GooglethinkingLevelmap is static, and effort-dial-less reasoners (thinking: undefined, e.g.xai-oauth/grok-build) short-circuitresolveOpenAiReasoningEffortwithout the removedmodelOmitsReasoningEffortpredicate. - Anthropic streaming retries now use a 10-retry budget with the Anthropic-compatible 0.5s exponential backoff capped at 8s with jitter; server
retry-afterhints still win, and retryable pre-content failures such as 502s no longer stop after three tries.
Fixed
- Fixed Ollama chat requests honoring
omitMaxOutputTokens, sendingthink: falsewhen reasoning is explicitly disabled, and preserving HTTP 400 response bodies in surfaced errors. - Fixed
AuthStorage.markUsageLimitReachedcollapsing "every sibling is momentarily blocked" into "no sibling exists": it now returnsUsageLimitMarkResultwith the earliest sibling block expiry (retryAtMs), so retry layers can wait out a short-lived block (60s post-401, 5-min usage-probe) instead of adopting the provider's multi-hour retry-after.rotateSessionCredentialand the auth-gateway adapt to the new shape. - Fixed Gemini streaming silently presenting truncated or blocked output as a successful
stop: in-band{"error":{...}}events andpromptFeedback.blockReasonchunks were never inspected, and a stream ending without anyfinishReasonkept the initializedstop— all three now surface as errors (both the API-key and gemini-cli/Antigravity consumers), and thetoolUsestop-reason override no longer masksSAFETY/MALFORMED_FUNCTION_CALLfinishes that arrive after a valid tool call. - Fixed Gemini/Bedrock error finishes reporting "An unknown error occurred": the raw finish/stop reason (
MALFORMED_FUNCTION_CALL,RECITATION,guardrail_intervened, …) is now recorded into the surfaced error message. - Fixed the Anthropic provider retry loop ignoring server
retry-afteron 429/529 — it now waitsmax(headerDelay, backoff)instead of hammering a rate-limited endpoint three times within ~14s of guaranteed failures. - Fixed in-stream Anthropic SSE
errorevents being thrown as raw JSON envelopes; the structurederror.type/messageis parsed out, keeping retry classification on the typed token instead of accidental regex hits. - Fixed transparent-reconnect tolerance duplicating content behind replaying proxies: after a duplicate
message_start, replayedcontent_block_startevents for already-closed indexes are now consumed silently instead of appending duplicate text/tool calls. - Fixed the Anthropic gateway accepting malformed known-type content blocks (e.g.
{type:"text", text:123}) through the unknown-block catch-all, corrupting history and surfacing later as an opaque TypeError — they now fail validation with a clean 400. The gateway's encode stream also emitspingkeepalives every 15s and a completemessage_start/message_delta/message_stopenvelope when the inner stream ends without a terminal event, so strict clients no longer classify slow or empty streams as protocol errors. - Fixed dotted-version Claude ids (
claude-opus-4.7/4.8on GitHub Copilot, Vercel AI Gateway, Zenmux) missing adaptive thinkingdisplaysupport — streamed reasoning stayed hidden on those entries because the display predicate only matched dash-form ids (same failure class as #1373). - Fixed the Mistral
requiresThinkingAsTextreplay path calling.unshift()on string assistant content — an unconditional TypeError that failed any same-model history turn carrying both thinking and text. - Fixed the Responses gateway stripping
encrypted_contentfrom inbound reasoning items (strip-mode schema), which broke codex-style stateless replay; the schema is now loose, restoring the symmetry the outbound encoder already preserved. Composite internalcallId|itemIdids are also split before hitting the wire so third-party clients that validatecall_idcharsets no longer reject them. - Ported the shared unfinished-tool-call sweep to the codex
response.completedhandler, so a lostoutput_item.donecan no longer persist a tool call with stale{}arguments and transient parser fields into session history. - Fixed live text freezing until item completion when a lossy proxy drops
content_part.added: the missing part is now synthesized on the firstoutput_text/refusaldelta (shared and codex decoders). - Fixed interleaved
content/tool_callsdeltas fragmenting a tool call into a truncated call plus a nameless phantom: text/thinking transitions no longer finish open tool-call blocks, so index-only continuation deltas re-find them. - Fixed the Azure chat-completions path ignoring
AZURE_OPENAI_DEPLOYMENT_NAME_MAP(only the Responses provider honored it), producing opaque 404s when deployment names differ from catalog model ids. - Fixed the chat gateway discarding inbound assistant
reasoning_content, which fed DeepSeek/Kimi exact-replay upstreams a placeholder instead of the model's actual reasoning; it now round-trips as a thinking block, andtoolcall_endemits a corrective id/name chunk when the streamed start carried empty values. - Fixed the auth retry loop minting OAuth tokens and firing a doomed request after the caller aborted, and stopped masking resolver failures (broker/network/refresh errors) as "No API key" — the actual cause is preserved.
- Fixed
EventStream.end()without a terminal result leaving.result()pending forever (reachable via extension streams and the lazy wrapper); it now rejects with a synthesized error. - Fixed the Copilot retry wrapper blind-retrying every retryable error with fixed 400ms delays: 429/5xx now honor
Retry-After(capped at 30s) and other statuses are not retried, while status-less transport blips keep the linear retry. - Fixed the OpenAI completions error path ending the stream without closing open text/thinking/tool-call blocks, leaving consumers with orphaned block lifecycles on every stream error or idle-timeout abort.
- Fixed DSML hold-back freezing display on any bare
<in model output for up to 256 chars: idle-state holding now only triggers on a strict DSML section-open prefix, and blowing the 1MB parameter cap no longer leaks the closing envelope tags as visible text; a capped parameter value also carries an explicit…[parameter truncated]marker instead of executing the tool with silently corrupted input. - Fixed schema normalization blanking DAG-shared subtrees to
{}: the visited-set cycle guard treated a subschema object reused across two properties as a cycle; path-trackingenter/exitnow allows sharing while still short-circuiting true cycles, frozen input schemas no longer throw, and the path counter no longer leaks depth on the cycle branch (which made every later normalization of the same object misreport a cycle). - Fixed shared in-flight Google token refreshes being bound to the first caller's
AbortSignal, failing every concurrent waiter when one parallel Vertex call was cancelled; callers now race their own signal against a detached refresh, which is bounded by its own 30s timeout so a hung fetch cannot pin the in-flight slot until process restart. - Fixed Gemini <3 multimodal tool results breaking the single-function-response-turn invariant for parallel tool calls (image turns are buffered and flushed after the merged functionResponse turn), and the gemini-cli consumer now defaults missing
functionCall.argsto{}like the shared consumer. - Fixed Bedrock dropping
toolConfigentirely whentoolChoiceis"none"while history still contains tool blocks — the Converse API rejects such requests, so tool specs are kept and only the choice is omitted. - Fixed AWS credential handling serving expired credentials until process restart: cache entries are invalidated on 401/403, file-sourced session-token credentials get a 5-minute TTL, and concurrent first requests single-flight instead of spawning duplicate
credential_process/SSO fetches — the shared resolution is detached from the first caller's abort signal (one cancelled request no longer fails every waiter) and bounded by its own 30s timeout. The eventstream reader also cancels the response body on abnormal exit instead of leaving the HTTP connection draining. - Fixed an unbounded, zero-backoff Codex WebSocket reconnect loop on
websocket_connection_limit_reached: the no-content reconnect path never consulted the retry budget and never waited, hammering the endpoint forever when the limit is account-scoped. Reconnects are now budgeted and delayed like every other WS retry path, falling back to a single SSE replay when exhausted. - Fixed the Codex whitespace-loop breaker not observing degenerate frames that arrive after their item closed (or before it opened) — those frames count as stream progress, so the idle watchdogs never fired and the turn hung forever, which is exactly the failure mode the breaker exists for. Whitespace-loop recovery now also refuses to replay the turn once a
toolcall_endwas delivered, surfacing the error instead of re-emitting the same tool calls. - Fixed the two remaining Codex retry paths (WS mid-stream reconnect and the empty-content SSE fallback) leaking blockless native output items (e.g.
web_search_call) from the failed attempt into the replayed turn'sproviderPayloadand append baseline. - Fixed Codex WebSocket failure handling closing whatever connection currently occupies the session slot — including a concurrent caller's in-flight CONNECTING handshake, whose rejection (
websocket closed before open) is classified fatal and disabled WebSockets for the whole session. Failure cleanup now skips CONNECTING sockets and the pool re-joins replacement handshakes (bounded). - Fixed the Codex request transformer not repairing orphan
custom_tool_call_outputitems (onlyfunction_call_outputwas folded into an assistant note) — a compaction splice that dropped anapply_patchcall while keeping its result produced a hard 400 on the default GPT-5 Codex toolset. - Fixed
processResponsesStreamfinalizing reasoning items via a bareitemIdcontent scan instead of the routed entry: with id-less reasoning items (local hosts), everyoutput_item.donematched the FIRST thinking block — the second item's text clobbered it and the second block was never finalized or signed. - Fixed
processResponsesStreamdropping tool calls and message text whoseoutput_item.addedevent was lost (lossy proxies):toolcall_endwas emitted with a dangling contentIndex while the call never enteredmessage.content, so the agent loop silently never executed it. The done handler now synthesizes the missing block; still-open tool-call blocks are also final-parsed atresponse.completedso thetoolUseoverride cannot hand the agent stale{}arguments. - Fixed
response.incompletewithincomplete_details.reason: "content_filter"being reported as a token-cap truncation (stopReason: "length") — the agent loop's length recovery then asked the model to "shorten" a filtered prompt. Content-filtered turns now surface as errors; usage is also populated fromresponse.failedevents, and an unknown terminal status degrades to"stop"with a logged anomaly instead of throwing away a fully-streamed response. - Fixed Copilot
premiumRequestsaccounting being dropped from failed/cancelled responses:populateResponsesUsageFromResponsereplacedusagewholesale and the error path threw before the success-path re-apply. The populate now preserves the field. - Fixed
deduplicateToolCallIdssuffixing the whole composite Responses id (callId|itemId) —normalizeResponsesToolCallIdextracts the first segment as the wirecall_idat encode time, so both copies collapsed back onto onecall_idand the request carried duplicate call/output pairs. The suffix and length budget now apply per segment. - Gated native history payload replay on api + model id in both Responses providers: after a mid-session model switch, reasoning items carrying encrypted content minted by the previous model were replayed verbatim under the new model. Replay now falls back to block re-encode (which already strips foreign signatures), matching
transformMessages' same-model trust rule. - Fixed Azure OpenAI Responses requests omitting
store: falsewhile requestingreasoning.encrypted_content(stateless-only per OpenAI), replaying custom tool calls paired with mismatchedfunction_call_outputitems (customCallIds was never threaded through), letting the SDK's internal retries (maxRetries 5) silently re-POST inside the explicit first-event deadline, and sending aprompt_cache_keywhen the caller opted out viacacheRetention: "none". - Fixed strict-pairing Responses backends (Azure, Copilot) silently discarding tool results whose call is absent from history — the result is now folded into an assistant note (same shape as orphan-output repair) so the model keeps the information.
- Fixed the OpenAI Responses first-event watchdog staying armed across the
onResponsenotification callback (a slow callback aborted an already-connected stream), Copilot transient-model retries re-attempting on an already-aborted signal (instant dead retry surfacing the scheduler's AbortError), CodexreasoningSummary: nullbeing coerced to"auto"(the documented omit-summary contract was unreachable), nested Codex error codes (response.error.code) being invisible to the connection-limit/previous-response recovery matchers, and the session id leaking unredacted intoPI_CODEX_DEBUGlogs via thex-client-request-idheader. - Fixed
processResponsesStream(shared byopenai-responsesandazure-openai-responses) ignoring the terminalresponse.incompleteevent: a max-output-tokens-truncated response ended withstopReason: "stop", zero usage, and no cost instead of"length"with the reported token counts.response.incompleteis now handled alongsideresponse.completedand counts as stream progress for the idle watchdogs. - Fixed custom tool-call content blocks keeping the transient
partialJsonaccumulation buffer (and a potentially stalearguments.input) afterresponse.output_item.donein the shared Responses stream processor — the function_call branch already cleaned these up. - Fixed two OpenAI Codex stream-retry paths (whitespace-loop recovery and retryable provider errors) leaking native output items from the abandoned attempt into the replayed turn's
providerPayload— stale reasoning items completed before the failure were re-sent as history input on subsequent requests alongside the retry's own items. - Fixed the Codex WebSocket queue wiping already-received frames when a transport error arrived: a
response.completedqueued just before an eager server close was discarded, turning a finished response into a spuriouswebsocket closedfailure and a full request replay. Errors now append behind pending data frames. - Fixed concurrent
getOrCreateCodexWebSocketConnectioncallers (prewarm racing the first request) tearing down each other's in-flight handshake — closing a CONNECTING socket rejected the other caller with a fatalwebsocket closed before open, disabling WebSockets for the entire session. Callers now join the pending handshake. - Stopped the Codex connection-limit recovery from replaying a turn over SSE after a
toolcall_endhad already been delivered to the consumer (canSafelyReplayWebsocketOverSseguard was bypassed, re-emitting the same tool calls); the error now surfaces instead. - Extended the Codex whitespace-only argument-delta circuit breaker to
custom_tool_call_input.deltaframes, which counted as stream progress and could keep a degenerate response alive forever with no cap on buffer growth. - Fixed Codex stream failures during transport open reporting a synthetic request dump (empty URL/body) instead of the real request, and a
response.createdevent resetting the recorded time-to-first-token. - Fixed the Codex WebSocket connect watchdog timer leaking (pinning the event loop for up to 10s) when the request signal aborted before or during the handshake.
- Fixed OpenRouter-hosted Anthropic adaptive reasoning models (Claude Fable/Mythos 5 and Opus 4.6+) so the catalog exposes
xhigh; Fable/Mythos and Opus 4.7+ requests now map userhigh/xhighonto OpenRouter's Anthropicxhigh/maxeffort scale. - Fixed an unknown Anthropic
stop_reasonfailing the whole turn after the response had fully streamed.mapStopReasonthrew on unrecognized values, and since the reason arrives on the trailingmessage_deltathe error was unretryable — the livemodel_context_window_exceededstop reason (default on Sonnet 4.5+) hit this path. It now maps tolength, and any future unknown reason degrades to a logged anomaly plus a normalstopinstead of an error. - Stopped clamping API-key Anthropic requests to Claude Code's 64k output cap. The
CLAUDE_CODE_MAX_OUTPUT_TOKENSclamp exists to match the OAuth wire fingerprint, butbuildParamsapplied it unconditionally, silently halving the output budget of 128k-output models (e.g. Opus 4.8) for API-key callers. OAuth requests keep the clamp. - Stopped a successful strict-tools fallback from shipping
errorMessageon astopReason: "stop"assistant message. After a grammar-too-large 400 triggered the non-strict retry, the original 400 text was kept on the final message even when the retry succeeded — consumers that treaterrorMessagepresence as failure (e.g. balance probes) misclassified the turn, and the stale text suppressed later refusal explanations. The fallback is now logged instead. - Fixed model-supplied
User-Agentheaders being silently dropped on non-OAuth Anthropic requests.enforcedHeaderKeysfiltered the header out ofmodelHeadersin every branch but only the OAuth branch set one back; the Cloudflare-gateway, bearer-gateway, andX-Api-Keybranches now forward the caller's value verbatim. - Stopped sending the
fast-mode-2026-02-01beta header once a session has learned the endpoint+model rejects fast mode (fastModeDisabledprovider state), matching the already-droppedspeedparam. - Stopped
buildAnthropicHeadersdefaulting API-key requests onto the full Claude Code OAuth beta list (oauth-2025-04-20,claude-code-20250219, …). TheclaudeCodeBetasdefault is now OAuth-gated, matching the streaming path — the web-search header builder was the only caller hitting the default, so API-key search requests now carry just their own betas (e.g.web-search-2025-03-05). An emptyanthropic-betaheader is omitted entirely instead of being sent as an empty string. - Fixed image-bearing
developermessages being upgraded to mid-conversationsystemturns on Opus 4.8+/Fable/Mythos 5. System content is text-only on the wire, so a developer turn carrying image blocks in an upgrade-eligible position produced a 400; it now stays ausermessage. - Fixed a spliced reconnect's second envelope overwriting the completed Anthropic message:
message_deltawas not gated by the terminal-stop flag (content events and duplicatemessage_startwere), so the splice'sstop_reason/usage replaced the finished turn's — atool_useturn could be relabeledstop, and the harness then never executed the streamed tool calls. Post-terminal deltas are now logged as envelope anomalies and skipped. - Fixed a
pingarriving beforemessage_startconsuming the Anthropic first-event watchdog: the stall was then classified as a terminal mid-stream idle timeout instead of a retryable first-event timeout. Pings no longer count as the first item but still refresh the idle deadline once content is flowing. - Fixed Anthropic-compatible proxies that omit
usage/deltaobjects frommessage_start/message_delta/content_block_*envelopes crashing the turn with an unretryableTypeError; the missing payloads now degrade to logged envelope anomalies like every other malformed-frame case. - Fixed
applyPromptCachingplacingcache_controlonthinking/redacted_thinkingblocks — Anthropic rejects that with a 400. A thinking-only assistant turn inside the trailing cache window (e.g. followed by the syntheticContinue.pad) no longer receives a breakpoint. - Fixed consecutive
assistantparams reaching the wire when an empty user/developer turn between two assistant turns was dropped by the converter (e.g. an empty "nudge" submission after a length-truncated reply); Anthropic 400s on non-alternating assistant turns, and the broken triple replayed on every subsequent request. Auser: "Continue."separator is now inserted, mirroring the trailing-prefill fallback. - Fixed
supportsAdaptiveThinkingDisplaymisparsing bare dated Opus ids:claude-opus-4-20250514(Opus 4.0) parsed as minor20250514≥ 4.7, which silently dropped theinterleaved-thinking-2025-05-14beta for API-key Opus 4.0 requests. - Fixed
output_config.effortshipping without theeffort-2025-11-24beta on thinking-off requests against adaptive-only Claude models (the effort:"low" pin), and the mid-conversationsystemrole shipping withoutmid-conversation-system-2026-04-07on API-key and OAuth-utility requests; both betas are now added whenever the request can carry the corresponding field. - Fixed GitHub Copilot anthropic-messages requests going out with no
Content-Typeand noanthropic-versionheader — the copilot branch builds its headers from scratch and Bun's fetch does not defaultContent-Typefor string bodies. Both headers are now pinned to match every other branch. - Fixed Anthropic client/provider retry multiplication: with the first-event watchdog disabled (
PI_STREAM_FIRST_EVENT_TIMEOUT_MS=0), the client's internalmaxRetries: 5reactivated and stacked with the provider loop's 3 retries — up to 24 wire attempts with double backoff. The provider now pins per-requestmaxRetries: 0unconditionally. - Fixed
AnthropicMessagesClientspreadingfetchOptionsafter the core request fields, letting a caller-suppliedsignal/method/bodysilently disconnect the timeout controller or corrupt the request. Transport extras (TLS) still pass through; core fields now always win. - Fixed Foundry mTLS/CA material being cached for the process lifetime when the env vars point at files: the cache key now folds in the file mtime so on-disk certificate rotation takes effect.
- Fixed the Claude Code fingerprint version drifting across surfaces: the usage endpoint (
claude-cli/2.1.160) and OAuth bootstrap (claude-code/2.1.160) pinned a stale version while/v1/messagesreported 2.1.165; both now derive fromclaudeCodeVersion. - Fixed a system prompt that merely mentions
x-anthropic-billing-header:mid-text suppressing the entire Claude Code system-block injection (billing header, instruction, and cch attestation); the resumed-session guard now anchors withstartsWith. - Fixed lone surrogates in cross-API tool-call arguments reaching Anthropic's strict UTF-8 validation: replayed OpenAI/Google-origin
tool_use.inputstring leaves are now deep-sanitized withtoWellFormed(), while same-API Anthropic arguments stay byte-identical to keep prompt-cache prefixes stable. - Bounded the many-image resize fan-out to 4 concurrent decodes (it previously decoded every oversized image at once, two encode pipelines each — multi-GB transient memory at the 20+-image threshold that activates the feature).
- Fixed
mergeHeadersmerging case-sensitively on the Copilot/client-options path, where a miscased user-configured header (e.g.authorizationnext to the synthesizedAuthorization) survived as two keys that theHeadersconstructor joins comma-separated on the wire. - Hardened the Anthropic stream lifecycle: prologue failures (e.g. a malformed Copilot credential in
buildCopilotDynamicHeaders) and error-finalization failures now surface as anerrorevent instead of an unhandled rejection that leftstream.result()hanging forever; the spurious "cch billing placeholder not patched" warning no longer fires when the placeholder only appears in user content.
Removed
- Removed the dead
iterateUntilAborthelper (superseded byiterateWithIdleTimeout); it leaked the upstream iterator when the consumer abandoned mid-yield and had no production call sites.
[15.10.10] - 2026-06-09
Added
- Exported
wrapFetchForCchso non-streaming OAuth callers (e.g. the web-search provider) can patch the Claude Code billing-headercchattestation into their request bodies instead of shipping thecch=00000placeholder.
Fixed
- Fixed an unbounded, zero-backoff Codex WebSocket reconnect loop on
websocket_connection_limit_reached: the no-content reconnect path never consulted the retry budget and never waited, hammering the endpoint forever when the limit is account-scoped. Reconnects are now budgeted and delayed like every other WS retry path, falling back to a single SSE replay when exhausted. - Fixed the Codex whitespace-loop breaker not observing degenerate frames that arrive after their item closed (or before it opened) — those frames count as stream progress, so the idle watchdogs never fired and the turn hung forever, which is exactly the failure mode the breaker exists for. Whitespace-loop recovery now also refuses to replay the turn once a
toolcall_endwas delivered, surfacing the error instead of re-emitting the same tool calls. - Fixed the two remaining Codex retry paths (WS mid-stream reconnect and the empty-content SSE fallback) leaking blockless native output items (e.g.
web_search_call) from the failed attempt into the replayed turn'sproviderPayloadand append baseline. - Fixed Codex WebSocket failure handling closing whatever connection currently occupies the session slot — including a concurrent caller's in-flight CONNECTING handshake, whose rejection (
websocket closed before open) is classified fatal and disabled WebSockets for the whole session. Failure cleanup now skips CONNECTING sockets and the pool re-joins replacement handshakes (bounded). - Fixed the Codex request transformer not repairing orphan
custom_tool_call_outputitems (onlyfunction_call_outputwas folded into an assistant note) — a compaction splice that dropped anapply_patchcall while keeping its result produced a hard 400 on the default GPT-5 Codex toolset. - Fixed
processResponsesStreamfinalizing reasoning items via a bareitemIdcontent scan instead of the routed entry: with id-less reasoning items (local hosts), everyoutput_item.donematched the FIRST thinking block — the second item's text clobbered it and the second block was never finalized or signed. - Fixed
processResponsesStreamdropping tool calls and message text whoseoutput_item.addedevent was lost (lossy proxies):toolcall_endwas emitted with a dangling contentIndex while the call never enteredmessage.content, so the agent loop silently never executed it. The done handler now synthesizes the missing block; still-open tool-call blocks are also final-parsed atresponse.completedso thetoolUseoverride cannot hand the agent stale{}arguments. - Fixed
response.incompletewithincomplete_details.reason: "content_filter"being reported as a token-cap truncation (stopReason: "length") — the agent loop's length recovery then asked the model to "shorten" a filtered prompt. Content-filtered turns now surface as errors; usage is also populated fromresponse.failedevents, and an unknown terminal status degrades to"stop"with a logged anomaly instead of throwing away a fully-streamed response. - Fixed Copilot
premiumRequestsaccounting being dropped from failed/cancelled responses:populateResponsesUsageFromResponsereplacedusagewholesale and the error path threw before the success-path re-apply. The populate now preserves the field. - Fixed
deduplicateToolCallIdssuffixing the whole composite Responses id (callId|itemId) —normalizeResponsesToolCallIdextracts the first segment as the wirecall_idat encode time, so both copies collapsed back onto onecall_idand the request carried duplicate call/output pairs. The suffix and length budget now apply per segment. - Gated native history payload replay on api + model id in both Responses providers: after a mid-session model switch, reasoning items carrying encrypted content minted by the previous model were replayed verbatim under the new model. Replay now falls back to block re-encode (which already strips foreign signatures), matching
transformMessages' same-model trust rule. - Fixed Azure OpenAI Responses requests omitting
store: falsewhile requestingreasoning.encrypted_content(stateless-only per OpenAI), replaying custom tool calls paired with mismatchedfunction_call_outputitems (customCallIds was never threaded through), letting the SDK's internal retries (maxRetries 5) silently re-POST inside the explicit first-event deadline, and sending aprompt_cache_keywhen the caller opted out viacacheRetention: "none". - Fixed strict-pairing Responses backends (Azure, Copilot) silently discarding tool results whose call is absent from history — the result is now folded into an assistant note (same shape as orphan-output repair) so the model keeps the information.
- Fixed the OpenAI Responses first-event watchdog staying armed across the
onResponsenotification callback (a slow callback aborted an already-connected stream), Copilot transient-model retries re-attempting on an already-aborted signal (instant dead retry surfacing the scheduler's AbortError), CodexreasoningSummary: nullbeing coerced to"auto"(the documented omit-summary contract was unreachable), nested Codex error codes (response.error.code) being invisible to the connection-limit/previous-response recovery matchers, and the session id leaking unredacted intoPI_CODEX_DEBUGlogs via thex-client-request-idheader. - Fixed
processResponsesStream(shared byopenai-responsesandazure-openai-responses) ignoring the terminalresponse.incompleteevent: a max-output-tokens-truncated response ended withstopReason: "stop", zero usage, and no cost instead of"length"with the reported token counts.response.incompleteis now handled alongsideresponse.completedand counts as stream progress for the idle watchdogs. - Fixed custom tool-call content blocks keeping the transient
partialJsonaccumulation buffer (and a potentially stalearguments.input) afterresponse.output_item.donein the shared Responses stream processor — the function_call branch already cleaned these up. - Fixed two OpenAI Codex stream-retry paths (whitespace-loop recovery and retryable provider errors) leaking native output items from the abandoned attempt into the replayed turn's
providerPayload— stale reasoning items completed before the failure were re-sent as history input on subsequent requests alongside the retry's own items. - Fixed the Codex WebSocket queue wiping already-received frames when a transport error arrived: a
response.completedqueued just before an eager server close was discarded, turning a finished response into a spuriouswebsocket closedfailure and a full request replay. Errors now append behind pending data frames. - Fixed concurrent
getOrCreateCodexWebSocketConnectioncallers (prewarm racing the first request) tearing down each other's in-flight handshake — closing a CONNECTING socket rejected the other caller with a fatalwebsocket closed before open, disabling WebSockets for the entire session. Callers now join the pending handshake. - Stopped the Codex connection-limit recovery from replaying a turn over SSE after a
toolcall_endhad already been delivered to the consumer (canSafelyReplayWebsocketOverSseguard was bypassed, re-emitting the same tool calls); the error now surfaces instead. - Extended the Codex whitespace-only argument-delta circuit breaker to
custom_tool_call_input.deltaframes, which counted as stream progress and could keep a degenerate response alive forever with no cap on buffer growth. - Fixed Codex stream failures during transport open reporting a synthetic request dump (empty URL/body) instead of the real request, and a
response.createdevent resetting the recorded time-to-first-token. - Fixed the Codex WebSocket connect watchdog timer leaking (pinning the event loop for up to 10s) when the request signal aborted before or during the handshake.
- Fixed OpenRouter-hosted Anthropic adaptive reasoning models (Claude Fable/Mythos 5 and Opus 4.6+) so the catalog exposes
xhigh; Fable/Mythos and Opus 4.7+ requests now map userhigh/xhighonto OpenRouter's Anthropicxhigh/maxeffort scale. - Fixed an unknown Anthropic
stop_reasonfailing the whole turn after the response had fully streamed.mapStopReasonthrew on unrecognized values, and since the reason arrives on the trailingmessage_deltathe error was unretryable — the livemodel_context_window_exceededstop reason (default on Sonnet 4.5+) hit this path. It now maps tolength, and any future unknown reason degrades to a logged anomaly plus a normalstopinstead of an error. - Stopped clamping API-key Anthropic requests to Claude Code's 64k output cap. The
CLAUDE_CODE_MAX_OUTPUT_TOKENSclamp exists to match the OAuth wire fingerprint, butbuildParamsapplied it unconditionally, silently halving the output budget of 128k-output models (e.g. Opus 4.8) for API-key callers. OAuth requests keep the clamp. - Stopped a successful strict-tools fallback from shipping
errorMessageon astopReason: "stop"assistant message. After a grammar-too-large 400 triggered the non-strict retry, the original 400 text was kept on the final message even when the retry succeeded — consumers that treaterrorMessagepresence as failure (e.g. balance probes) misclassified the turn, and the stale text suppressed later refusal explanations. The fallback is now logged instead. - Fixed model-supplied
User-Agentheaders being silently dropped on non-OAuth Anthropic requests.enforcedHeaderKeysfiltered the header out ofmodelHeadersin every branch but only the OAuth branch set one back; the Cloudflare-gateway, bearer-gateway, andX-Api-Keybranches now forward the caller's value verbatim. - Stopped sending the
fast-mode-2026-02-01beta header once a session has learned the endpoint+model rejects fast mode (fastModeDisabledprovider state), matching the already-droppedspeedparam. - Stopped
buildAnthropicHeadersdefaulting API-key requests onto the full Claude Code OAuth beta list (oauth-2025-04-20,claude-code-20250219, …). TheclaudeCodeBetasdefault is now OAuth-gated, matching the streaming path — the web-search header builder was the only caller hitting the default, so API-key search requests now carry just their own betas (e.g.web-search-2025-03-05). An emptyanthropic-betaheader is omitted entirely instead of being sent as an empty string. - Fixed image-bearing
developermessages being upgraded to mid-conversationsystemturns on Opus 4.8+/Fable/Mythos 5. System content is text-only on the wire, so a developer turn carrying image blocks in an upgrade-eligible position produced a 400; it now stays ausermessage. - Fixed a spliced reconnect's second envelope overwriting the completed Anthropic message:
message_deltawas not gated by the terminal-stop flag (content events and duplicatemessage_startwere), so the splice'sstop_reason/usage replaced the finished turn's — atool_useturn could be relabeledstop, and the harness then never executed the streamed tool calls. Post-terminal deltas are now logged as envelope anomalies and skipped. - Fixed a
pingarriving beforemessage_startconsuming the Anthropic first-event watchdog: the stall was then classified as a terminal mid-stream idle timeout instead of a retryable first-event timeout. Pings no longer count as the first item but still refresh the idle deadline once content is flowing. - Fixed Anthropic-compatible proxies that omit
usage/deltaobjects frommessage_start/message_delta/content_block_*envelopes crashing the turn with an unretryableTypeError; the missing payloads now degrade to logged envelope anomalies like every other malformed-frame case. - Fixed
applyPromptCachingplacingcache_controlonthinking/redacted_thinkingblocks — Anthropic rejects that with a 400. A thinking-only assistant turn inside the trailing cache window (e.g. followed by the syntheticContinue.pad) no longer receives a breakpoint. - Fixed consecutive
assistantparams reaching the wire when an empty user/developer turn between two assistant turns was dropped by the converter (e.g. an empty "nudge" submission after a length-truncated reply); Anthropic 400s on non-alternating assistant turns, and the broken triple replayed on every subsequent request. Auser: "Continue."separator is now inserted, mirroring the trailing-prefill fallback. - Fixed
supportsAdaptiveThinkingDisplaymisparsing bare dated Opus ids:claude-opus-4-20250514(Opus 4.0) parsed as minor20250514≥ 4.7, which silently dropped theinterleaved-thinking-2025-05-14beta for API-key Opus 4.0 requests. - Fixed
output_config.effortshipping without theeffort-2025-11-24beta on thinking-off requests against adaptive-only Claude models (the effort:"low" pin), and the mid-conversationsystemrole shipping withoutmid-conversation-system-2026-04-07on API-key and OAuth-utility requests; both betas are now added whenever the request can carry the corresponding field. - Fixed GitHub Copilot anthropic-messages requests going out with no
Content-Typeand noanthropic-versionheader — the copilot branch builds its headers from scratch and Bun's fetch does not defaultContent-Typefor string bodies. Both headers are now pinned to match every other branch. - Fixed Anthropic client/provider retry multiplication: with the first-event watchdog disabled (
PI_STREAM_FIRST_EVENT_TIMEOUT_MS=0), the client's internalmaxRetries: 5reactivated and stacked with the provider loop's 3 retries — up to 24 wire attempts with double backoff. The provider now pins per-requestmaxRetries: 0unconditionally. - Fixed
AnthropicMessagesClientspreadingfetchOptionsafter the core request fields, letting a caller-suppliedsignal/method/bodysilently disconnect the timeout controller or corrupt the request. Transport extras (TLS) still pass through; core fields now always win. - Fixed Foundry mTLS/CA material being cached for the process lifetime when the env vars point at files: the cache key now folds in the file mtime so on-disk certificate rotation takes effect.
- Fixed the Claude Code fingerprint version drifting across surfaces: the usage endpoint (
claude-cli/2.1.160) and OAuth bootstrap (claude-code/2.1.160) pinned a stale version while/v1/messagesreported 2.1.165; both now derive fromclaudeCodeVersion. - Fixed a system prompt that merely mentions
x-anthropic-billing-header:mid-text suppressing the entire Claude Code system-block injection (billing header, instruction, and cch attestation); the resumed-session guard now anchors withstartsWith. - Fixed lone surrogates in cross-API tool-call arguments reaching Anthropic's strict UTF-8 validation: replayed OpenAI/Google-origin
tool_use.inputstring leaves are now deep-sanitized withtoWellFormed(), while same-API Anthropic arguments stay byte-identical to keep prompt-cache prefixes stable. - Bounded the many-image resize fan-out to 4 concurrent decodes (it previously decoded every oversized image at once, two encode pipelines each — multi-GB transient memory at the 20+-image threshold that activates the feature).
- Fixed
mergeHeadersmerging case-sensitively on the Copilot/client-options path, where a miscased user-configured header (e.g.authorizationnext to the synthesizedAuthorization) survived as two keys that theHeadersconstructor joins comma-separated on the wire. - Hardened the Anthropic stream lifecycle: prologue failures (e.g. a malformed Copilot credential in
buildCopilotDynamicHeaders) and error-finalization failures now surface as anerrorevent instead of an unhandled rejection that leftstream.result()hanging forever; the spurious "cch billing placeholder not patched" warning no longer fires when the placeholder only appears in user content.
[15.10.9] - 2026-06-09
Added
- Added
antigravityRankingStrategyand registered it as the defaultCredentialRankingStrategyforgoogle-antigravity, so multi-account selection consumes the per-counter Antigravity usage reports (sorted ascending byremainingFractioninfetchAntigravityUsage) before falling back to round-robin — preventing the exhausted-counter credential from being chosen first when an unblocked sibling has headroom (#2187). - Added Claude Fable 5 to the first-party Anthropic catalog, seeded directly via
ANTHROPIC_CURATED_FALLBACK_MODELSrather than waiting on models.dev (1M context / 128k output, adaptive thinking, $10/$50 per MTok). The model parser recognizes thefablekind so effort tiers (low→max), adaptive thinking, and Opus-4.7-style sampling restrictions apply; token limits and pricing are pinned inapplyAnthropicCatalogPolicy.
Fixed
- Fixed
google-antigravitynot rotating to another stored OAuth account when Cloud Code Assist returns429 You have exhausted your capacity on this model. Your quota will reset after ….parseRateLimitReasonmatched the literalcapacitybefore thequota will resetsuffix and downgraded the failure toMODEL_CAPACITY_EXHAUSTED(45–75 s backoff), andisUsageLimitErrorreturned false for the same message — somarkUsageLimitReachedwas never invoked and the agent kept hammering the exhausted credential while the retry layer bailed on the multi-hourretry-after. Both paths now treat the Antigravity phrasing asQUOTA_EXHAUSTED/ usage-limit, blocking the current credential until reset and letting the session pick an unblocked sibling (#2187). - Fixed OpenRouter Anthropic chat-completions requests placing
cache_controlon empty assistant tool-call content. The cache marker now skips empty text and attaches to the most recent non-empty text part, avoiding HTTP 400 payloads with{type:"text", text:"", cache_control:...}. - Fixed Fable-only Anthropic request shaping to cover Claude Mythos 5, and added Mythos 5 to the first-party Anthropic catalog seed. Adaptive display, sampling suppression, mid-conversation system messages, forced-tool-choice downgrade, and Bedrock adaptive metadata now handle both model families.
- Fixed adaptive-only Claude models (Opus 4.6+, Sonnet 4.6+, Fable/Mythos 5) returning HTTP 400
"thinking.type.disabled" is not supported for this modelwhenever thinking was turned off (utility calls and forced-tool turns route through the disable path). These models accept onlythinking.type: "adaptive"; the request builder now omits the thinking field and pins the lowest adaptive effort instead of emittingtype: "disabled". - Widened the OpenAI-completions first-event watchdog floor from 120s to 300s for DeepSeek V4 reasoning models hosted on the official DeepSeek API. The reasoner emits no SSE bytes until its private chain-of-thought finishes, which routinely takes longer than the generic 100s first-event budget under load — every chat then aborted with
OpenAI completions stream timed out while waiting for the first eventand silently retried. Mirrors the existing GLM coding-plan widening (#2177).
[15.10.8] - 2026-06-09
Added
- Added optional
fetchtransport override (fetch?: FetchImpl) to Google, Ollama, and OpenAI-compatible model-manager options so dynamic model discovery and metadata lookups can use a caller-supplied HTTP client instead of only globalfetch - Added optional
fetchon OAuth controller and API-key validation/login flows so token exchange, refresh, and device/PKCE login requests can be routed through a customfetchimplementation - Added optional
fetchsupport to usage polling context, allowing usage providers to execute usage checks using an injected HTTP client - Added
AssistantMessage.upstreamProvider, capturing the upstream provider an aggregator routed the request to (OpenRouter reports it via a top-levelproviderfield on every chunk, e.g."Anthropic"). Surfaced from the OpenAI-completions stream alongsideresponseId.
Fixed
- Fixed a degenerate OpenAI Codex stream (the model emits whitespace-only
function_call_arguments.deltaframes forever — commonly seen right after atodotool call) terminating the turn with an error instead of recovering. The whitespace-loop circuit-breaker now (a) stops aborting the shared per-requestAbortController—requestSignalis anAbortSignal.anyover it, so aborting latched it and made every reopen on the reusedrequestSetupimpossible — and (b) drops the half-built junk tool call and replays the request from scratch, bounded byCODEX_WHITESPACE_LOOP_RETRY_LIMIT(2). Sampling nondeterminism usually clears the loop on a fresh attempt; once the budget is exhausted the error is surfaced as before, but without the junk tool call polluting the message. - Capped requested output tokens at 64k (
OPENAI_MAX_OUTPUT_TOKENS, mirroring Anthropic'sCLAUDE_CODE_MAX_OUTPUT_TOKENS) on OpenAI-family wires with a known upstream output cap — theopenai-completionsrequest builder (non-OpenRouter) and the shared responses sampling helper (openai-responses,azure-openai-responses). A model's catalogmaxTokensoften tracks its context window rather than the upstream's per-request output cap, so requesting the full ceiling 400'd (e.g.z-ai/glm-4.7asking for 131072 output exceeded the upstream's 131072-token total context). Output is nowmin(requested, model.maxTokens, 64000). - Stopped sending
max_tokens/max_completion_tokenson OpenRouter (openrouter.ai) completions requests. OpenRouter filters out any upstream whose advertised output cap is below the requestedmax_tokens, so a value derived from the catalog (which reflects the highest-cap provider) silently excluded lower-cap upstreams —provider.order: ["cerebras"]forz-ai/glm-4.7fell through to DeepInfra because Cerebras's ~40k output cap is below the request, whileonly: ["cerebras"](no fallback target) bypassed the filter and worked. Omitting the field lets each upstream self-cap and keeps provider routing (only/order) honored. Kimi via OpenRouter stays exempt — it derives TPM rate limits frommax_tokens.
[15.10.7] - 2026-06-08
Fixed
- Fixed first-party Anthropic requests returning HTTP 400 "Invalid
signatureinthinkingblock" after interrupting the model during its visible output.transformMessagesstripped the signature from everythinkingblock of anaborted/errorturn, including blocks that had already finished streaming — Anthropic delivers a block's signature atcontent_block_stopbefore the next block starts, so a thinking block followed bytext/tool_useis fully signed. The valid signature was then replayed empty (signature: ""), which signature-enforcing Anthropic rejects, including when the provider is routed through an LLM gateway baseUrl. Only the single mid-stream block at the abort point is now treated as untrustworthy; completed thinking blocks keep their replayable signatures (#2144). - Pinned a regression test against issue #2123: OAuth requests to adaptive-thinking Claude Opus models (4.6+) ship a
context_management.edits[clear_thinking_20251015]block paired with thethinkingfield, but the eager-todo prelude (and other paths that forcetool_choicetotool/anyon the first user turn) route throughdisableThinkingIfToolChoiceForced, which would stripparams.thinkingwhile leaving the orphancontext_managementbehind. The Anthropic API then rejected the request with400 ... clear_thinking_20251015 strategy requires thinking to be enabled or adaptive. The fix that lands in [15.10.5] now drops both fields together; the new test locks the contract so the strategy can never outlive its enablingthinkingpayload again. - Fixed Antigravity usage counters so exhausted Google/Gemini quota renders as
0% freewhile separate Anthropic/OpenAI-backed Antigravity model counters remain visible independently, without replaying stale pre-fix cached usage reports.
[15.10.6] - 2026-06-08
Added
- Added AIML API as an OpenAI-compatible provider preset with
AIMLAPI_API_KEYdiscovery (#2105).
[15.10.5] - 2026-06-08
Breaking Changes
- Renamed the OAuth subpath export
@oh-my-pi/pi-ai/utils/oauth→@oh-my-pi/pi-ai/oauth(and@oh-my-pi/pi-ai/utils/oauth/*→@oh-my-pi/pi-ai/oauth/*, e.g.oauth/types,oauth/callback-server,oauth/openai-codex) after relocating the OAuth implementation out ofutils/oauth/intoregistry/oauth/. The high-level OAuth API (getOAuthProviders,refreshOAuthToken,getOAuthApiKey,registerOAuthProvider,unregisterOAuthProviders,getOAuthProvider) and theOAuth*types stay exported from the package root, unchanged.
Changed
- Changed Anthropic retry handling to avoid retrying 4xx responses other than 408 and 429
- Optimized the Anthropic
cchattestation patch to locate the billing-header placeholder with nativeBuffer.indexOf(memmem) instead of a hand-rolled byte loop. The marker sits ~99% through the body (messagesserializes beforesystem), so the old scan walked almost the entire payload; output bytes are unchanged but the patch is ~7.5x faster (563µs -> 75µs on a 1MB body). - Refactored provider configuration to a single-source registry (
registry/, renamed fromprovider-registry/with itsproviders/subdir flattened up). TheKnownProvider/OAuthProvidertype unions,PROVIDER_DESCRIPTORS,DEFAULT_MODEL_PER_PROVIDER, theserviceProviderMapenv-key fallbacks, the/loginprovider list (builtInOAuthProviders), and therefreshOAuthToken/AuthStorage.logindispatch are all derived from it. Provider defs live directly underregistry/; thin provider-specific login flows are inlined into the def file, while heavier provider-local OAuth flows and the shared OAuth flow infra (callback-server,pkce,google-oauth-shared,types, runtimeindex) now live together underregistry/oauth/(previously split acrossprovider-registry/providers/oauth/andutils/oauth/). The non-OAuth API-key paste/validation helpers (api-key-login,api-key-validation) sit beside the defs inregistry/. Adding a provider that reuses an existing wire API is now one new provider def plus one registry entry in the common case. ExposesPROVIDER_REGISTRY,getProviderDefinition,ProviderDefinition, andPASTE_CODE_LOGIN_PROVIDERS.
Fixed
- Disabled OpenAI Codex Responses stream obfuscation by sending
stream_options.include_obfuscation=false, reducing raw WebSocket/SSE debug noise and bandwidth. - Interrupted OpenAI Codex Responses streams that emit long runs of whitespace-only tool-call argument deltas, preventing degenerate WebSocket/SSE responses from filling the raw stream buffer indefinitely.
- Preserved streaming responses when Anthropic emits unrecognized content_block envelopes by ignoring unknown blocks and continuing to emit known content
- Applied cache control to the most recent tool result block when building Anthropic OAuth payloads without a preceding text block, enabling ephemeral caching for tool-result-only messages
- Kept Anthropic sampling parameters (temperature, top_p, top_k) when thinking is explicitly disabled
- Fixed raw Anthropic SSE handling by parsing event frames with strict JSON parsing and matching event-type validation, surfacing malformed frames as stream errors instead of repairing them
- Fixed Anthropic stream envelope handling to reject duplicate
content_block_startindexes and block deltas/stops for unopened blocks, preventing malformed envelope states from producing partial output - Fixed Anthropic image conversion to normalize
image/jpgtoimage/jpegand emit a placeholder for unsupported image MIME types - Fixed Anthropic thinking request preparation by clamping
max_tokensto provider/model limits and adjusting thinking budgets to a valid value - Fixed Anthropic request shaping around forced tool choice, unsigned thinking replay, prompt-cache marker placement, non-Anthropic bearer gateways, Foundry TLS loading, and strict tool-schema normalization so malformed or incompatible request payloads are rejected locally or shaped consistently before streaming
- Fixed the Anthropic stream parser shipping a truncated tool call as a completed turn. When a transport drop cut the SSE stream mid-
tool_useand a transparent reconnect spliced a fresh message envelope onto the same stream, the duplicatemessage_startwas deduped but the orphaned tool block — which never received itscontent_block_stop— survived in the assistant message with its seed{}(or partially-parsed) arguments. The terminal stop signal from the reconnect then let it flow through as a normal tool call, so e.g. areaddispatched with{}failed downstream validation (path: expected string, received undefined). The parser now treats any tool block left open at stream end as a truncated envelope and routes it through the existing retry/error path instead of emitting bogus arguments. - Fixed the Zhipu Coding Plan login prompt advertising a misleading
sk-...placeholder. Zhipu API keys are formatted<id>.<secret>(nosk-prefix), so the placeholder now matches the actual format instead of suggesting the wrong shape. (#2106) - Fixed Moonshot
kimi-k2.6(and any futurekimi-k2.x) discovered viaMOONSHOT_API_KEYstalling on first turn with no output. ThemoonshotModelManagerOptionsdiscovery mapper only marked ids containing"thinking"asreasoning: true, so dynamickimi-k2.6entries fell through withreasoning: false; the openai-completions z.ai branch was then skipped and the request reached Moonshot with nothinkingparameter at all. Moonshot K2.6 requires an explicitthinking: {type}field (the same native-API wire shape #1838 introducedthinking.keepfor), so the server held the stream silently. The mapper now stampsreasoning: true, vision input, and defaultthinkingmetadata on everykimi-k2.xid, restoring the explicitthinking: {type: "disabled"|"enabled"}wire body the Moonshot endpoint expects. (#2113)
[15.10.4] - 2026-06-08
Added
- Added
anthropic-client-platform(desktop_app) andanthropic-client-version(1.11187.4) headers to the Anthropic request fingerprint for OAuth sessions
Changed
- Changed non-built-in tool names sent to Anthropic from
proxy_prefixing to_prefixing (for examplebashto_bash) while built-in tool names remain unchanged - Updated the Anthropic OAuth stealth fingerprint to track Claude Code 2.1.165:
claudeCodeVersionbumped to2.1.165(flows into both thecc_versionbilling header and theclaude-cli/<version>user-agent),claudeCodeSystemInstructionchanged to"You are a Claude agent, built on Anthropic's Claude Agent SDK.", and the billing-headercc_entrypointchanged fromclitolocal-agent. - Clamped the Anthropic request
max_tokenstoMath.min(CLAUDE_CODE_MAX_OUTPUT_TOKENS, options.maxTokens || model.maxTokens)(64k) so OAuth requests match Claude Code's requested output cap instead of sending the model's full ceiling (e.g. 128k for Opus 4.8).
[15.10.3] - 2026-06-08
Removed
- Removed the synthetic
<turn-aborted>developer guidance note thattransformMessagesinjected after an aborted/errored assistant turn (and itsturn-aborted-guidance.mdprompt). The per-call synthetic"aborted"tool results already tell the model the turn's tools were terminated, so the extra "verify current state before retrying" note was redundant — and it biased the model toward second-guessing a deliberate user interrupt when the turn was resumed. - Removed the legacy Anthropic first-user-message skip for
<system-reminder>blocks now that synthetic reminders no longer travel as user messages.
[15.10.2] - 2026-06-08
Added
- Added support for
impersonated_service_accountApplication Default Credentials (ADC) in Vertex AI to enable chained impersonation without failing via 401invalid_client. - Added
AuthStorage.getCredentialOrigin(provider)(returning a structuredCredentialOrigin/CredentialOriginKind) andgetEnvApiKeyName(provider), so callers can render where a provider's auth comes from — runtime override, config, stored OAuth/api-key, env var (with the backing variable name), or fallback resolver — without parsing the prose ofdescribeCredentialSource.
Changed
- Changed
onSseEventrecording for OpenAI Responses, Azure OpenAI Responses, OpenAI Completions, and Anthropic stream providers to emit reconstructed SSE events from decoded SDK stream items instead of wrapping raw fetch responses - Changed OpenAI Completions SSE diagnostics to include
event: "chat.completion.chunk"inonSseEventrecords for chunked responses - Changed the default Anthropic model in
DEFAULT_MODEL_PER_PROVIDERfromclaude-sonnet-4-6toclaude-opus-4-6, so sessions that fall back to the provider default (no configureddefaultrole, no--model, no restored session) now start on Claude Opus 4.6.
Fixed
- Fixed duplicate upstream
tool_call_idvalues collapsing distinct tool calls during message transformation, preserving one call/result pairing per emitted tool call before provider replay and keeping generated duplicate IDs distinct after OpenAI/Mistral wire-length caps. (#2055) - Fixed the Anthropic provider retrying persistent account usage/quota limits (e.g.
429 "This request would exceed your account's rate limit",usage_limit_reached) as if they were transient. Because the error text contains "rate limit",isProviderRetryableErrormatched it and the stream retry loop looped through its 2s/4s/8s backoff (then thestreamSimplea/b/c policy re-minted the credential and ran the whole thing again) before surfacing the failure — even though the server'sretry-afterparked the account for minutes-to-hours. These errors are now recognized viaisUsageLimitErrorand surfaced immediately to the credential-rotation layer, so e.g.omp dry-balance --benchreports a rate-limited account as failed at once instead of appearing to hang. - Fixed MiniMax-compatible OpenAI-completions hosts losing tool-call argument content when
function.argumentsis streamed as an object across more than one delta. The accumulator added in #1776 wroteblock.partialArgs = rawArgsper chunk, so every chunk but the last was overwritten — for aneditcall this surfaced as a tail-slice of the patch text being applied (e.g. a single-linereplace 91..91:body extending the deletion across the surrounding rows). Chunks are now shallow-merged; for shared string keys,startsWithdistinguishes cumulative restatements (take the latest) from per-chunk-delta fragments (concatenate). Per-chunktoolcall_deltaemission for the object branch is suppressed (the previous code emittedJSON.stringify(rawArgs)per chunk, which fed downstream concat consumers —packages/agent/src/proxy.ts,openai-chat-server,openai-responses-server,anthropic-messages-server— an invalid sequence like{"input":"a"}{"input":"b"}); the merged object is flushed instead as a single concat-safe delta infinishToolCallBlockbeforetoolcall_end, so accumulators reconstruct the args correctly. The single-chunk shape covered by the existing #1776 regression test stays correct end-to-end. (#2080) - Fixed the OpenAI Responses compatibility server misrouting late
toolcall_deltaevents for earlier parallel tool calls after a latertoolcall_start. The encoder now keeps OpenFunctionCall state by content index, allocates output indexes at item start, and closes each tool item by its owntoolcall_end, preserving deferred MiniMax object-argument flushes for the matching call. (#2080)
[15.10.1] - 2026-06-07
Breaking Changes
- Removed the
onAuthErroroption from stream request options and shifted auth retry handling to resolver-basedapiKeybehavior, requiring callers using custom auth-retry hooks to migrate
Added
- Added
ApiKeyResolverandApiKeyauth helpers, includingisApiKeyResolver,isAuthRetryableError,resolveApiKeyOnce, andwithAuth, and exported them from the package root - Added support for a function-valued
apiKeyinSimpleStreamOptionsso a single stream request can refresh or rotate credentials during retry - Added
forceRefreshcredential option toAuthStorage.getApiKeyandrotateSessionCredentialsupport for session-level credential rotation after auth failures - Added
AuthStorage.resolver(provider, options)method that builds anApiKeyResolverimplementing the a/b/c auth-retry policy directly on the storage instance
Changed
- Changed gateway and stream auth flows to share the a/b/c retry policy, refreshing the same session credential first and then switching to a sibling credential on repeated auth failures
Fixed
- Fixed streaming auth retries to handle
401and usage-limit errors before replay-unsafe content is emitted, including failures surfaced only viaerrorStatus - Fixed tool argument validation to coerce singleton non-string values into arrays when the schema expects an array, preventing Anthropic-compatible models that emit
todo.opsas an object from getting stuck in repeated validation-error loops. (#2026) - Fixed streaming retries to buffer and suppress partial
startevents from failed auth attempts so only clean retried events are delivered - Fixed the HTTP 400 raw-request dumper (
appendRawHttpRequestDumpFor400) littering the real~/.omp/logs/http-400-requestsdirectory during tests. Provider suites exercise the 400 error path with mockedfetchresponses, which the dumper could not distinguish from genuine failures; it now skips persistence under the Bun test runner (isBunTestRuntime()). - Fixed Anthropic Opus requests unnecessarily forcing
tool_choice.disable_parallel_tool_use, allowing Claude Opus to use the provider's default parallel tool-calling behavior again. - Fixed parallel
function_callitems losing arguments against llama.cpp's OpenAI Responses endpoint (/v1/responses), where every call but the last finalized with{}and the agent rejected them withpath: Invalid input: expected string, received undefined. llama.cpp'sto_json_oaicompat_respemitsoutput_item.addedwith onlyitem.call_id(noitem.id, nooutput_index) while the matchingfunction_call_arguments.deltacarriesitem_id: "fc_<call_id>".processResponsesStreamnow registers function-call and custom-tool-call items underitem.call_idas a secondary lookup key (alongsideitem.id/output_index) so identifier-deviant hosts route deltas and done events to the right block. (#2015) - Fixed
PI_REQ_DEBUGresponse recording truncating the captured body when a streamed response was cancelled mid-flight. The response tee inwrapResponsecould callFileRequestDebugResponseLog.close()from both thecancelcallback and the resumedpull(which observesdoneonce the source reader is cancelled); the second caller saw the handle already nulled and returned before the first caller's pending write flushed, so the.res.loglost the already-buffered chunk.close()now memoizes its flush-and-close promise so every caller awaits the same completion.
[15.10.0] - 2026-06-06
Added
- Added a dependency-free
@oh-my-pi/pi-ai/effortmodule exporting theEffortenum andTHINKING_EFFORTS, split out ofmodel-thinkingso hot-path consumers can import the thinking levels without pulling inmodel-thinkingand its provider-compat dependency graph. The package barrel still re-exports both names, so existing imports are unaffected.
Fixed
- Fixed Antigravity usage provider emitting one bar per model instead of deduplicating by tier — a single account's 15+ model entries now collapse to one bar per tier, matching the shared-quota reality of the upstream API.
- Fixed Antigravity usage reports missing
emailandaccountIdin metadata, so the/usagedisplay and the deduplicator can associate reports with their credentials. - Fixed usage-report dedup ignoring
projectIdfor Google Cloud providers, preventing duplicate credential entries from being recognized as the same account. - Fixed Cloud Code Assist (Antigravity / Gemini CLI) rejecting the
githubtool with HTTP 400 when theprparameter schema containedanyOf: [string, array]. The CCA mixed-type combiner collapse picked the first non-null type (string) but indiscriminately copied type-specific keys from variant branches —itemsfrom the array variant leaked onto the string-typed result, producing{type: "string", items: {...}}which Google's API rejects as invalid. The collapse now filters merged variant fields against the winning type's allowed key set. (#2002) - Fixed OpenAI Responses-family providers (Codex, OpenAI Responses, Azure Responses) rejecting requests with
400 No tool output found for function call …after the user branched/navigated the session tree to a node that ends on a tool call (the tool-result child is dropped from the reconstructed history) or after a turn was aborted/crashed between the call streaming and its result persisting. The converters now synthesize a placeholderfunction_call_output/custom_tool_call_outputimmediately after any unpairedfunction_call/custom_tool_call, symmetric to the existing orphan-output repair, so the model still sees the call and can recover instead of the whole request 400ing. - Fixed Anthropic-compatible reasoning endpoints losing prior-turn reasoning on continuation requests when they emit unsigned
thinkingblocks.convertAnthropicMessagestreated unknown endpoints as signature-enforcing and demoted unsigned reasoning totype: "text", which destabilized tool-call argument serialization on the next turn — the upstream symptom behind theargs?.ops?.map is not a functioncrash reported against thetodotool. Officialapi.anthropic.comkeeps the conservative text fallback; non-officialanthropic-messagesreasoning models now replay unsigned reasoning as nativetype: "thinking"(#2005).
[15.9.67] - 2026-06-06
Fixed
- Fixed llama.cpp/OpenAI Responses parallel tool calls losing arguments when
function_call_arguments.doneevents omitoutput_indexanditem_id, by routing those identifierless final-argument events through the open function calls in item order. (#1970) - Fixed local Ollama (
openai-responses) turns failing with HTTP 400invalid reasoning value: "minimal"when a discovered model ran withminimal(orxhigh) thinking. Ollama's OpenAI-compatiblereasoning.effortonly acceptshigh|medium|low|max|none, so discovered reasoning-capable Ollama models now carry acompat.reasoningEffortMapremappingminimal → lowandxhigh → max; non-reasoning models are left untouched.
[15.9.2] - 2026-06-05
Added
- Added an AES-256-GCM auth-broker snapshot cache module and
RemoteAuthCredentialStoreOptions.onSnapshotso broker clients can persist broker-sourced full snapshots without blocking startup on every run. - Added
Model.omitMaxOutputTokensso providers (notably Ollama proxies fronting cloud catalogs) can suppressmax_output_tokens(Responses) andmax_tokens/max_completion_tokens(Completions) on the wire while still using the catalogmaxTokensfor local budgeting. Without it,applyCommonResponsesSamplingParamsunconditionally sent the catalog cap and HTTP-400'd against upstream APIs whose true output limit was unknown to OMP. (#1881)
Changed
- Changed usage-ranked OAuth credential selection to pick deterministic session-sticky weighted buckets instead of always choosing the top-ranked account, capping the best account at 2x the baseline session likelihood while keeping equal-priority accounts evenly balanced.
Fixed
- Fixed parallel
function_callitems on the OpenAI Responses API losing arguments on every call except the last when the upstream server interleaves their stream events (observed against llama.cpp and other local Responses-compat hosts).processResponsesStreamno longer routesfunction_call_arguments.{delta,done},output_item.done, content_part/text/refusal/reasoning events through a singletoncurrentItem/currentBlockreference; it now tracks every open item in registries keyed byoutput_indexanditem_idso each event is folded into the matching block and the emittedtoolcall_endcarries the correctcontentIndex. (#1880)
[15.9.1] - 2026-06-04
Added
- Added regional Xiaomi Token Plan login/provider entries (
xiaomi-token-plan-sgp,xiaomi-token-plan-ams,xiaomi-token-plan-cn) soomp logincan store token-plan keys against the selected region. (#1846)
Fixed
- Removed the
context-1m-2025-08-07(1M long-context) beta from the Anthropic agent request headers, the OAuth model-discovery header, and the Claude usage-API header. Sending it caused subscription/OAuth requests without long-context credits to fail with429 Usage credits are required for long context requests, breaking Sonnet. The remaining betas are unchanged. - Fixed Kimi K2.x
maxTokenson Fireworks and Fire Pass (fireworks/kimi-k2.5,fireworks/kimi-k2.6,firepass/kimi-k2.6-turbo) being inherited from Fireworks/v1/modelsdiscovery (max_completion_tokens: 65536) rather than the published Kimi-on-Fireworks output budget, which let callers (and the openai-completions default-injection safety net) ship a budget the router cannot honor and made runaway reasoning traces more likely. The Fireworks resolver now clamps every Kimi K2.x id (public catalog ids and the canonicalaccounts/fireworks/{models,routers}/kimi-k2…wire form) to 32,768 output tokens, and the generator applies the same cap as a post-processing safety net so thefirepassstatic fallback and the bundledfireworksentries stay in sync across regens. (#1849) - Fixed Xiaomi Token Plan MiMo OpenAI-compatible tool-call continuations omitting required
reasoning_contentreplay. (#1846) - Fixed Anthropic prompt caching for OpenAI-compatible Claude proxies by honoring
compat.cacheControlFormat: "anthropic"outside OpenRouter. (#1845) - Fixed Moonshot Kimi K2.6 silently pausing for many seconds between tool calls because the server discarded the
reasoning_contentthat omp was already sending with every assistant tool-call replay. The K2.6thinkingparameter takes an extrakeepfield whose default (null) ignores historical reasoning, so K2.6 had to re-derive its full chain-of-thought from the user prompt on every iteration of the agent loop. The Moonshot direct (api.moonshot.ai) and Kimi Code (api.kimi.com) wire bodies now sendthinking: { type: "enabled", keep: "all" }forkimi-k2.6requests with reasoning enabled, matching Moonshot's documented best practice for multi-step tool-calling agents. The flag is gated on the K2.6 id and the two native hosts because earlier Moonshot models (K2.5 and below) 400 on the unknown field and every Kimi gateway (OpenRouter, OpenCode, Kilo, Fireworks, …) speaks its own thinking shape. (#1838) - Fixed Alibaba DashScope (Bailian) compatible-mode endpoint
400 InternalError.Algo.InvalidParameter: The provided messages input is invalid. The error info is [Unexpected item type in content.]when a screenshot or other image-producing tool result was folded into a known text-only Qwen turn (e.g.qwen3.7-max,qwen-max,qwen3-coder-*) hosted atdashscope.aliyuncs.com/compatible-mode/v1.convertMessagesinopenai-completionsno longer forwardsimage_urlcontent parts for those text-only id families even when a misconfigured custom provider claimsinput: ["text", "image"]; multimodal compatible-mode ids such asqwen3.7-plusandqwen-vl-maxstill rely on the cataloginputfield. The tool-result branch and the user-content branch both fall back to the standard[image omitted: model does not support vision]placeholder for text-only ids so the model still sees the attachment intent. (#1859)
[15.9.0] - 2026-06-04
Fixed
- Fixed MiniMax-compatible OpenAI-completions hosts (e.g.
minimax-code-cn/MiniMax-M3) losing tool-call arguments when the stream deliversfunction.argumentsas a complete object instead of the OpenAI JSON-string contract. The streaming buffer previously concatenated the object into a string, coercing it to[object Object]and leavingbash/editcalls with empty or malformed inputs; the tool-call block now holds the object payload directly. (#1776) - Fixed Cloud Code Assist (Gemini / Antigravity) rejecting tool schemas with
Invalid JSON payload received. Unknown name "propertyNames"(HTTP 400) when a tool exposed a property literally namedproperties(e.g. the Resend MCPcreate_contacttool). The schema normalizer'sinsidePropertiesflag was re-asserted when descending into such a property's value schema, so Google-unsupported keywords (propertyNames,additionalProperties, …) nested inside it were never stripped. The flag is now only set when entering a realpropertiesmap from a schema node, not from within anotherpropertiesmap. - Fixed local/self-hosted providers leaking machine-specific endpoints into the bundled
models.json. Agenerate-modelsrun on a machine with a LiteLLM proxy baked 1202litellmmodels pinned tohttp://localhost:4000/v1into the committed catalog.litellm(andlm-studio) now joinollama/vllmin the generator's discovery-only exclusion set, so local providers are never fetched during generation nor written tomodels.json— they are discovered dynamically at runtime instead. LiteLLM model discovery now enriches metadata against models.dev (the same reference source the other gateway providers use) rather than a bundled reference map. Added a regression test pinning the invariant (no local provider blocks, no loopback/private-networkbaseUrls in the bundled catalog).
[15.8.2] - 2026-06-03
Fixed
- Fixed
opencode-zen/minimax-m3-free(and forward-compatopencode-zen/minimax-m3) andopencode-go/minimax-m3being routed toanthropic-messagesdespite the OpenCode Zen/Go gateways only serving these ids at/v1/chat/completions, which surfaced raw MiniMax/tool-call markup (<invoke name="bash">,<tool_call>,<description>,<cwd>,<|minimax|>) in the UI. Resolver overrides now pin these ids toopenai-completionsand the bundledmodels.jsonentries are flipped to match. (#1617) - Fixed MiniMax Coding Plan China login opening the international
platform.minimax.iosubscription page instead of the Chinaplatform.minimaxi.compage.
[15.8.0] - 2026-06-02
Added
- Added
AnthropicMessagesClientand related Anthropic wire types/errors viaanthropic-clientexport so callers can build a standalone Anthropic Messages client without depending on@anthropic-ai/sdk - Added
parseClaudeRateLimitHeadersandAuthStorage.ingestUsageHeadersso Anthropic rate-limit response headers can warm the per-credential usage cache with throttling while preserving per-tier data from the last full usage report.
Changed
- Changed Anthropic request handling to use the package-local
AnthropicMessagesClientimplementation instead of@anthropic-ai/sdkas the default transport - Updated the
AnthropicOptions.clientsurface to accept anyAnthropicMessagesClientLikeimplementation withmessages.create, enabling custom compatible clients - Changed generated OAuth metadata
user_idto use a deterministicdevice_idderived from the install ID instead of a random value claudeCodeVersionbumped to2.1.148to match current Claude Code release.X-Stainless-Package-Versionupdated to0.94.0(matches the bundled@anthropic-ai/sdkversion);X-Stainless-Runtime-Versionpinned tov24.3.0(Bun version bundled with CC 2.1.148);X-Stainless-Osheader key corrected toX-Stainless-OS.createClaudeBillingHeadernow emits a deterministic billing header (cc_version=<claudeCodeVersion>.<suffix>; cc_entrypoint=cli; cch=00000;), where<suffix>is the first 3 hex chars ofSHA-256(salt + msg[4] + msg[7] + msg[20] + version)instead of random bytes. The fingerprint seed is taken from the first user message (skipping synthetic/developer injections), mirroring Claude Code'scomputeFingerprintFromMessages.cchattestation implemented:cch=00000is a placeholder that, for OAuth requests,wrapFetchForCchrewrites on the wire toXXHash64(body, 0x4D659218E32A3268) & 0xFFFFFformatted as 5 lowercase hex chars, computed in-place viaBun.hash.xxHash64. The rewrite is anchored to thesystem[0]billing-header prefix so user content is never mutated, and is installed only when a billing-header prefix is present (OAuth turns).anthropic-betaheader set for OAuth model discovery and Claude usage-API requests expanded to addcontext-1m-2025-08-07,redact-thinking-2026-02-12,mid-conversation-system-2026-04-07,advanced-tool-use-2025-11-20,effort-2025-11-24, andextended-cache-ttl-2025-04-11. The usage-APIuser-agentis bumped toclaude-cli/2.1.158 (external, cli).- Reasoning models now append
effort-2025-11-24to the per-requestAnthropic-Betaheader (matches Claude Code). buildAnthropicSystemBlocks(CC-instruction mode) now emits the same 3-block layout as Claude Code: billing header (never cached), system instruction (cached), all user content merged into one block with\n\n(cached). Previously emitted one block per item with cache only on the last, which fingerprinted the caller by block count.applyPromptCachingnow matches Claude Code's breakpoint layout: 2 system (instruction + merged content) + 2 message, with no tool breakpoint. The tool breakpoint was redundant — tools follow system in the token sequence, so when system changes the tool cache prefix also changes. The instruction block (system[1]) is stable across every request and now gets its own guaranteed-hit breakpoint.applyPromptCachingnow caches the last two messages regardless of role instead of the last two user messages. The penultimate assistant message (tool calls + response from the previous turn) is larger and more recently created than the penultimate user message, making it the higher-value cache target.- OAuth scope set expanded: added
user:sessions:claude_code,user:mcp_servers,user:file_upload.AUTHORIZE_URLstays atclaude.ai/oauth/authorizeandTOKEN_URLstays atapi.anthropic.com/v1/oauth/token— theplatform.claude.comequivalents are CC's console-credential flow and do not grantuser:inference, which OMP requires for direct OAuth-token inference. - Token refresh POST now sends
anthropic-beta: oauth-2025-04-20andUser-Agent: anthropic-sdk-typescript/0.94.0 userOAuthProvider(CC sends these on refresh but not on the initial code exchange).
Fixed
- Fixed tool argument validation to wrap a plain string in a singleton array when the schema requires an array, allowing tool-level path/list normalization to recover from bare string arguments.
- Restored
eager_input_streamingand strict flags on OAuth Anthropic tool definitions when model compatibility allows eager streaming. - Fixed OAuth stream calls with injected custom clients missing a
betaclient by falling back toclient.messages.createinstead of requiringclient.beta.messages.create - Fixed direct use of internal API client typing so retry/timeouts and malformed-error classification remain compatible while not requiring the external SDK
- Fixed Cursor provider requests failing with
Cannot send empty user message to Cursor APIafter tool-result history by selecting the latest user/developer turn instead of assuming the final context message is the active user turn. - Fixed Anthropic web search dropping
ANTHROPIC_CUSTOM_HEADERSwhenCLAUDE_CODE_USE_FOUNDRYwas unset, causing 401s from corporate API gateways.resolveAnthropicCustomHeadersForBaseUrlnow forwards the parsed headers whenever the base URL is non-Anthropic (or Foundry is enabled), andbuildAnthropicSearchHeadersthreads them throughbuildAnthropicHeadersso the search and streaming paths behave identically (#1693). - Fixed OpenCode Go Anthropic-format models such as
qwen3.7-maxsending AnthropicX-Api-Keyauth alongside the OpenCode bearer token, avoiding spurious Alibaba401 Invalid API-key providederrors. (#1661) - Fixed OAuth token exchange and refresh flows to fetch Claude CLI bootstrap identity when token responses omit account information, so
accountIdandemailare now recovered when available - Fixed Anthropic thinking traces being lost on direct OAuth requests. OAuth requests no longer send
redact-thinking-2026-02-12unless thinking is explicitly hidden, Opus 4.7+ adaptive thinking opts intodisplay: "summarized", and the top user-facing thinking tier now sends Anthropic'soutput_config.effort = "max"rather than the next-lower"xhigh"tier.
Removed
- Removed the
@anthropic-ai/sdkruntime dependency. The Anthropic provider now uses the package-localAnthropicMessagesClientand hand-maintained wire types inproviders/anthropic-wire.ts; the SDK was only ever used for URL assembly, auth-header injection, bounded retries, the pre-response timeout, and HTTP-error-to-status mapping, all of which are reproduced with identical observable behavior.
[15.7.5] - 2026-06-01
Added
- Added Anthropic task budget support, forwarding
taskBudgetasoutput_config.task_budgetwith the requiredtask-budgets-2026-03-13beta header and accepting Anthropic gateway requests that sendoutput_config.task_budget.
Fixed
- Fixed OpenAI-family first-event timeouts so
PI_OPENAI_STREAM_IDLE_TIMEOUT_MScannot be undercut by a lower genericPI_STREAM_FIRST_EVENT_TIMEOUT_MSwhile local OpenAI-compatible servers are still processing large prompts.PI_OPENAI_STREAM_FIRST_EVENT_TIMEOUT_MSis now available for an explicit OpenAI-specific first-event override. (#1603)
[15.7.4] - 2026-05-31
Fixed
- Fixed Anthropic stream idle-timeout retries after the provider stream has already begun.
- Fixed Xiaomi MiMo
/loginrejecting token-plan (tp-) keys with401 Invalid API Key. The validation request was still sending the legacy Anthropicx-api-keyheader against the OpenAI-compatible/v1/chat/completionsendpoint; switched toAuthorization: Bearer, matching the runtime path. (#1580) - Fixed OpenAI-compatible tool-call replay to send empty assistant content instead of
null, avoiding strict custom backends that crash withstr/NoneTypeconcatenation after subagent tool results. (#1585)
[15.7.3] - 2026-05-31
Changed
- Throttled per-delta streaming JSON re-parsing of OpenAI Responses/Codex tool-call arguments (bounding mid-stream parse cost from O(N²) to O(N)). Finalization via
response.output_item.donenow writes the authoritative full arguments back to the persisted assistant-message block, so tool calls finalized without a trailingresponse.function_call_arguments.doneno longer retain stale/empty ({}) arguments. (#1507)
[15.6.0] - 2026-05-30
Fixed
- Fixed Anthropic adaptive-thinking replay preserving signed thinking blocks on the latest abandoned tool-use assistant message, avoiding
thinking blocks in the latest assistant message cannot be modified400s. (#1531)
[15.5.15] - 2026-05-30
Added
- Added
PI_REQ_DEBUG=1request/response recording for provider transports. Each request writesrr-session-N.json; each received response writesrr-session-N.res.logwith response headers followed by raw body bytes.
Fixed
- Fixed OpenCode-Go dynamic model refresh downgrading
qwen3.7-maxfrom Anthropic Messages to OpenAI-compatible transport, which caused401 Model qwen3.7-max is not supported for format oa-compatafter/v1/modelscache refreshes.
[15.5.12] - 2026-05-29
Removed
- Removed ANTML stream markup healing for
antml:function_callsandantml:thinkingenvelopes, so Anthropic-compatible providers no longer parse those tags intotoolCall/thinkingevents
Fixed
- Fixed GLM-5.x coding-plan OpenAI-compatible streams to use a longer default watchdog window, avoiding spurious
OpenAI completions stream stalled while waiting for the next eventerrors during slowglm-5.1thinking/output phases. (#1494) - Fixed
zhipu-coding-planmodel discovery and credential validation to use the dedicated GLM Coding Plan endpoint (https://open.bigmodel.cn/api/coding/paas/v4) instead of the general BigModel endpoint, preventing requests from consuming ordinary account balance. (#1494) - Fixed DeepSeek tool calls failing on NanoGPT (e.g.
nanogpt/deepseek/deepseek-v4-prowith reasoning enabled) by routing tool-bearing DeepSeek requests through NanoGPT's:toolsmodel route and addingnanogptto the DSML leak allowlist so streamed<|DSML|tool_calls>...</|DSML|tool_calls>envelopes are healed into structured tool calls instead of being passed through as visible text. (#1488) - Fixed DeepSeek tool calls failing on NanoGPT (e.g.
nanogpt/deepseek/deepseek-v4-prowith reasoning enabled) by addingnanogptto the DSML leak allowlist so streamed<|DSML|tool_calls>...</|DSML|tool_calls>envelopes are healed into structured tool calls instead of being passed through as visible text. The:toolsmodel suffix is no longer appended on NanoGPT; that route triggered NanoGPT's server-side tool-call parser and 502'd withcode: "malformed_tool_call"on complex tool schemas (todo_write) — the default route forwardsdelta.content(including DSML envelopes) which is healed client-side. (#1488) - Fixed OpenAI-compatible streamed parallel tool calls losing indexed argument deltas by tracking active tool-call blocks by the provider's
tool_calls[].index; this keeps parallel NanoGPTreadcalls from merging or dropping theirpatharguments. (#1488)
[15.5.11] - 2026-05-29
Added
- Added mid-conversation
systemmessage support for Anthropic Messages by upgrading eligibledeveloperturns torole: "system"on first-party Claude API with Claude Opus 4.8+ and newer - Added
supportsMidConversationSystemto Anthropic compatibility settings so consumers can opt in to or disable mid-conversationsystemrole handling per model - Added
anthropic.claude-opus-4-8model metadata in the model registry for Bedrock Converse streaming with effort-based thinking support throughxhigh
Changed
- Changed Anthropic adaptive-thinking effort mapping for Opus 4.7+ on the Messages API to use the model's full five-tier scale: user-facing efforts now shift up one notch (
minimal→low,low→medium,medium→high,high→xhigh,xhigh→max) so the top tier reaches the genuinemaxlevel andhighlands on Anthropic's recommendedxhighcoding/agentic default. Older adaptive models (Opus 4.6) and Bedrock Converse keep the four-tier legacy mapping wherexhighaliases tomax.
Fixed
- Fixed OpenCode Zen
400 thinking is enabled but reasoning_content is missing in assistant tool call messagefor every model behindopencode-go/opencode-zen(Kimi K2.x, DeepSeek V4 Pro/Flash, GLM-5.x, Qwen3.x, MiMo, MiniMax) by reactivatingrequiresReasoningContentForToolCallsand pinning the wire field toreasoning_contentfor any opencode request in thinking mode. The static compat default still omits the field for thinking-disabled turns to preserve theExtra inputs are not permittedguard from #1071; forced-tool turns also stay off because the existingdisableReasoningOnForcedToolChoiceguard strips thinking from the wire body. (#1484)
[15.5.8] - 2026-05-28
Added
- Added
CheckCredentialsOptions.completionProbe(andcompletionTimeoutMs) soAuthStorage.checkCredentialscan additionally exercise each credential against the provider's chat-completion endpoint after refresh-on-expiry. Result lands onCredentialHealthResult.completion({ok, reason?, modelId?, latencyMs?}) without disturbing the usageokfield. Public types:CompletionProbe,CompletionProbeInput,CompletionProbeCredential,CredentialCompletionResult. The probe is invoked even when noUsageProvideris registered for the row, and is skipped when OAuth refresh fails (the stale bytes would only mask the upstream failure). - Added Wafer Pass and Wafer Serverless providers (
wafer-pass,wafer-serverless). OpenAI-compatible (https://pass.wafer.ai/v1), bearer auth,wfr_…keys./login wafer-passand/login wafer-serverlesspaste-and-validate the key against/v1/models.WAFER_PASS_API_KEYandWAFER_SERVERLESS_API_KEYenvironment variables wired intogetEnvApiKey. Bundled catalog seedswafer-pass/{GLM-5.1, Qwen3.5-397B-A17B}andwafer-serverless/{GLM-5.1, Kimi-K2.6, Qwen3.5-397B-A17B, Qwen3.6-35B-A3B, qwen3.7-max, deepseek-v4-flash, deepseek-v4-pro}; dynamic discovery via/v1/modelsoverlays additional models at runtime. Pass-tier discovery filterswafer.tier === "pass_included". Pass-SKU costs are seeded at0(flat-rate subscription, no per-token charge — matcheskimi-code/firepass/alibaba-coding-plan). Serverless costs are the wafer.ai retail rate, derived from the*_cents_per_millionenvelope viavalue × 125 / 10000(e.g. GLM-5.1120→ $1.50/M, Kimi-K2.688→ $1.10/M). Reasoning entries get a thinking compat picked from thewafer.providerenvelope:zai/moonshotai→ zai-stylethinking: { type },qwen→ top-levelenable_thinking,deepseekand unknown upstreams stay unset sodetectOpenAICompatcan pickreasoning_effortfrom the id pattern at request time.
Changed
- Changed auth-gateway credential resolution to use per-conversation
promptCacheKey/sessionIdwhen callingAuthStorage.getApiKey, so repeated turns can keep the same credential until it becomes unavailable - Changed auth-gateway and pi-native request handling to align
sessionIdwith prompt/context identity before credential lookup - Changed Anthropic prompt preparation to downscale image blocks over 2000px when a request includes 20+ images, reducing oversized payloads automatically
- Changed OpenAI chat request parsing to accept
nameontoolmessages and fall back to the matching assistanttool_callsname, so parsed tool results now carry a proper tool name when the wire omits it - Changed
checkCredentialsto skip runningcompletionProbewhen OAuth refresh fails, so stale bearer tokens are never probed and the refresh failure remains the returnedreason - Changed completion reporting to return
completion: { ok: null, reason: ... }when a credential has no usable bearer bytes instead of attempting the probe - Refactored
AuthStorage.checkCredentialsso OAuth refresh-on-expiry runs up-front and the refreshed credential is shared between the usage probe and the new completion probe; rows without a registeredUsageProviderno longer short-circuit before the completion probe runs.
Fixed
- Fixed DeepSeek DSML tool-call envelope leaks on Ollama Cloud and OpenAI-compatible streams by healing leaked envelopes into structured tool calls without displaying raw DSML markers. (#1462)
- Fixed auth-gateway to classify usage-limit messages such as
usage_limit_reached,resource_exhausted, and Codex-styleTry again in ~X mintext as 429rate_limit_errorresponses - Fixed auth-gateway usage-limit handling to honor parsed retry hints and switch to a sibling credential via
markUsageLimitReachedinstead of invalidating the rate-limited credential - Fixed
streamSimpleto retry on usage-limit errors (including message-only error events) before any content is emitted, soonAuthErrorcan rotate credentials automatically - Fixed auth-gateway error classification to extract embedded status codes and use word-boundary matching, so
GenerateContentRequestand similar messages are no longer misreported as rate-limit errors - Fixed
checkCredentialsto handlecompletionProbeexceptions by recording the failure inCredentialHealthResult.completion.reasonwhile still returning the usage probe result - Fixed Google Vertex's bundled model list to use the authoritative models.dev catalog, including MaaS entries such as
deepseek-ai/deepseek-v3.2-maasand removing retired Gemini 1.5 fallbacks. (#1456)
[15.5.7] - 2026-05-27
Added
SimpleStreamOptions.openrouterVariant("nitro","floor","online","exacto", …) — when set, appends:<variant>to OpenRouter model IDs at request time, leaving ids that already carry an explicit:suffixuntouched. Plumbed throughopenai-completionsand the pi-native gateway forwarder.- xAI Grok OAuth (SuperGrok Subscription) provider in
/login. Loopback PKCE flow on127.0.0.1:56121; the token unlocks Grok-4.x chat. Ported from NousResearch/hermes-agent (MIT). - OpenRouter provider in
/login. API-key paste flow validated againsthttps://openrouter.ai/api/v1/auth/key(the/modelsendpoint is public and cannot validate auth). The pasted key is stored under the existingopenrouterprovider id used byOPENROUTER_API_KEY. XAI_OAUTH_TOKENenvironment variable accepted as a headless fallback for the xAI Grok OAuth provider.
Changed
OpenAIResponsesOptionsgains four optional, provider-agnostic fields that adapter wrappers can use to compose provider-specific behavior on top of the generic transport:includeEncryptedReasoning(gatesinclude: ["reasoning.encrypted_content"]; defaulttrue, preserves current behavior),filterReasoningHistory(strips replayedtype: "reasoning"items from conversation history; defaultfalse),headers(merged onto the client's default headers), andextraBody(merged into the request payload).- The existing
XAI_API_KEYpath is unchanged — it continues to use the OpenAI-completions transport.
Fixed
- Fixed OpenRouter DeepSeek V4 tool-call follow-up requests replaying normalized
reasoningas-is instead of DeepSeek's requiredreasoning_content, which caused HTTP 400 errors in thinking mode. (#1445)
[15.5.6] - 2026-05-27
Added
- Added
PI_CODEX_WEBSOCKET_MAX_IDLE_REUSE_MSto control how long an idle Codex WebSocket stays eligible for reuse, with0disabling the check
Fixed
- Fixed reused Codex WebSocket connections that had gone silent without activity to be dropped and replaced with a fresh handshake after the idle-reuse threshold, preventing stalled next requests
- Fixed stale response frames left in the websocket queue from a completed turn so subsequent requests no longer process terminal frames from the previous response
- Fixed websocket dead-socket detection to fail a stale connection when no inbound traffic or pong is observed after a ping timeout, improving recovery on runtimes that do not emit pong events
[15.5.5] - 2026-05-27
Added
- Added
PI_CODEX_WEBSOCKET_PING_INTERVAL_MSto configure the interval for Codex WebSocket protocol ping heartbeats - Added
PI_CODEX_WEBSOCKET_PONG_TIMEOUT_MSto configure the Codex WebSocket pong timeout used to detect unresponsive connections - Added
PI_CODEX_WEBSOCKET_MESSAGE_QUEUE_CAPACITYto configure the maximum buffered Codex WebSocket inbound queue size before transport fallback
Changed
- Improved Codex WebSocket timeout diagnostics to include last event type and time since last progress event
- Enhanced Codex WebSocket error classification to recognize ping, pong, send, and queue-overflow failures as retryable
Fixed
- Fixed Codex WebSocket send failures by wrapping socket.send() in try-catch and surfacing errors as retryable transport errors
- Fixed Codex WebSocket inbound queue overflow by adding capacity bounds and triggering fallback to SSE when exceeded
- Fixed Codex WebSocket pong timeout detection by tracking pong events and failing the connection when no pong is received within the configured timeout
- Fixed Anthropic streaming to suppress hallucinated meta-prompt thinking blocks (the recent "I don't see any current rewritten thinking..." regression). When the marker phrase
rewritten thinkingappears in a streamed thinking summary the block is collapsed to a plainThinking...placeholder and its signature is dropped so subsequent turns can't re-anchor on the garbled chain. - Fixed Codex WebSocket silent stalls by adding protocol pings, inbound queue bounding, clearer idle-timeout diagnostics, and SDK retry clamping for first-event timeouts.
[15.5.0] - 2026-05-26
Added
- Added
zhipu-coding-planprovider for Zhipu (智谱) BigModel's domestic coding-plan SKU athttps://open.bigmodel.cn/api/coding/paas/v4, with dynamic model discovery (ZHIPU_API_KEY), zai-format thinking,reasoning_contentfield, and OAuth login flow (#1340).
Removed
- Removed the
pi-aiCLI binary (packages/ai/src/cli.ts) and itsbinentry. Use the in-process equivalent in the omp coding-agent CLI:omp auth-broker login [provider],omp auth-broker logout [provider], andomp auth-broker list. The library API (AuthStorage.login(),getOAuthProviders(), etc.) is unchanged.
Fixed
- Fixed delayed
toolResultemissions so real tool results are emitted in the correct assistanttoolCallwindow after handoff/compaction, preventing out-of-order or orphaned tool results - Fixed delayed
toolResulthandling for aborted calls so a late real result is emitted instead of a syntheticabortedresult for the sametoolCallId - Fixed usage polling to disable credentials when OAuth refresh fails definitively (for example
invalid_grant) and clear cached last-good usage data so stale reports no longer remain visible
[15.4.3] - 2026-05-26
Fixed
- Fixed Google Vertex model discovery to use the project-scoped OpenAI-compatible model list so Vertex Model Garden models such as GLM and Claude are available through ADC auth (#1412).
[15.4.2] - 2026-05-26
Fixed
- Fixed OpenCode Zen
big-picklefollow-up requests replaying assistant tool-call turns without DeepSeek-requiredreasoning_content, which caused HTTP 400 errors in thinking mode.
[15.4.1] - 2026-05-26
Added
- Added
isOpenAICompletionsProgressChunkexport to identify real progress chunks vs. keepalives in OpenAI completions streams - Added per-provider stream watchdog overrides via
getStreamIdleTimeoutMs(fallbackMs)andgetStreamFirstEventTimeoutMs(idleTimeoutMs, fallbackMs)to allow providers like Google Gemini CLI to extend first-event timeouts without affecting global defaults - Added
promptCacheKeytoStreamOptionsand passed it through stream option mapping so callers can specify an explicit prompt-cache key separate fromsessionId - Added
promptCacheKeysupport to the native server option whitelist sopromptCacheKeyis accepted bypi-native-serverstreams - Restored the per-provider stream watchdog (
iterateWithIdleTimeout) on top of the abortable iterator. The lazy stream forwarder inregister-builtinsnow wraps every provider's event stream with the first-event + steady-state idle watchdog (PI_STREAM_FIRST_EVENT_TIMEOUT_MS,PI_STREAM_IDLE_TIMEOUT_MS; aliases honored), and Anthropic / OpenAI Completions / OpenAI Responses / Azure OpenAI Responses / Codex SSE re-emit their per-provider progress predicates so empty keepalive frames cannot keep a stalled stream alive. Reverts the partial regression from #1392 that left Codex WebSocket subagent runs hanging silently for hours when the broker dropped frames between deltas. The Codex WebSocket transport additionally now resetslastProgressAtonly on progress events (not keepalives), giving the 300s WS-internal idle ceiling the same liveness semantics as the SSE path.
Changed
- Enabled OpenAI Codex WebSocket streams to apply
streamIdleTimeoutMsandstreamFirstEventTimeoutMsfromStreamOptionsper request instead of fixed internal defaults - Changed stream idle watchdog implementation from
iterateUntilAborttoiterateWithIdleTimeout, which now enforces maximum idle gaps between streamed events and distinguishes between first-event and steady-state timeouts - Changed Anthropic, OpenAI Responses, OpenAI Completions, Azure OpenAI Responses, and OpenAI Codex Responses providers to use the new idle-timeout iterator with per-provider progress predicates so empty keepalive frames cannot keep a stalled stream alive
- Changed Codex WebSocket transport to reset
lastProgressAtonly on progress events (not keepalives), giving the 300s WS-internal idle ceiling the same liveness semantics as the SSE path - Changed Google Gemini CLI stream forwarding defaults to use a 5-minute first-event floor via per-provider lazy-stream limits to avoid premature first-event timeouts on slow startup
- Changed OpenAI Responses and OpenAI Codex request handling to keep
sessionIdfor provider routing and conversation headers whilepromptCacheKeycontrols theprompt_cache_keypayload independently - Changed
StreamOptions.streamIdleTimeoutMsdocumentation to clarify it is now wired into every built-in provider and the lazy stream forwarder, and thatstreamFirstEventTimeoutMsis honored at both the SDK-request layer and the iterator-watchdog layer - Changed OpenAI Responses and OpenAI Codex request handling so
sessionIdcontinues to drive provider routing and state whilepromptCacheKeycontrols theprompt_cache_keypayload - Changed Google Gemini CLI stream forwarding defaults to use a 5-minute first-event floor to avoid premature first-event timeouts on slow startup
- Changed auth-gateway request mapping to preserve incoming
prompt_cache_keyas bothpromptCacheKeyandsessionIdwhen routing OpenAI-compatible sessions - Un-deprecated
StreamOptions.streamIdleTimeoutMs; the option is wired into every built-in provider and the lazy stream forwarder again.streamFirstEventTimeoutMsis now honored at both the SDK-request layer (viacreateSdkStreamRequestOptions) and the iterator-watchdog layer, in cooperation.
Removed
- Removed
installH2Fetchand thefetchpatch that forced HTTP/2 on HTTPS requests; callers now use the default Bunfetchtransport
Fixed
- Fixed first-item timeout handling so
iterateWithIdleTimeoutno longer keeps first-event timers active after the source throws or the consumer stops before semantic progress - Fixed silent multi-hour hangs on Codex WebSocket subagent runs when the broker dropped frames between deltas by restoring per-provider stream watchdogs with progress-event filtering
- Fixed z.ai/GLM-via-OpenRouter subagent stalls where no-op keepalive chunks reset the idle watchdog indefinitely by filtering non-progress items before resetting the deadline
[15.4.0] - 2026-05-26
Breaking Changes
- Removed
findAnthropicAuthfromanthropic-authand replaced store-driven auth discovery withbuildAnthropicAuthConfig, requiring callers to provide an already-resolved API key before building Anthropic auth config
Added
- Added
PI_CODEX_WEBSOCKET_FIRST_EVENT_TIMEOUT_MSandPI_CODEX_WEBSOCKET_IDLE_TIMEOUT_MSoptions to tune Codex WebSocket timeout behavior before fallback - Added
AuthStorage.getOAuthAccessto return a refreshed OAuth access token with identity metadata (accountId,email,projectId,enterpriseUrl) for callers that need bearer-token headers together - Added Codex WebSocket forwarding to the
onSseEventobserver so the raw provider-stream debug viewer captures the inbound JSON frames and the outbound request frame from the WS transport using the same synthesized SSE-wire shape (event:+data:lines, prefixed with a: ws ← <type>(inbound) or: ws → <type>(outbound) comment).
Changed
- Changed OAuth selection in
AuthStorageto treat credentials as stale when they are within 60 seconds of expiry and rotate them preemptively - Changed Google Gemini CLI, Google Gemini usage, Antigravity usage, and Kimi usage flows to stop refreshing OAuth tokens directly and rely on
AuthStoragefor token rotation
Deprecated
- Deprecated
streamIdleTimeoutMsinStreamOptionsas a compatibility-only field that is no longer used by providers
Removed
- Removed provider-local OAuth refresh helpers from Google Gemini CLI and Google/Kimi/Antigravity usage probes, preventing direct refresh calls from those usage paths
Fixed
- Dropped truncated, thinking-only assistant turns with only
thinking/redacted_thinkingblocks and notextortoolcontent during message transformation, preventing Anthropic requests from sending consecutive assistant messages after amax_tokens/error/abortedinterruption - Fixed Amazon Bedrock bearer-token authentication to honor
AWS_BEARER_TOKEN_BEDROCKbefore resolving AWS profiles or runningcredential_process, matching Bedrock API-key precedence. (#1399) - Updated
isRetryableErrorto treat Bun HTTP/2 transport errors (HTTP2StreamReset,HTTP2RefusedStream) as retryable so transient stream-reset failures can be retried - Fixed Codex WebSocket streaming to recover from stalled sessions by falling back to SSE when the first event or subsequent progress is delayed beyond the configured websocket timeout
- Fixed expired OAuth handling so provider-level paths no longer attempt direct token refresh calls for expired credentials and instead rely on
AuthStoragefor rotation - Fixed provider streams aborting slow-but-valid first tokens or silent inter-event gaps with OMP-owned first-event/idle watchdog errors. Built-in lazy streams, OpenAI/Anthropic/Azure/Codex SSE, and Codex WebSocket streams now wait for provider output, provider/socket errors, caller aborts, or explicit request-layer timeouts instead of treating provider silence as failure (#1392).
- Fixed Claude Opus 4.7 on Amazon Bedrock streaming no reasoning output (and appearing to hang on long reasoning runs) because Anthropic silently switched the adaptive-thinking display default to
"omitted". The Bedrock provider now sendsthinking.display = "summarized"by default on Opus 4.7+ adaptive models and on budget-based Claude models, mirroring the existing direct-Anthropic behavior.BedrockOptions.thinkingDisplay("summarized" | "omitted") is exposed for callers that want to opt out, andhideThinkingSummarynow wires through to the Bedrock case (#1373). - Fixed Cursor Composer resume/tool-continuation turns failing with
Cannot send empty user message to Cursor API. Empty current user turns now use Cursor'sresumeActioninstead of constructing an invaliduserMessageAction(#1376). - Fixed
pi-ai login moonshotfailing withinvalid temperature: only 1 is allowed for this model(HTTP 400) because the API-key validator probedkimi-k2.5withtemperature: 0. Moonshot login now validates againstGET /v1/models, matching the DeepSeek/Fireworks/NanoGPT/ZenMux pattern and authenticating the key without invoking model-specific parameter restrictions.
[15.3.2] - 2026-05-25
Added
- Added
GET /v1/snapshot/streamfor live auth-broker snapshot updates via SSE withsnapshot,entry, andremovedevent frames - Added
AuthBrokerClient.openSnapshotStream()for consuming SSE snapshot streams from/v1/snapshot/stream - Added
streamSnapshotsoption toRemoteAuthCredentialStore(defaulttrue) to enable or disable SSE-based snapshot synchronization - Added
streamKeepaliveMstostartAuthBroker()to tune heartbeat frequency for the SSE stream - Added
AuthStorage.checkCredentials({ signal?, timeoutMs?, baseUrlResolver? })that returns a per-credentialCredentialHealthResultwith tri-stateok(true/false/null-unverifiable), the credential's identity (provider, type, email/accountId, broker-refresh flag), and the upstream error string when the probe fails. Iterates sequentially overlistAuthCredentials(), exercises OAuth refresh on expiry, then calls the per-providerUsageProvider.fetchUsagewithout swallowing errors — so callers can identify which row in a multi-account broker is producing 401s instead of getting a silently-deduplicatedfetchUsageReportslist. - Added
GET /v1/credentials/checktostartAuthGateway()that forwards toAuthStorage.checkCredentialsand returns{ generatedAt, credentials }. Gated by the same bearer as the rest of the gateway.
Changed
- Changed
RemoteAuthCredentialStoreto prefer SSE snapshot streaming and automatically fall back to long-polling when a broker returns 404 for/v1/snapshot/stream - Changed snapshot write-refresh flow so
RemoteAuthCredentialStoreskips immediate/v1/snapshotrefreshes when SSE streaming is active - Changed broker SSE stream behavior to keep connections open with periodic keepalives and an increased server idle timeout
[15.3.0] - 2026-05-25
Added
- Added DeepSeek to the built-in API-key login provider catalog so
omp login deepseekstores a reusableDEEPSEEK_API_KEYcredential for the bundled DeepSeek models.
[15.2.4] - 2026-05-22
Fixed
- Fixed ChatGPT Plus/Pro (Codex) OAuth login returning
Token exchange failed: 403on Windows. When port 1455 was in use, the callback server silently fell back to a random port; OpenAI's authorization endpoint accepts any localhost redirect URI (loose validation), so the browser callback succeeds and shows "Authentication Successful", but the token endpoint rejects the non-registered port with 403. TheOpenAICodexOAuthFlownow enforces a fixedredirectUrioption so a busy port immediately surfaces as "port unavailable" instead of producing a confusing 403 (#1277). - Improved
exchangeCodeForTokenerror diagnostics: the 403 response body (error/error_descriptionfields) is now included in the thrown message, matching the existingrefreshOpenAICodexTokenbehaviour.
Added
- Added
ChatGPT Plus/Pro (Codex, headless/device)(openai-codex-device) as an alternative login method for the Codex provider. Uses OpenAI's device-code flow (/api/accounts/deviceauth/usercode→ poll/api/accounts/deviceauth/token), which avoids a local callback server and port 1455 entirely. Credentials are stored under the existingopenai-codexprovider key so all models and tooling continue to work without reconfiguration (#1277).
[15.2.2] - 2026-05-22
Fixed
- Fixed
gemini-3.1-pro-highandgemini-3.1-pro-lowon thegoogle-antigravityprovider always returning HTTP 400 from Cloud Code Assist. TheANTIGRAVITY_SYSTEM_INSTRUCTIONidentity header was not injected for these models because the internal check matched the string"gemini-3-pro-high"(hyphen) instead of the versioned"gemini-3.1-pro-..."form. The guard now matches allgemini-3model variants (#1274).
[15.2.0] - 2026-05-21
Fixed
- Fixed
/login(and/logout, plus anyAuthStorage.set/removecall) against a remote auth-broker throwingRemoteAuthCredentialStore is read-only on the client. Use 'omp auth-broker login <provider>' to mutate credentials.Added three optional async write hooks toAuthCredentialStore(upsertAuthCredentialRemote,replaceAuthCredentialsRemote,deleteAuthCredentialsRemote);RemoteAuthCredentialStoreimplements them via the broker'sPOST /v1/credentialandPOST /v1/credential/:id/disableendpoints and applies the broker's authoritative post-write entries to the local snapshot.AuthStorageroutes through the hooks when present, so OAuth and API-key logins (and logouts) initiated from a broker-backed client now persist server-side and surface immediately without waiting for the long-poll snapshot tick.
[15.1.9] - 2026-05-21
Fixed
- Fixed Ollama named tool forcing to send only the requested tool when the caller passes a named
toolChoice, preservingtool_choice: "required"while preventing local models from selecting a different tool. (#1236) - Fixed
/btw(and IRC background replies) returning aBedrockException400 (The toolConfig field must be defined when using toolUse and toolResult content blocks.) on LiteLLM → Bedrock once the session has tool-call history. Two source fixes inbuildParams: (1)if (context.tools)→if (context.tools?.length)so an explicitcontext.tools = [](the /btw opt-out) never routes throughconvertToolsand never emits an empty"tools"array; (2)else if (hasToolHistory(...))→else if (context.tools === undefined && hasToolHistory(...))so the Anthropic-proxy sentinel that injectstools: []for tool-history turns is suppressed when the caller explicitly opted out, preventing it from re-introducing the empty array. As defence-in-depth,tool_choice: "none"is also dropped when the resolved tools list is missing or empty. (#1227)
[15.1.8] - 2026-05-20
Added
- Added Fireworks Fire Pass as a separate
firepassprovider with API-key login flow, bundledkimi-k2.6-turbomodel entry (Kimi K2.6 Turbo), and wire-id translation from the friendly catalog id to theaccounts/fireworks/routers/kimi-k2p6-turborouter endpoint. Fire Pass keys (fpk_…) authorize only the dedicated router and reject/v1/models, so login validation pings chat completions against the router id directly. Extended the openai-completions Kimi-family safety net so the firepass entry inherits the per-Fireworks-docs "always sendmax_tokens" default (Kimi K2 guide); the router's acceptedreasoning_effortset includesxhigh, so it is forwarded verbatim rather than remapped. See https://docs.fireworks.ai/firepass.
Fixed
- Fixed DeepSeek V4 direct API requests with tools to keep documented thinking mode instead of dropping reasoning: lower OMP efforts now map to DeepSeek's supported
high,tool_choiceis omitted,thinking: { type: "enabled" }andmax_tokensare sent, and partial userreasoningEffortMapoverrides merge with DeepSeek defaults. (#1207) - Fixed model cache schema v2 databases so offline refreshes preserve cached provider discoveries after upgrading to schema v3 and subsequent online refreshes can overwrite the cache. (#1219)
- Fixed Perplexity OAuth credentials being treated as expired one hour after login.
getJwtExpirywas fabricatingexpires = now + 1hwhenever the JWT had noexpclaim (the common case — Perplexity sessions are server-side). Once the hour elapsed,getOAuthApiKeywould mark the cred expired and the search provider's loader would silently skip it, surfacing as "logged out". Logins with noexpnow persist a far-future sentinel;getOAuthApiKeyalso normalizes any staleexpireswritten by older builds.
[15.1.7] - 2026-05-19
Added
- Added Anthropic realization of
serviceTier: "priority". The anthropic-messages provider now setsspeed: "fast"on the request and appends thefast-mode-2026-02-01beta toAnthropic-Betawhenever the caller passesserviceTier: "priority". When the server rejects an unsupported model withinvalid_request_error, the provider transparently retries the same turn without the fast-mode signal (mirroring the strict-tools fallback pattern), persists the disable via a newproviderSessionState.fastModeDisabledflag so subsequent requests in the session skip the field, and surfaces the action via the newAssistantMessage.disabledFeaturesarray (id"priority") so callers can sync user-facing toggles. A newclearAnthropicFastModeFallback(providerSessionState)helper lets callers re-arm priority after the auto-fallback fired. - Added scoped
ServiceTiervalues:"openai-only"(priority onopenai/openai-codex, ignored elsewhere) and"claude-only"(priority on directanthropic, ignored on Bedrock/Vertex Claude and elsewhere). A newresolveServiceTier(serviceTier, provider)helper computes the effective tier for the provider; existing OpenAI/Anthropic provider code routes through it, soservice_tierand Anthropic fast-mode emission both respect scope.getPriorityPremiumRequestsnow counts Anthropic+priority as one premium request (previously zero) and continues to ignore providers that drop the field on the wire.
Fixed
- Fixed Anthropic fast mode (
serviceTier: "priority") looping on 429rate_limit_error: "Extra usage is required for fast mode."for accounts without the extra-usage entitlement.isAnthropicFastModeUnsupportedErrornow matches the 429 phrasing in addition to the 400invalid_request_error"does not support thespeedparameter" case, so the provider dropsspeed: "fast"on the in-turn retry, setsproviderSessionState.fastModeDisabledfor the remainder of the session, and surfacesdisabledFeatures: ["priority"]to the caller instead of retrying with the same payload untilPROVIDER_MAX_RETRIESis exhausted.
[15.1.6] - 2026-05-19
Fixed
- Fixed
{}(empty JSON Schema, the wire representation ofz.unknown()) being passed verbatim to grammar-constrained samplers (llama.cpp, etc.) inadditionalProperties,items, and other schema-valued positions across every provider (OpenAI, Anthropic, Google, Ollama, Bedrock, Cursor). Grammar builders treat{}as "generate an empty object" rather than "any JSON value", causing open-typed fields (e.g.extra.titlefromz.record(z.string(), z.unknown())) to always emit{}instead of the intended string/number/etc.toolWireSchemanow applies a newnormalizeEmptySchemaspass (exported) to both the Zod and TypeBox/raw-JSON-Schema branches, converting{}→true(semantically identical per JSON Schema draft 2020-12 §4.3.1) in all schema-valued positions. Strict-mode opt-out is preserved across all providers: OpenAI'shasUnrepresentableStrictObjectMaphits the=== truebranch instead of theisJsonObject({})branch (same result); Anthropic'snormalizeAnthropicStrictSchemaNodeopts out viaadditionalProperties !== false(still true fortrue); Google'snormalizeSchemaForGooglestripsadditionalPropertiesregardless (pre-existing). (#1179) - Fixed
pi-ai login <provider>crashing withUnknown providerfor providers that only theauth-storagelogin()switch knew about (perplexity, alibaba-coding-plan, gitlab-duo, huggingface, opencode-zen/go, lm-studio, ollama, cerebras, fireworks, qianfan, synthetic, venice, litellm, moonshot, together, cloudflare/vercel ai gateways, vllm, qwen-portal, nvidia, xiaomi, and any custom OAuth provider). The CLI now delegates toSqliteAuthCredentialStore.login()instead of duplicating a smaller switch, so the auth-brokeromp auth-broker login <provider>flow works for every registered OAuth provider.
[15.1.4] - 2026-05-19
Changed
- Updated auth-gateway format and pi-native request handling to invalidate the failed API key and retry the provider request with a replacement key when authentication fails
Fixed
- Fixed OpenAI Responses and Codex tool schema normalization to emit
properties: {}for no-argument object schemas without rewriting literal payloads. (#1147) - Fixed Anthropic 400 (
unexpected tool_use_id found in tool_result blocks ... Each tool_result block must have a corresponding tool_use block in the previous message) when handoff/compaction folds an assistanttool_useinto the handoff summary string but leaves the matching user-sidetool_resultmessage in the history.transformMessagesnow indexes everytool_useid surviving the first pass and drops orphantool_resultmessages whose originator was compacted away, preserving the text payload as a user-level<stale-tool-result>note so the model still sees what the tool returned. The note is emitted withrole: "user"rather thanrole: "developer"so providers that elevate developer-role messages (Ollama:developer→system; OpenAI chat-completions reasoning models:developer→developer) cannot lift stale tool output to an instruction-priority tier above the surrounding user/developer messages. - Fixed streaming authentication retry to trigger when a provider emits a 401
errorevent after astartevent but before any replay-unsafe content is emitted - Added
credential_processsupport to the Bedrock provider's AWS credential resolver so profiles delegating to external brokers (aws-vault,granted, in-house tools) resolve instead of falling through toUnable to resolve AWS credentials. Parses the AWS SDKVersion: 1JSON envelope, honorsExpirationin the per-profile cache, propagatesAbortSignalto the spawned helper, routes Windows.cmd/.bathelpers throughcmd.exe /c, and ships a POSIX-shell-style tokenizer that preserves backslashes inside double quotes so Windows paths survive (#1142)
[15.1.3] - 2026-05-17
Breaking Changes
- Changed
AuthBrokerClient.fetchSnapshot()to return status-based results (200or304) instead of always returning a raw snapshot body, so callers now need to branch onstatus - Renamed public schema utilities in
@oh-my-pi/pi-ai/utils/schemaby replacingsanitizeSchemaForGoogle,sanitizeSchemaForCCA,prepareSchemaForCCA, andsanitizeSchemaForMCPwithnormalizeSchemaForGoogle,normalizeSchemaForCCA, andnormalizeSchemaForMCP - Added MCP schema normalization via
normalizeSchemaForMCPfor compatibility checks - Removed the
StringEnumhelper from@oh-my-pi/pi-ai/utils/schema. Usez.enum([...])directly; Zod's emitted JSON Schema is already wire-compatible with Google and other providers. - Renamed the concrete SQLite credential store class from
AuthCredentialStoretoSqliteAuthCredentialStore.AuthCredentialStoreis now the persistence interface implemented by both the SQLite store and the newRemoteAuthCredentialStore. Updatenew AuthCredentialStore(db)/AuthCredentialStore.open(...)call-sites toSqliteAuthCredentialStore; type-position uses (store: AuthCredentialStore) continue to work unchanged.
Added
- Added
onAuthErrortoStreamOptionsand wiredstreamSimple()to retry once with a replacement API key when the first provider response is a 401 before any assistant events are emitted - Added generation-aware snapshot metadata (
generation,serverNowMs,refresher, androtatesInMs) to auth-broker snapshot responses to support client-side credential-rotation planning - Added
transport: "pi-native"onModeland the matchingstreamPiNativeclient. Whenmodel.transport === "pi-native",streamSimpleshort-circuits the per-provider dispatch and POSTs the canonicalContextto the auth-gateway'sPOST /v1/pi/streamendpoint. The response is SSE-framedAssistantMessageEvents parsed byreadSseJsonand pushed verbatim into the localAssistantMessageEventStream— no wire-format translation, no partial-stripping reconstruction. Used by containerized omp installs (robomp slots, swarm extension, etc.) to route every LLM call through a credential-holding sidecar; the slot itself never sees the real provider tokens. Server-controlled fields (apiKey,signal,fetch, lifecycle callbacks, the provider-session map) are stripped from the wire body —apiKeyrides in theAuthorizationheader as the gateway bearer. - Added
POST /v1/pi/streamto the auth-gateway. Same auth + abort + model-resolution + codex-compat + prefix-cache plumbing as the foreign-wire routes; only the wire-format translation is skipped. Request body is{ modelId, context, options?, stream? }wherecontextis the canonical pi-aiContextandoptionsisSimpleStreamOptionswith non-serializable fields stripped. Response is SSE-framedAssistantMessageEvent(terminated bydata: [DONE]) when streaming, or{ message: AssistantMessage }JSON whenstream: false. - Added Vertex AI authentication via Google Application Default Credentials from
GOOGLE_APPLICATION_CREDENTIALS,~/.config/gcloud/application_default_credentials.json, or metadata server tokens, with token caching and refresh skew control viaGOOGLE_VERTEX_REFRESH_SKEW_MS - Added support for Anthropic image message parts with
type: "url"andtype: "file"sources - Added
stopSequencesandfrequencyPenaltyto shared stream options and wired them through to OpenAI request translation - Added optional request cancellation support to auth-broker interactions by propagating
AbortSignalinto health, snapshot, usage, and refresh calls - Added
AuthStorage.setConfigApiKey/removeConfigApiKey/clearConfigApiKeysfor config-sourced per-provider bearers (e.g.models.ymlproviders.<name>.apiKey). The new tier sits between runtime--api-keyand stored credentials ingetApiKey/peekApiKeyresolution, so a bearer pinned in config now beats the broker's OAuth access token. Also suppresses OAuthaccount_uuidattribution when active, since outbound auth is the explicit config bearer, not OAuth.describeCredentialSourcereports"config override (models.yml)"for visibility. - Added per-model
additional_rate_limitsparsing toopenaiCodexUsageProvider. The Codexwham/usageendpoint surfaces a separateGPT-5.3-Codex-Sparkrate limit (metered_feature: codex_bengalfox) on Pro accounts; these now emit dedicatedopenai-codex:spark:{primary,secondary}UsageLimitentries withscope.tier = "spark", mirroring how Anthropic exposesanthropic:7d:sonnetseparately from the umbrellaanthropic:7dbucket. The osx-widgets client already keyed spark detection offlimit.id.includes("spark"); this populates that contract end-to-end. - Added
GET /v1/usageto the auth-broker API to expose aggregated usage reports fromAuthStorage.fetchUsageReports - Added auth-broker usage polling response handling that returns normalized usage reports plus generation timestamp for clients (5-min per-credential cache via
AuthStorage) - Added the auth-broker subsystem (
@oh-my-pi/pi-ai/auth-broker) for sharing OAuth credentials across machines without leaking refresh tokens. startAuthBroker(...)boots aBun.serveHTTP server exposingGET /v1/healthz,GET /v1/snapshot,POST /v1/credential(upsert),POST /v1/credential/:id/refresh, andPOST /v1/credential/:id/disable.AuthBrokerClientis the matching HTTP client used by remote clients.RemoteAuthCredentialStoreis a client-sideAuthCredentialStorethat mirrors a broker snapshot in memory; mutating methods (replace*,upsert*,delete*ForProvider) throw because writes are server-side only.AuthBrokerRefresheris the background refresh loop that pre-refreshes credentials withinrefreshSkewMsand disables on definitive failure (invalid_grant/ non-network 401-403).- Added
AuthStorage.exportSnapshot(),AuthStorage.upsertCredential(provider, credential),AuthStorage.forceRefreshCredentialById(id), andAuthStorage.disableCredentialById(id, cause)public methods consumed by the auth-broker server. - Added
AuthStorageOptions.refreshOAuthCredentialoverride so a remote-store client can route every OAuth refresh through the broker instead of the local OAuth endpoint. - Added
REMOTE_REFRESH_SENTINEL("__remote__") — the wire placeholder substituted for OAuth refresh tokens in broker snapshots; clients never see the real refresh token. - Exposed the OAuth provider catalog (
getOAuthProviders,OAuthProvider,OAuthProviderInfo) andrefreshOAuthTokenthrough the package barrel so the coding-agent CLI can target them without reaching intoutils/oauth. - Added the auth-gateway subsystem (
@oh-my-pi/pi-ai/auth-gateway) — a forward-proxy that sits between unauthenticated clients (the macOS usage widget, llm-git, robomp containers, …) and the broker. Clients send standard provider-format requests; the gateway parses them into omp's canonicalContext, dispatches through pi-ai'sstreamSimple(), and translates the canonical event stream back to the matching wire format.Authorizationis injected server-side so access tokens never leave the gateway host. Wire surface: GET /healthz— unauth liveness.GET /v1/usage— aggregated provider usage; 5-min per-credential cache viaAuthStorage.fetchUsageReports.GET /v1/models— model catalog (scoped to providers with credentials).POST /v1/chat/completions— OpenAI chat-completions in/out.POST /v1/messages— Anthropic messages in/out (text + thinking + tool_use blocks, SSE event taxonomy preserved).POST /v1/responses— OpenAI Responses in/out (reasoning items + function_call output items, SSE pass-through).- Added exports from
@oh-my-pi/pi-ai/auth-gateway:startAuthGateway,AuthGatewayServerOptions,AuthGatewayBootOptions,AuthGatewayServerHandle,ModelResolver,DEFAULT_AUTH_GATEWAY_BIND. Per-formatparseRequest/encodeResponse/encodeStreamtriples are reachable via the./providers/*subpath asopenai-chat-server,anthropic-messages-server, andopenai-responses-server. - Added
listProvidersWithEnvKey()to enumerate every provider with an env-var fallback (used by the new migrate command in coding-agent).
Changed
- Changed
GET /v1/snapshotto support generation-based polling withIf-None-Matchandwaitfor long-poll updates and to return304when no snapshot changes are available - Changed Bedrock credential resolution for streaming calls to prefer environment keys, AWS profile/SSO credentials, and IMDSv2 fallback when available
- Changed auth-gateway parsing for OpenAI chat-completions and Responses to ignore unsupported SDK-only fields instead of rejecting requests
- Changed auth-gateway protocol handling to include CORS headers on responses and support browser-origin requests
- Changed prompt-cache handling to resolve cache keys from request metadata and headers and preserve them through protocol translation
- Changed Anthropic messages parsing to forward request
metadatathrough to downstream execution - Changed usage report caching to use a 5-minute per-credential TTL with jittered refresh timing to reduce usage endpoint rate-limit collisions
- Changed usage polling failure handling so transient errors continue serving the last known report instead of returning null and dropping the credential from usage aggregates after cache expiry
- Changed
sanitizeSchemaForGoogleto normalize snake_case schema keys (such asany_ofandadditional_properties) to camelCase and auto-generatepropertyOrderingfor multi-property objects - Changed strict-mode sanitization to resolve
$refnodes with sibling keys by inlining and merging referenced local definitions - Changed strict-mode sanitization to flatten single-entry
allOfnodes and remove theallOfwrapper - Changed Anthropic tool schema normalization to preserve supported metadata keywords such as
$ref,$defs,$schema,enum,const,default,title, andnullableinstead of stripping them - Changed string schema processing to retain only supported
formatvalues (date-time,time,date,duration,email,hostname,uri,ipv4,ipv6,uuid) and demote unsupportedformatvalues todescriptionhints
Fixed
- Fixed OAuth credential refresh flow so concurrent manual and background refreshes now share one in-flight attempt per credential, and
RemoteAuthCredentialStorenow re-synchronizes before using near-expiring OAuth credentials - Fixed stale-credential handling after auth failures by waiting for updated broker snapshots and refreshing suspect credentials through broker endpoints before continuing
- Fixed Google Generative AI startup behavior to throw a clear API-key-required error when no key is configured
- Fixed AWS Bedrock image message serialization to preserve base64
source.bytespayloads instead of decoding and rebuilding them - Fixed Google provider error handling to extract the API-reported
error.messagefrom JSON response bodies when available - Fixed
RemoteAuthCredentialStore.getUsageReportto return the matching credential-specific usage report and coalesce parallel callers into one broker/v1/usagefetch - Fixed auth-broker credential upload validation to reject the remote refresh-token sentinel and prevent storing a non-refresh value
- Fixed OpenAI Responses streaming output to emit
reasoning_summary_textevents and parse/sendsummary_textreasoning payloads - Fixed Anthropic stop-sequence handling by trimming requests to the API limit of four entries before forwarding
- Fixed prompt caching behavior across protocol translations so cached-token usage is preserved when Anthropic and OpenAI requests are routed through each other
- Fixed Claude usage fetching to retry transient
429and5xxresponses with exponential backoff, respectingRetry-Afterbefore returning failure - Fixed auth-gateway request translation to preserve OpenAI Responses string/system message content, reasoning replay payloads, completed item text in stream item-done events, Anthropic tool-result ordering, and OpenAI Chat/Responses cached-token usage totals
- Fixed auth-gateway failure handling so unsupported request controls, upstream terminal errors, non-streaming aborts, and already-aborted client requests fail explicitly instead of being accepted, ignored, or encoded as successful HTTP 200 responses
- Fixed Gemini CLI / Antigravity tool schema normalization to run the full Cloud Code Assist pipeline, matching shared Google schema handling for union/object merging and nullable extraction
- Fixed stripped validation hints to be preserved as description spill text (
{key: value}blocks) whennormalizeSchemaForGoogleandnormalizeSchemaForCCAdrop unsupported schema keywords - Fixed
sanitizeSchemaForGoogleto collapse nullability forms (type:'null'and null-bearinganyOfvariants) intonullablewhile preserving remaining variants - Fixed
sanitizeSchemaForGoogleto inline local$defsreferences instead of dropping$ref/$defsstructure during Google schema sanitization - Fixed
normalizeAnthropicToolSchemato handle self-referential schemas without infinite recursion - Fixed object schema normalization so explicit open-map declarations (
additionalProperties: trueand schema-valuedadditionalProperties) are preserved instead of being converted to closed objects - Fixed unsupported schema constraints on arrays and strings (
maxItems,uniqueItems,pattern,minLength,maxLength, andminItemswhen greater than 1) by demoting them intodescriptionrather than dropping them
Security
- Hardened auth-gateway bearer-token checks with constant-time comparison to avoid timing-side-channel leaks
[15.1.2] - 2026-05-15
Breaking Changes
- Rejected draft-07 tuple and dependency keywords (
itemsarrays,dependencies,additionalItems) in JSON Schema validation
Added
- Added
responseHeaders,responseStatus, andresponseRequestIdfields toMockResponseso mock providers can provide syntheticProviderResponseMetadata - Added
onResponsemetadata emission for mocks that sends lowercased headers and a default status of 200 before streaming when response headers are configured - Added recursive strict-mode sanitization for array
prefixItemsentries so tuple schemas now enforce object constraints per item
Changed
- Normalized legacy draft-07 JSON Schema constructs used in tool parameters (
itemsarrays,additionalItems,definitions,dependencies) to draft 2020-12 before OpenAI/Google/CCA sanitization, wire conversion, and argument validation - Reworked OpenAI response schema adaptation to rewrite
oneOfintoanyOfwhile preserving existinganyOfbranches - Changed tuple array validation to validate per-index schemas from
prefixItemsand applyitemsonly to remaining elements
Fixed
- Fixed validation of plain JSON Schema tool arguments that omitted a
$schemaURI so draft-07-shaped schemas now pass validation instead of being rejected - Fixed tuple-array validation for legacy JSON Schema tool schemas to enforce
additionalItems: falseand per-position constraints after automatic draft upgrade - Fixed Anthropic tool schema normalization to recurse into
prefixItemsso unsupported constraints inside tuple items are stripped in the generated input schema - Fixed Anthropic tool-schema normalization stripping the body of explicit open
additionalProperties(e.g. Zod'sz.record(z.string(), z.unknown())compiling toadditionalProperties: {}) by unconditionally overwriting it withfalse, which closed record-style fields and prevented models from supplying any key. The coding-agent'sresolvetool exposes plan-approval titles via such a field, so Kimi K2 (and any other Anthropic-shaped provider) could not passextra: { title }, blocking plan mode entirely (#1104) - Fixed Anthropic strict tool planning to leave tools with open
additionalPropertiesmaps non-strict instead of sending schemas Anthropic rejects.
[15.1.0] - 2026-05-15
Breaking Changes
- Removed TypeBox root exports (
Type,Static, andTSchema) from the package entrypoint, so callers importing those symbols from@oh-my-pi/pi-aimust migrate tozodor@oh-my-pi/pi-ai/types
Added
- Added support for defining tool schemas with Zod (
z.object,z.string, etc.) by allowingTool.parametersto be either Zod schemas or legacy JSON Schema objects and converting them to provider wire format automatically - Added package-level schema helpers in the
zod/v4style by exportingzandZodTypefrom the root entrypoint - Added a
mockAPI provider viacreateMockModelto buildModel<"mock">instances for fully in-memory, deterministic assistant streams in tests - Added
streamMockandregisterMockApiso mock responses can be consumed throughstream()and the global custom API registry without an external model backend - Added async/sync response scripting with optional context-based handlers, and new
push()/reset()controls to drive multi-turn mock interactions and inspect per-call invocation state - Added support in mock responses for simulating tool calls, usage metadata, custom stop reasons, delayed emissions, and terminal error/aborted outcomes
Changed
- Changed Azure OpenAI Responses tool schema conversion to sanitize tool parameter schemas and rewrite
oneOfbranches asanyOfso tool calls remain compatible with Azure's schema expectations - Changed
Static<S>to extract a schema object’sstatictype when present, improving inferred tool argument types for non-Zod parameter definitions - Changed
Statictyping behavior so it now infers argument types from Zod schemas and defaults tounknownfor non-Zod JSON Schema parameter definitions - Restored the default steady-state stream idle timeout to 120s (regressed in 15.0.0). 30s was too aggressive for reasoning models, slow proxies, and tool-call planning gaps, surfacing as repeated
Provider stream stalled while waiting for the next eventerrors. ExistingPI_STREAM_IDLE_TIMEOUT_MS/PI_OPENAI_STREAM_IDLE_TIMEOUT_MSoverrides are unchanged.
Fixed
- Preserved top-level unknown fields in validated tool-call arguments so extra root properties are retained after schema coercion
- Fixed coercion for Zod
recordfields by parsing JSON-stringified record arguments into objects - Validated legacy draft-07 JSON Schema tool parameters directly instead of converting through Zod, improving support for features like
$ref,definitions,nullable, anduniqueItems - Fixed Cloud Code Assist schema preparation to strip unsupported
propertyNamesand fall back to a minimal tool schema when schema meta-validation detects malformed keywords - Fixed OpenAI Completions streaming to avoid treating non-output chunks (including role-only preambles) as progress events so idle-timeout watchdog behavior no longer hangs on no-op streamed chunks
- Fixed Cloud Code Assist schema compatibility checks by replacing strict AJV meta-schema validation with structural JSON Schema validation to avoid rejecting structurally valid tool schemas
- Fixed lazy built-in provider streams (
anthropic-messages,bedrock-converse-stream,cursor-agent,google-*,ollama-chat,openai-*) prematurely aborting slow first-token responses withProvider stream stalled while waiting for the next event. The lazy-stream watchdog wrapper was treating the syntheticstartevent (yielded immediately by every provider before the model emits any tokens) as the first real item, which caused the watchdog to drop fromfirstItemTimeoutMs(100s) toidleTimeoutMs(30s) before the upstream model had produced anything. The sharediterateWithIdleTimeoutnow keepsawaitingFirstItemtrue until a real progress item arrives, and the lazy-stream wrapper marksstartas a non-progress keepalive (#1073 regression). - Heal leaked Kimi K2 chat-template tool-call tokens (
<|tool_calls_section_begin|>…<|tool_call_argument_begin|>…<|tool_calls_section_end|>) that some hosts (nativekimi-codeAPI, OpenRouter, Fireworks, etc.) emit intodelta.contentinstead of structuredtool_calls. The OpenAI-completions stream consumer now strips the markers from visible text, reconstructs the embedded calls as propertoolCallcontent blocks (stream-aware, token-boundary-safe), and promotesfinish_reason: stoptotoolUsewhen calls were healed. - Fixed OpenAI-completions Kimi K2 healed-call promotion clobbering non-stop terminal finish reasons (
error,length,aborted); promotion now only fires when the prior stop reason is the natural-completionstop - Fixed OpenAI-completions duplicate Kimi tool calls when a single chunk delivers both leaked markers and a structured
delta.tool_calls; the healer now strips visible markers but discards its synthesized calls so structured payloads remain the single source of truth - Fixed Kimi tool-call healer synthesizing a bogus empty call when assistant text mentions a literal
<|tool_call_end|>(or<|tool_call_begin|>/<|tool_call_argument_begin|>) outside an active<|tool_calls_section_begin|>…<|tool_calls_section_end|>section; the tokens now survive as text - Fixed OpenAI-completions ignoring per-request
StreamOptions.streamFirstEventTimeoutMswhen configuring the underlying OpenAI SDK HTTP timeout, causing slow-before-headers providers to be aborted at the env default before the wrapping watchdog armed - Fixed JSON Schema validator silently accepting values that violate
propertyNames,patternProperties,dependentRequired,dependencies,if/then/else,contains, andprefixItems; the in-tree validator now enforces these keywords instead of falling through.unevaluatedProperties/unevaluatedItemsremain permissive but log a one-time warning so tool authors are not surprised. - Fixed recursive
$refschemas being treated as universally valid: the validator previously short-circuited on the second occurrence of any ref it had already seen, so nested values violating the referenced sub-schema passed. Cycle detection now keys on (ref, value-identity) pairs with a depth cap for primitive values, so genuine sub-tree violations are still caught. - Fixed JSON Schema meta-validator accepting malformed
if/then/elseanddependencieskeywords; each conditional sub-schema is now structurally validated and draft-07dependenciesaccepts either a schema or a string array of dependent keys. - Fixed Zod-emitted wire schemas dropping null-valued unknown root fields before
preserveUnknownRootFieldscould snapshot them, so callers liketask.simpleno longer lose aschema: nullargument and downstream rejection paths fire as intended. - Fixed mock provider partial
Usageto recomputetotalTokens(andcost.totalwhen cost components are supplied) when omitted, instead of reporting 0 - Fixed mock provider auto-generated tool-call IDs to use a per-instance counter (now reset by
reset()), so test order no longer affects IDs acrosscreateMockModel()instances
[15.0.2] - 2026-05-15
Fixed
- Fixed
StreamOptions.fetchtyping to accept fetch-compatible override functions that do not exposepreconnect, allowing custom fetch implementations to be used without type errors across runtimes - Fixed Moonshot Kimi K2.6 forced tool calls to send
thinking: { type: "disabled" }, avoidingtool_choice 'specified' is incompatible with thinking enabled400s while preserving the requested named tool (#1077).
[15.0.1] - 2026-05-14
Breaking Changes
- Increased the minimum Bun runtime version to
>=1.3.14for the@aws-?package
Added
- Added
installH2Fetchto patchglobalThis.fetchso HTTPS requests attempt HTTP/2 over ALPN with automatic HTTP/1.1 fallback when HTTP/2 is unsupported - Added priority service-tier traffic to the
premiumRequestsaccounting on OpenAI and OpenAI Codex providers. SendingserviceTier: "priority"now incrementsusage.premiumRequestsby 1 per request, matching the existing GitHub Copilot premium-request budget semantics so downstream consumers (e.g. theomp stats"Premium Reqs" card and/usage) reflect priority traffic alongside Copilot premium calls.
[15.0.0] - 2026-05-13
Added
- Added
AuthStorage.onCredentialDisabled(listener)— a multi-subscriberon/offAPI forcredential_disabledevents. Returns an unsubscribe function; calling it more than once is a no-op. Multiple subscribers all receive every disable event, with synchronous and async exceptions isolated per-listener so a misbehaving subscriber cannot starve the rest of the chain. Buffer-and-replay semantics are preserved: events emitted while no listener is subscribed are buffered (FIFO, capped at 32) and replayed once to the listener that triggers the empty→non-empty transition. After every subscriber unsubscribes, subsequent disable events buffer again until the next subscribe.
Fixed
- Fixed OAuth credentials being silently disabled when two omp processes (or any two
AuthStorageinstances sharing aagent.db) race on token refresh. Anthropic rotates refresh tokens on every use, so the loser'sinvalid_grantresponse previously soft-deleted the row that the winner just rotated, forcing the user to/loginagain.#tryOAuthCredentialnow re-reads the row from disk before declaring a definitive failure: if the persistedrefreshdiffers from the snapshot it tried, the peer-rotated credential is reloaded and the request retries against the fresh token instead of disabling the live row. - Closed a remaining race window in OAuth refresh-failure handling: between re-reading the credential row to check for peer rotation and the subsequent soft-delete, another process could still complete a refresh and rotate the row, leaving us to disable the freshly-rotated credential by
id. The disable now runs as a single CAS update conditioned on the row'sdatastill matching the snapshot we tried to refresh, and ondisabled_cause IS NULL. If the CAS reports 0 rows changed (peer rotation, or row already disabled by a concurrent failure on the same snapshot), we reload from disk and retry instead of mutating the wrong row or emitting a spuriouscredential_disabledevent.
[14.9.3] - 2026-05-10
Fixed
- Anthropic provider now retries generic transient connect failures (
unable to connect,fetch failed,connection error, etc.) by falling back to the sharedisRetryableErrorallowlist after the provider-specific patterns. Previously these errors bypassed the hand-curated regex inisProviderRetryableErrorand aborted the stream on the first attempt, while the OpenAI SDK and CodexfetchWithRetrypaths already handled them.
[14.9.0] - 2026-05-10
Fixed
- Fixed silent forwarding of image content (for example Python plot output rendered in the terminal) to models without vision support, which produced opaque 404 errors from upstream. Image blocks are now stripped and replaced with a
[image omitted: model does not support vision]placeholder for non-vision models, including tool-result payloads (#967, #968). - Added
AuthStorageonCredentialDisabledcallback (sync or async) so embedders can react when a credential is automatically disabled (e.g. OAuth refresh fails withinvalid_grant) — useful for surfacing a banner or auto-launching a re-login flow instead of letting the credential silently disappear. Sync throws and async rejections are both caught and logged so a misbehaving subscriber cannot break the disable path. - Added Anthropic OAuth
account.uuidandaccount.email_addressextraction from the/v1/oauth/tokenexchange and refresh responses; bothAnthropicOAuthFlow.exchangeToken()andrefreshAnthropicToken()now populateOAuthCredentials.{accountId, email}so downstream consumers can attribute requests to the authenticated account without a separate/api/oauth/profileround-trip. - Added
onSseEventstream diagnostics so HTTP SSE providers can expose raw SSE frames without changing parsed model output. - Added
streamIdleTimeoutMsoption (andPI_STREAM_IDLE_TIMEOUT_MSenv override;PI_OPENAI_STREAM_IDLE_TIMEOUT_MSremains a backward-compatible alias) for a steady-state inter-event watchdog. Set to0to disable. - Added a semantic-progress predicate to OpenAI Responses and Codex SSE/WebSocket transports so
response.in_progress-style keepalives no longer reset the idle deadline on stalled tool calls.
Changed
- Anthropic streams now enforce a steady-state idle timeout (defaults to 120s, same control as
PI_STREAM_IDLE_TIMEOUT_MS) in addition to the first-event watchdog. Long-running responses that go fully silent between events will now surface asAnthropic stream stalled while waiting for the next eventinstead of hanging. - Fixed
resolveAnthropicMetadataUserId()to accept JSON-formatuser_idvalues that match real Claude Code's payload shape ({ device_id, account_uuid, session_id, ... }fromservices/api/claude.ts:getAPIMetadata). Previously only the syntheticuser_<hex>_account_<uuid>_session_<uuid>cloaking format was accepted on OAuth, which caused stable session-keyed metadata supplied by callers to be discarded and replaced with fresh random entropy on every request — defeating session-count attribution on the Claude OAuth path.
[14.8.0] - 2026-05-09
Fixed
- Fixed Gemini 3 Pro thinking metadata so
mediumeffort is rejected with the expected error instead of being silently accepted:ThinkingConfignow carries an optional explicitlevelslist that survivesexpandEffortRange, letting non-contiguous supported sets (e.g.[low, high]) round-trip through enrichment. - Fixed Kimi Code OAuth expiry handling to refresh access tokens 5 minutes before server expiry, avoiding daily 401s from using tokens right up to the cutoff.
[14.7.6] - 2026-05-07
Added
- Added
hideThinkingSummaryoption toSimpleStreamOptions. When true,streamSimplerequests that the underlying provider omit reasoning/thinking summaries: Anthropic receivesthinking.display = "omitted"(where supported), and OpenAI Responses / Azure / Codex providers leavereasoning.summaryunset so the server skips emitting the human-readable summary stream entirely.
Changed
- Changed OpenAI Responses, Azure OpenAI Responses, and OpenAI Codex providers to omit
reasoning.summaryfrom requests whenreasoningSummaryis explicitlynull(previously fell back to"auto").
[14.7.5] - 2026-05-07
Added
- Added
OpenAICompat.supportsMultipleSystemMessagesso chat-completions hosts can opt out of separate leading system blocks. Auto-detected astruefor OpenAI, Azure, OpenRouter, Cerebras, Together, Fireworks, Groq, DeepSeek, Mistral, xAI, Z.ai, GitHub Copilot, and Zenmux;falsefor MiniMax, Alibaba Dashscope, and Qwen Portal whose chat templates reject follow-up system messages. Unknown OpenAI-compatible hosts (custom vLLM/local) default tofalse; users can opt back in viacompat.supportsMultipleSystemMessages: true.
Fixed
- Fixed strict-template OpenAI-compatible hosts (e.g. Qwen 3.5+ via vLLM, MiniMax) rejecting follow-up
system/developermessages by coalescing ordered system prompts into a single block joined by\n\nwhencompat.supportsMultipleSystemMessagesis false. Canonical hosts continue to receive separate blocks so KV-cache reuse stays effective when only the trailing prompt changes (#958).
[14.7.2] - 2026-05-06
Fixed
- Fixed VLLM model discovery to use
max_model_lenas the context window when the endpoint reports it. - Fixed custom Ollama Cloud/local-proxy model aliases (for example
deepseek-v4-pro:cloud) to inherit bundled cache-pricing metadata when the upstream model is known (#937). - Fixed local Ollama model discovery to apply
/api/showthinking and vision capabilities in addition to native context windows (#928).
[14.7.0] - 2026-05-04
Breaking Changes
- Changed
Context.systemPromptfrom a string tostring[], so callers must now pass an array of prompts instead of a single string - Changed behavior will throw at runtime for non-array system prompts because request builders now normalize system prompts as an array
Added
- Added support for multiple system prompts by changing
Context.systemPromptto an ordered string array and preserving provider-appropriate instruction precedence
Changed
- Changed request builders for Anthropic, OpenAI, Bedrock, Azure, Cursor, Google, and Ollama to propagate every non-empty system prompt entry without demoting durable instructions into ordinary conversation turns
Fixed
- Filtered out empty normalized system prompts so blank entries are no longer sent to providers
- Removed blank system prompt strings from provider payloads to avoid unnecessary empty instruction messages
[14.6.6] - 2026-05-04
Added
- Added always-on OpenRouter response caching (1h TTL) by sending
X-OpenRouter-Cache: trueandX-OpenRouter-Cache-TTL: 3600on every OpenRouter request — identical requests replay from OpenRouter's edge cache for free. https://openrouter.ai/docs/features/response-caching
[14.6.4] - 2026-05-03
Fixed
- Fixed OpenAI Codex websocket continuations to retry with full context when
previous_response_idexpires server-side instead of surfacingprevious_response_not_found.
[14.6.2] - 2026-05-03
Added
- Added
EventStream.fail(err)method to terminate the async iterator with an error, enabling consumers to catch stream-level failures viafor awaitwithout hanging
Fixed
- Fixed OpenAI Responses tool schema conversion to rewrite non-strict
oneOfunions toanyOfbefore sending tools to the Responses API (#920)
[14.6.0] - 2026-05-02
Added
- Added
disableReasoningto stream and OpenAI completion options to force reasoning off for models that support it, sendingreasoning: { enabled: false }for OpenRouter-compatible requests - Added
thinkingDisplayoption to Anthropic options to control whether adaptive and explicit reasoning is returned assummarizedoromitted - Added Anthropic model compatibility flags
supportsEagerToolInputStreamingandsupportsLongCacheRetentionfor API-capability-specific request behavior
Changed
- Changed Anthropic request payloads to send
thinking: { type: "disabled" }whenthinkingEnabledis explicitlyfalseon reasoning-enabled models - Changed Anthropic cache retention handling so
cacheRetention: "long"now usesttl: "1h"only for canonical Anthropic endpoints with long-cache support - Changed Anthropic tool schema generation to include
eager_input_streamingonly on models that advertise support - Changed Anthropic OAuth login flow to include browser fallback guidance and richer error context when token exchange or refresh fails
Fixed
- Fixed Anthropic non-thinking requests to include the caller-provided
temperaturevalue in request payloads - Fixed Anthropic
claude-opus-4-7non-thinking payloads to omit sampling fields (temperature,top_p, andtop_k) - Fixed OpenAI Codex base URL normalization so configured base URLs with or without
/codexor/codex/responsesnow resolve to/codex/responses - Fixed OpenAI Codex websocket handling to parse JSON from non-string message payloads including
ArrayBuffer, typed arrays, andBlobvalues - Fixed OpenAI Codex websocket handshakes to replace stale
openai-betavalues with the websocket beta and avoid sending request-body headers over websocket transport - Fixed abort tracking so caller-initiated cancellations are treated as user aborts even after local watchdog timeouts, preventing unintended automatic retries
- Fixed Anthropic stream handling to parse raw SSE envelopes directly, ignore unrelated events, and repair malformed JSON in SSE payloads
- Fixed Anthropic streaming to emit an explicit error when the SSE stream ends without a
message_stopevent - Fixed OpenAI Codex websocket continuations to send true
previous_response_iddeltas forstore: falsetranscripts, expose request stats, and default text verbosity tolowunless explicitly overridden. - Fixed OpenAI Codex websocket append reuse after
response.completedterminal events.
[14.5.14] - 2026-05-01
Added
- Added package-level
google-gemini-headersexports (getGeminiCliHeaders,getGeminiCliUserAgent,getAntigravityHeaders,extractRetryDelay, andANTIGRAVITY_SYSTEM_INSTRUCTION) for header and retry handling reuse without importing full Google providers
Changed
- Changed package exports and streaming/provider wiring to load heavy Google/Kimi/GitLab/synthetic provider modules lazily through
register-builtins, reducing startup import overhead from optional provider SDKs
Fixed
- Fixed DeepSeek V4 tool-call follow-up 400 errors from three root causes:
- Mapped
reasoning_effort"xhigh" to "max" for DeepSeek-family models on any provider (NVIDIA, OpenCode-Go, etc.), not justdeepseek - Recovered
reasoning_contentfrom thinking blocks with valid signatures that were filtered by the non-empty-text check
- Mapped
- Added empty-string fallback when
reasoning_contentis genuinely absent (e.g. proxy-stripped) but the provider requires the field
[14.5.13] - 2026-05-01
Breaking Changes
- Removed
utils/oauthre-exports from the package entrypoint, so OAuth helper imports from the root module must be updated
[14.5.10] - 2026-04-30
Added
- Added provider response metadata callbacks for Anthropic and OpenAI streaming requests.
[14.5.9] - 2026-04-30
Added
- Added
usage.reasoningTokensto OpenAI and Google usage output when providers report reasoning/thinking tokens - Added
usage.cttl.ephemeral5mandusage.cttl.ephemeral1hto report Anthropic cache-write TTL token buckets - Added
usage.server.webSearchandusage.server.webFetchto report Anthropic server tool-call request counts
Fixed
- Fixed OpenAI usage attribution to avoid double-counting
reasoning_tokensin output totals - Fixed Anthropic streaming usage handling so a previously populated cache TTL breakdown is preserved when later events omit
cache_creation
[14.5.4] - 2026-04-28
Changed
- Changed OpenAI custom Lark grammar payloads to strip comments and blank lines before sending provider requests.
Fixed
- Fixed OpenAI Codex GPT model pricing by inheriting matching OpenAI catalog rates for zero-priced discovered Codex entries.
[14.5.3] - 2026-04-27
Added
- Added
fireworksas a supported provider with API key login flow and credential storage - Added Fireworks model catalog support with
fireworks-scoped openai-completions modelsglm-5,glm-5.1,kimi-k2.5,kimi-k2.6, andminimax-m2.7 - Added built-in discovery wiring so providers with base URL
api.fireworks.aiare recognized as OpenAI-compatible and can use streaming token control
Changed
- Updated the built-in model catalog to use corrected
contextWindowandmaxTokensvalues for many existing models instead of placeholder limits - Updated several model cost entries, including cache-read pricing, to corrected values
Fixed
- Fixed Fireworks request formatting by translating between public model IDs and API wire IDs when sending OpenAI-completions requests
- Fixed OpenAI-compatible model parameter handling for Fireworks by allowing
max_tokensto be sent during requests
[14.5.1] - 2026-04-26
Fixed
- Fixed NVIDIA NIM DeepSeek-V4 models leaking chat-template tool-call markers (e.g.
<|DSML|tool_calls|>) into visible response text by stripping the special tokens from streameddelta.content(#798)
[14.4.0] - 2026-04-26
Added
- Added an
examplesoption toStringEnumto include example values in the generated schema
Changed
- Changed Anthropic tool schema generation to strip unsupported schema fields (including
patternProperties), addadditionalProperties: falsefor object types, and apply Anthropic strict-mode limits when marking tools as strict - Changed Anthropic strict tool planning to cap strict
toolsat twenty entries and convert excess optional/union parameters to nullable schemas to stay within provider constraints
Fixed
- Fixed Anthropic tool schema compilation failures by keeping the
writetool out of the strict-tool allowlist when the full coding-agent tool set is active - Fixed Anthropic 400
tools.*.custom: For 'object' type, property 'minItems' is not supportedby strippingminItemsfrom object-shaped JSON schema nodes (array nodes still keep supportedminItemsvalues) - Fixed Anthropic tool schemas that used tuple-style arrays by stripping unsupported
maxItemsand only preserving provider-supportedminItemsvalues - Fixed Anthropic and OpenRouter Anthropic tool calls that previously failed with
compiled grammar is too largeby retrying automatically without strict tool schemas and reusing non-strict mode for subsequent requests in the same provider session - Fixed parsing of JSON tool arguments containing raw control characters inside string values (such as embedded newlines) by escaping them before JSON parsing
- Fixed
validateToolArgumentsto accept stringified objects and arrays that include literal control characters inside string fields - Fixed OpenAI Codex Spark OAuth selection to fall back to non-Pro accounts when no ChatGPT Pro account is connected, so users without a Pro account can still attempt Spark requests in case the server permits access.
[14.3.0] - 2026-04-25
Added
- Added support for Claude Opus 4.7 (
claude-opus-4-7) model (#726)- Suppresses sampling parameters (temperature/top_p/top_k) that Opus 4.7 rejects
- Enables
display: "summarized"for adaptive thinking to restore visible thinking content
Fixed
- Fixed Cursor provider losing conversation history on follow-up turns (model responding "this appears to be the start of our session") by populating
ConversationStateStructure.rootPromptMessagesJsonwith JSON blob IDs for the system prompt plus prior user/assistant/tool-result messages. Cursor's server builds the model prompt fromrootPromptMessagesJson, not from the protobufturns[]tree, so sending only the system prompt there caused prior turns to be dropped - Fixed Cursor provider multi-turn conversations failing with
Connect error internal: Blob not foundon the second message by storingConversationStateStructure.turns,AgentConversationTurnStructure.user_message, andAgentConversationTurnStructure.stepsas content-addressed blob IDs in the KV store (matching the existing handling forrootPromptMessagesJson) rather than sending the raw serialized bytes inline (#678)
[14.2.1] - 2026-04-24
Fixed
- Fixed OpenAI Codex Spark OAuth selection to require a verified ChatGPT Pro account instead of falling back to Plus or unknown-plan accounts.
[14.2.0] - 2026-04-23
Added
- Added
gpt-5.5to the built-in model catalog for both OpenAI Responses (openai) and locallitellm(openai-completions) providers - Added
gpt-image-2to thelitellmbuilt-in model catalog - Added
isCopilotTransientModelError()andcallWithCopilotModelRetry()helpers inutils/retrythat detect GitHub Copilot's intermittentHTTP 400 model_not_supportedresponses for preview models (gpt-5.3-codex,gpt-5.4,gpt-5.4-mini, ...) and retry the request up to three times with backoff. OpenAI Responses, OpenAI Completions, and Anthropic provider paths now participate in this retry when the model is served through Copilot. - Added OpenAI Responses custom-tool grammar support for Codex-style
apply_patchcalls, including freeform streaming, history replay, and forced tool-choice mapping to the custom wire name.
Changed
- Updated built-in model metadata with revised
contextWindow,maxTokens, and pricing values for existing entries - Changed generated model policies to assign
applyPatchToolType: "freeform"for first-party GPT-5 OpenAI Responses and Codex models, so regeneratedmodels.jsonpreserves theapply_patchcustom-tool metadata. - Renamed
rewriteCopilotAuthErrortorewriteCopilotErrorand extended it to rewriteHTTP 400 model_not_supportedafter retries are exhausted with guidance about Copilot's OAuth-client-specific rollout gap (see opencode#13313).
Fixed
- Fixed Amazon Bedrock proxy handling to honor lowercase
http_proxy,https_proxy, andall_proxyenvironment variables when using HTTP/1 fallback - Fixed Amazon Bedrock streaming behind corporate HTTP proxies by using a proxy-aware HTTP/1 transport when
HTTPS_PROXY,HTTP_PROXY, orALL_PROXYis configured, including AWS SSO credential calls. - Fixed Amazon Bedrock requests to retry once with HTTP/1 when the AWS SDK's default HTTP/2 transport fails before streaming begins.
- Fixed OpenAI Responses streaming to display thinking tokens from local providers (llama.cpp, etc.) that send raw
reasoning_text.deltaevents and emptysummaryarrays inoutput_item.done. Previously, thinking content was silently dropped during streaming while non-streaming mode worked correctly. - Synced the bundled OpenCode Go catalog with the current docs so
kimi-k2.6,mimo-v2.5, andmimo-v2.5-proappear in offline/default model lists.
[14.1.3] - 2026-04-17
Fixed
- Preserved user-provided
session_idandx-client-request-idheaders in OpenAI Responses requests instead of overriding them with automatic session-derived values - Stopped sending
session_idandx-client-request-idheaders for OpenAI Responses requests whencacheRetentionis set tonone - Fixed direct OpenAI Responses requests to send
session_idandx-client-request-idfrom the same session-derived value asprompt_cache_key, improving prompt cache affinity for append-only sessions
[14.1.1] - 2026-04-14
Added
- Added
toolStrictModecompatibility option ("all_strict"or"none") to OpenAI-compatible model config to force tool schemas to be sent uniformly strict, uniformly non-strict, or keep mixed per-tool behavior
Changed
- Changed Cerebras OpenAI-compatible providers to default
toolStrictModeto"all_strict"unless explicitly overridden
Fixed
- Fixed OpenAI Completions handling for providers that reject mixed
strictflags by automatically retrying with non-strict tool schemas when an initial all-strict tool request fails with strict-format 400/422 errors - Fixed OpenAI-completions error reporting by including captured JSON error body details such as type, param, and code when a request fails without a body in the thrown SDK error
- Fixed shell execution failure responses to preserve all result fields when sanitizing, preventing truncated metadata in stream results
- Fixed context overflow detection to recognize
model_context_window_exceededfrom z.ai / GLM providers, preventing infinite retry loops when context window is exceeded (#638) - Fixed strict tool schema enforcement to preserve
additionalProperties: falseand required keys for reused nested object schemas, preventing invalidtodo_writefunction schemas in Codex/OpenAI requests
[14.1.0] - 2026-04-11
Added
- Added
accountIdto usage report metadata
Changed
- Changed usage parsing to emit a usage report with available fields when parsing fails, rather than returning null
Fixed
- Fixed
planTyperesolution to fall back to the raw payloadplan_typewhen parsed value is absent - Fixed usage metadata
rawfallback to preserve the original payload when parsed raw output is missing
[14.0.5] - 2026-04-11
Changed
- Replaced GitHub Copilot authentication from VSCode extension impersonation to the opencode OAuth flow, eliminating TOS concerns. Existing users will need to re-authenticate once with
/login github-copilot. - Simplified Copilot token handling: GitHub OAuth token is used directly for all API requests (no JWT exchange or refresh cycle).
- Changed GitHub Copilot API base URL from
api.individual.githubcopilot.comtoapi.githubcopilot.com. - Updated default OpenAI stream idle timeout to 120,000 milliseconds to keep stream generation alive longer
Fixed
- Fixed duplicate synthetic tool results being generated when a real tool result appears later in message history
- Fixed GitHub Copilot
/modelsdiscovery to unwrap structured OAuth credentials before sending the bearer token, preserving dynamic catalog refresh for OAuth-backed callers.
Removed
- Removed Copilot JWT proxy-ep base URL resolution (no longer needed with opencode auth).
[14.0.3] - 2026-04-09
Fixed
- Fixed Ollama discovery cache normalization so cached models upgrade to the OpenAI Responses transport after the provider change
[14.0.0] - 2026-04-08
Breaking Changes
- Removed
coerceNullStringsfunction and its automatic null-string coercion behavior from JSON parsing
Added
- Added support for OpenRouter provider with strict mode detection
- Added automatic cleaning of literal escape sequences (
\n,\t,\r) in JSON parsing to handle LLM encoding confusion - Added support for healing JSON with trailing junk after balanced containers (e.g.,
]\n</invoke>) - Added
CODEX_STARTUP_EVENT_CHANNELconstant andCodexStartupEventtype for monitoring Codex provider initialization status - Added automatic healing of malformed JSON with single-character bracket errors at the end of strings, improving LLM tool argument parsing robustness
[13.19.0] - 2026-04-05
Fixed
- Fixed GitHub Copilot model context window detection by correcting fallback priority for maxContextWindowTokens and maxPromptTokens
- Fixed Gemini 2.5 Pro context window detection in GitHub Copilot model limits test
- Fixed Claude Opus 4.6 context window detection in GitHub Copilot model limits test
- Fixed Anthropic streaming to suppress transient SDK console errors for malformed SSE keep-alive frames so the TUI only shows surfaced provider errors
- Added environment-based credential fallback for the OpenAI Codex provider.
[13.17.6] - 2026-04-01
Fixed
- Fixed Anthropic first-event timeouts to exclude stream connection setup from the watchdog, preserve timeout-specific retry classification after local aborts, and reset retry state cleanly between attempts
[13.17.5] - 2026-04-01
Changed
- Increased default first-event timeout from 15s to 45s to better accommodate longer request setup times
- Modified first-event watchdog to inherit idle timeout when it exceeds the default, ensuring consistent timeout behavior across different configurations
Fixed
- Fixed first-event watchdog initialization timing so it no longer starts before the actual stream request is created, preventing premature timeouts during request setup
- Fixed first-event watchdog timing so OpenAI-family providers no longer count slow request setup against the first streamed event timeout, and raised the default first-event timeout to avoid false aborts after long tool turns
[13.17.2] - 2026-04-01
Fixed
- Fixed OpenAI-family first-event timeouts to preserve provider-specific timeout errors for retry classification instead of flattening them to generic aborts (#591)
[13.17.1] - 2026-04-01
Added
- Added
thinkingSignaturefield to thinking content blocks to preserve the original reasoning field name (e.g.,reasoning_text,reasoning_content) for accurate follow-up requests - Added first-event timeout detection for streaming responses to abort stuck requests before user-visible content arrives
- Added
PI_STREAM_FIRST_EVENT_TIMEOUT_MSenvironment variable to configure first-event timeout (defaults to 15 seconds or idle timeout, whichever is lower) - Added Vercel AI Gateway to
/loginproviders for interactive API key setup
Changed
- Changed thinking block handling to track and distinguish between different reasoning field types, enabling proper field name preservation across multiple turns
Fixed
- Fixed Anthropic stream timeout errors to be properly retried by recognizing first-event timeout messages
- Fixed stream stall detection to distinguish between first-event timeouts and idle timeouts, enabling faster recovery for stuck connections
- Fixed
omp commitfailing with HTTP 400 errors when using reasoning-enabled models on OpenAI-compatible endpoints that don't support thedeveloperrole (e.g., GitHub Copilot, custom proxies). Now falls back tosystemrole whendeveloperis unsupported.
[13.17.0] - 2026-03-30
Changed
- Bumped zai provider default model from glm-4.6 to glm-5.1
[13.16.5] - 2026-03-29
Added
- Added Gemma 3 27B model support for Google Generative AI
Changed
- Updated Kwaipilot KAT-Coder-Pro V2 model display name and pricing information
- Updated Kwaipilot KAT-Coder-Pro V2 context window from 222,222 to 256,000 tokens and max tokens from 8,888 to 80,000
Fixed
- Fixed normalizeAnthropicBaseUrl returning empty string instead of undefined when baseUrl is empty
[13.16.4] - 2026-03-28
Added
- Added support for Groq Compound and Compound Mini models with extended context window (131K tokens) and configurable thinking levels
- Added support for OpenAI GPT-OSS-Safeguard-20B model with reasoning capabilities across multiple providers
- Added support for Kwaipilot KAT-Coder-Pro V2 model across Kilo, NanoGPT, and OpenRouter providers
- Added support for GLM-5.1 model with extended context window (200K tokens) and max output of 131K tokens
- Added support for Qwen3.5-27B-Musica-v1 model
- Added support for zai-org/glm-5.1 model with reasoning capabilities
- Added support for Sapiens AI Agnes-1.5-Lite model with multimodal input (text and image) and reasoning
- Added support for Venice openai-gpt-54-mini model
Changed
- Updated Qwen QwQ 32B max tokens from 16,384 to 40,960 across multiple providers
- Updated OpenAI GPT-OSS-Safeguard-20B model name to 'Safety GPT OSS 20B' and enabled reasoning capabilities
- Updated OpenAI GPT-OSS-Safeguard-20B context window from 222,222 to 131,072 tokens and max tokens from 8,888 to 65,536
- Updated OpenRouter Qwen QwQ 32B pricing: input from 0.2 to 0.19, output from 1.17 to 1.15, cache read from 0.1 to 0.095
- Updated OpenRouter Claude 3.5 Sonnet pricing: input from 0.45 to 0.42, cache read from 0.225 to 0.21
[13.16.3] - 2026-03-28
Changed
- Modified OAuth credential saving to preserve unrelated identities instead of replacing all credentials for a provider
- Updated credential identity resolution to use provider context for more accurate email deduplication
Fixed
- Fixed OAuth credential updates to replace matching credentials in-place rather than creating disabled rows, preventing unbounded accumulation of soft-deleted credentials
[13.15.0] - 2026-03-23
Added
- Added
isUsageLimitError()torate-limit-utilsas a single source of truth for detecting usage/quota limit errors across all providers
Fixed
- Fixed lazy stream forwarding to properly handle final results from source streams with
result()methods - Fixed lazy stream error handling to convert iterator failures into terminal error results instead of silently failing
- Fixed
parseRateLimitReasonto recognize "usage limit" in error messages and correctly classify them asQUOTA_EXHAUSTED - Fixed Codex
fetchWithRetryretrying 429 responses forusage_limit_reachederrors for up to 5 minutes instead of returning immediately for credential switching - Removed
usage.?limitfromTRANSIENT_MESSAGE_PATTERNin retry utils since usage limits are not transient and require credential rotation - Fixed
parseRateLimitReasonnot recognizing "usage limit" in Codex error messages, causing incorrect fallback toUNKNOWNclassification instead ofQUOTA_EXHAUSTED
[13.14.2] - 2026-03-21
Changed
- Updated thinking configuration format from
levelsarray tominLevelandmaxLevelproperties for improved clarity - Corrected context window from 400000 to 272000 tokens for GPT-5.4 mini and nano variants on Codex transport
- Normalized GPT-5.4 variant priority handling to use parsed variant instead of special-casing raw model IDs
- Added support for
minivariant in OpenAI model parsing regex
Fixed
- Fixed inconsistent thinking level configuration across multiple model definitions
[13.14.0] - 2026-03-20
Fixed
- Fixed resumed OpenAI Responses sessions to avoid replaying stale same-provider native history on the first follow-up after process restart (#488)
Added
- Added bundled GPT-5.4 mini model metadata for OpenAI, OpenAI Codex, and GitHub Copilot, including low-to-xhigh thinking support and GitHub Copilot premium multiplier metadata
- Added bundled GPT-5.4 nano model metadata for OpenAI and OpenAI Codex, including low-to-xhigh thinking support
[13.13.2] - 2026-03-18
Changed
- Modified tool result handling for aborted assistant messages to preserve existing tool results when already recorded, instead of always replacing them with synthetic 'aborted' results
[13.13.0] - 2026-03-18
Changed
- Changed tool argument validation to always normalize optional null values before type coercion, ensuring consistent handling of LLM-generated 'null' strings
Fixed
- Fixed tool argument validation to properly handle string 'null' values from LLMs on optional fields by stripping them during normalization
- Improved type safety of
validateToolCallandvalidateToolArgumentsfunctions by returning properly typedToolCall["arguments"]instead ofany
[13.12.9] - 2026-03-17
Changed
- Extracted OpenAI compatibility detection and resolution logic into dedicated
openai-completions-compatmodule for improved maintainability and reusability
Fixed
- Fixed
openai-responsesmanual history replay to strip replay-only item IDs and preserve normalized toolcall_idvalues for GitHub Copilot follow-up turns (#457)
[13.12.0] - 2026-03-14
Added
- Added support for
qwen-chat-templatethinking format to enable reasoning viachat_template_kwargs.enable_thinking - Added
reasoningEffortMapoption toOpenAICompatfor mapping pi-ai reasoning levels to provider-specificreasoning_effortvalues - Added
extraBodytoOpenAICompatto support provider-specific request body routing fields in OpenAI-completions requests - Added support for reading token usage from choice-level
usagefield as fallback when root-level usage is unavailable - Added new models: DeepSeek-V3.2 (Bedrock), Llama 3.1 405B Instruct, Magistral Small 1.2, Ministral 3 3B, Mistral Large 3, Pixtral Large (25.02), NVIDIA Nemotron Nano 3 30B, and Qwen3-5-9b
- Added
close()method toAuthStoragefor properly closing the underlying credential store - Added
initiatorOverrideoption in OpenAI and Anthropic providers to customize message attribution
Changed
- Changed assistant message content serialization to always use plain string format instead of text block arrays to prevent recursive nesting in OpenAI-compatible backends
- Changed Bedrock Opus 4.6 context window from 1M to 1M and added max tokens limit of 128K
- Changed OpenCode Zen/Go Sonnet 4.0/4.5 context window from 1M to 200K
- Changed GitHub Copilot context windows from 200K to 128K for both gpt-4o and gpt-4o-mini
- Changed Claude 3.5 Sonnet (Anthropic API) pricing: input from $0.5 to $0.25, output from $3 to $1.5, cache read from $0.05 to $0.025, cache write from $0 to $1
- Changed Devstral 2 model name from '135B' to '123B'
- Changed ByteDance Seed 2.0-Lite to support reasoning with effort-based thinking mode and image inputs
- Changed Qwen3-32b (Groq) reasoning effort mapping to normalize all levels to 'default'
- Changed finish_reason 'end' to map to 'stop' for improved compatibility with additional providers
- Changed Anthropic reference model merging to prioritize bundled metadata for known models while using models.dev for newly discovered IDs
Fixed
- Fixed reasoning_effort parameter handling to use provider-specific mappings instead of raw effort values
- Fixed assistant content serialization for GitHub Copilot and other OpenAI-compatible backends that mirror array payloads
- Fixed token usage calculation to properly extract cached tokens from both root and nested
prompt_tokens_detailsfields - Fixed stop reason mapping to handle string values and unknown finish reasons gracefully
- Fixed resource cleanup in
AuthCredentialStore.close()to properly finalize all prepared statements before closing the database
[13.11.1] - 2026-03-13
Fixed
- Added
llama.cppas local provider - Fixed auth schema V0-to-V1 migration crash when the V0 table lacks a
disabledcolumn
[13.11.0] - 2026-03-12
Added
- Added support for Parallel AI provider with API key authentication
- Added
PARALLEL_API_KEYenvironment variable support for Parallel provider configuration - Added automatic websocket reconnection handling for connection limit errors, with fallback to SSE replay when content has already been emitted
Changed
- Enhanced
CodexProviderStreamErrorto include an optional error code field for better error categorization and handling
Fixed
- Improved retry logic to handle HTTP/2 stream errors and internal_error responses from Anthropic API
[13.9.16] - 2026-03-10
Added
- Support for
onPayloadcallback to replace provider request payloads before sending, enabling request interception and modification - Support for structured text signature metadata with phase information (commentary/final_answer) in OpenAI and Azure OpenAI Responses providers
- Support for OpenAI Codex Spark model selection with plan-based account prioritization
- Added
modelIdoption togetApiKey()to enable model-specific credential ranking
Changed
- Enhanced
onPayloadcallback signature to accept model parameter and support async payload replacement - Improved error messages for
response.failedevents to include detailed error codes, messages, and incomplete reasons - Refactored OpenAI Codex response streaming to improve code organization and maintainability with extracted helper functions and type definitions
- Enhanced websocket fallback logic to safely replay buffered output over SSE when websocket connections fail mid-stream
- Improved error recovery for websocket streams by distinguishing between fatal connection errors and retryable stream errors
- Updated credential ranking strategy to prioritize Pro plan accounts when requesting OpenAI Codex Spark models
Fixed
- Fixed websocket stream recovery to properly reset output state and clear buffered items when falling back to SSE after partial output
- Fixed handling of malformed JSON messages in websocket streams to trigger immediate fallback to SSE without retry attempts
[13.9.13] - 2026-03-10
Added
- Added
isSpecialServiceTierutility function to validate OpenAI service tier values
[13.9.12] - 2026-03-09
Added
- Added Tavily web search provider support with API key authentication
Fixed
- Fixed OpenAI-family streaming transports to fail with an explicit idle-timeout error instead of hanging indefinitely when the provider stops sending events mid-response
- Fixed OpenAI Codex OAuth refresh and usage-limit lookups to respect request timeouts instead of waiting indefinitely during account selection or rotation
- Fixed OpenAI Codex prewarmed websocket requests to fall back quickly when the socket connects but never starts the response stream
[13.9.10] - 2026-03-08
Added
- Added
identity_keycolumn to auth credentials storage for improved credential deduplication - Added schema versioning system to auth credentials database for safer migrations
- Added automatic backfilling of identity keys during database schema migrations
Changed
- Changed credential deduplication logic to use single identity key instead of multiple identifiers for better performance
- Changed database schema to store normalized identity keys alongside credentials
- Changed auth schema migration to support upgrading from legacy database versions with automatic data backfill
Fixed
- Fixed API key credential matching to correctly identify when the same key is re-stored, preventing unnecessary row duplication on re-login
- Fixed credential deduplication to correctly handle OAuth accounts with matching emails but different account IDs
- Fixed API key replacement to reuse existing stored rows instead of accumulating disabled duplicates
- Fixed auth storage to preserve newer recorded schema versions when opened by older binaries
[13.9.8] - 2026-03-08
Fixed
- Fixed WebSocket stream fallback logic to safely replay buffered output over SSE when WebSocket fails after partial content has been streamed
[13.9.4] - 2026-03-07
Changed
- Simplified API key credential storage to always replace existing credentials on re-login instead of accumulating multiple keys
- Updated Kagi API key placeholder from
kagi_...toKG_...to match current API key format - Updated Kagi login instructions to clarify Search API access is beta-only and provide support contact
- Disabled usage reporting in streaming responses for Cerebras models due to compatibility issues
Fixed
- Fixed Cerebras model compatibility by preventing
stream_optionsusage requests in chat completions
[13.9.3] - 2026-03-07
Breaking Changes
- Changed
reasoningparameter fromThinkingLevel | undefinedtoEffort | undefinedinSimpleStreamOptions; 'off' is no longer valid (omit the field instead) - Removed
supportsXhigh()function; checkmodel.thinking?.maxLevelinstead - Removed
ThinkingLevelandThinkingEfforttypes; useEffortenum - Removed
getAvailableThinkingLevels()andgetAvailableThinkingEfforts()functions - Changed
transformRequestBody()signature to requireModelparameter as second argument for effort validation - Removed
thinking.tsmodule export; import frommodel-thinking.tsinstead
Added
- Added
incrementalflag toOpenAIResponsesHistoryPayloadto support building conversation history from multiple assistant messages instead of replacing it - Added
dtflag toOpenAIResponsesHistoryPayloadfor transport-level metadata - Added
ThinkingConfiginterface to models for canonical thinking transport metadata with min/max effort levels and provider-specific mode - Added
thinkingfield toModeltype containing per-model thinking capabilities used to clamp and map user-facing effort levels - Added
Effortenum (minimal, low, medium, high, xhigh) as canonical user-facing thinking levels replacingThinkingLevel - Added
enrichModelThinking()function to automatically populate thinking metadata on models based on their capabilities - Added
mapEffortToAnthropicAdaptiveEffort()function to map user effort levels to Anthropic adaptive thinking effort - Added
mapEffortToGoogleThinkingLevel()function to map user effort levels to Google thinking levels - Added
requireSupportedEffort()function to validate and clamp effort levels per model, throwing errors for unsupported combinations - Added
clampThinkingLevelForModel()function to clamp thinking levels to model-supported range - Added
applyGeneratedModelPolicies()andlinkSparkPromotionTargets()exports from model-thinking module - Added
serviceTieroption to control OpenAI processing priority and cost (auto, default, flex, scale, priority) - Added
providerPayloadfield to messages and responses for reconstructing transport-native history - Added Gemini usage provider for tracking quota and tier information
- Added
getCodexAccountId()utility to extract account ID from Codex JWT tokens - Added email extraction from OpenAI Codex OAuth tokens for credential deduplication
Changed
- Changed credential disabling mechanism from boolean
disabledflag todisabled_causetext field for tracking why credentials were disabled - Changed
deleteAuthCredential()anddeleteAuthCredentialsForProvider()methods to require adisabledCauseparameter explaining the reason for disabling - Changed Gemini model parsing to strip
-previewsuffix for consistent model identification - Changed OpenAI Codex websocket error handling to detect fatal connection errors and immediately fall back to SSE without retrying
- Changed OpenAI Codex to always use websockets v2 protocol (removed v1 support)
- Changed
reasoningparameter type fromThinkingLeveltoEffortinSimpleStreamOptions, removing 'off' value (callers should omit the field instead) - Changed thinking configuration to use model-specific metadata instead of hardcoded provider logic for effort mapping
- Changed OpenAI Codex request transformer to accept
Modelparameter for effort validation instead of string model ID - Changed Anthropic provider to use model thinking metadata for determining adaptive thinking support instead of model ID pattern matching
- Changed Google Vertex and Google providers to use shorter variable names for thinking config construction
- Moved thinking-related utilities from
thinking.tsto newmodel-thinking.tsmodule with expanded functionality - Moved model policy functions from
provider-models/model-policies.tstomodel-thinking.ts - Moved
googleGeminiCliUsageProviderfromproviders/google-gemini-cli-usage.tstousage/gemini.ts - Changed default OpenAI model from gpt-5.1-codex to gpt-5.4 across all providers
- Changed
UsageFetchContextto remove cache and now() dependencies—usage fetchers now use Date.now() directly - Removed
resetInMsfield from usage windows; consumers should calculate fromresetsAttimestamp - Changed OpenAI Codex credential ranking to deduplicate by email when accountId matches
- Improved OpenAI Codex error handling with retryable error detection
Removed
- Removed
thinking.tsmodule; usemodel-thinking.tsinstead - Removed
provider-models/model-policies.tsmodule; functionality moved tomodel-thinking.ts - Removed
supportsXhigh()function from models.ts; use model.thinking metadata instead - Removed
ThinkingLevelandThinkingEfforttypes; useEffortenum instead - Removed
getAvailableThinkingLevels()andgetAvailableThinkingEfforts()functions - Removed
model-policiesexport fromprovider-models/index.ts - Removed hardcoded thinking level clamping logic from OpenAI Codex request transformer; now uses model metadata
- Removed
UsageCacheandUsageCacheEntryinterfaces—caching is now handled internally by AuthStorage - Removed
google-gemini-cli-usageexport; use newgeminiusage provider instead - Removed
resetInMscomputation from all usage providers - Removed cache TTL constants and cache management from usage fetchers (claude, github-copilot, google-antigravity, kimi, openai-codex, zai)
Fixed
- Fixed credential purging to respect disabled credentials when deduplicating by email, preventing re-enablement of intentionally disabled credentials
- Fixed OpenAI Codex websocket error reporting to include detailed error messages from error events
- Fixed conversation history reconstruction to support incremental updates from multiple assistant messages while maintaining backward compatibility with full-snapshot payloads
- Fixed OpenAI Codex to reject unsupported effort levels instead of silently clamping them, providing clear error messages about supported efforts
- Fixed model cache normalization to properly apply thinking enrichment when loading cached models
- Fixed dynamic model merging to apply thinking enrichment to merged model results
- Fixed OpenAI Codex streaming to properly include service_tier in SSE payloads
- Fixed type safety in OpenAI responses by removing unsafe type casts on image content blocks
- Fixed credential purging to respect disabled credentials when deduplicating by email
[13.9.2] - 2026-03-05
Added
- Support for redacted thinking blocks in Anthropic messages, enabling secure handling of encrypted reasoning content
- Preservation of latest Anthropic thinking blocks and redacted thinking content during message transformation, even when switching between Anthropic models
Changed
- Assistant message content now includes
RedactedThinkingContenttype alongside existing text, thinking, and tool call blocks - Message transformation logic now preserves signed thinking blocks and redacted thinking for the latest assistant message in Anthropic conversations
Fixed
- Fixed Unicode normalization to consistently apply
toWellFormed()to all text content, including thinking blocks, ensuring proper handling of malformed UTF-16 sequences
[13.9.1] - 2026-03-05
Breaking Changes
- Removed
THINKING_LEVELS,ALL_THINKING_LEVELS,ALL_THINKING_MODES,THINKING_MODE_DESCRIPTIONS, andTHINKING_MODE_LABELSexports - Renamed
formatThinking()togetThinkingMetadata()with changed return type from string toThinkingMetadataobject - Renamed
getAvailableThinkingLevel()togetAvailableThinkingLevels()and added default parameter - Renamed
getAvailableThinkingEffort()togetAvailableThinkingEfforts()and added default parameter
Added
- Added
ThinkingMetadatatype to provide structured access to thinking mode information (value, label, description)
[13.9.0] - 2026-03-05
Added
- Exported new thinking module with
ThinkingEffort,ThinkingLevel, andThinkingModetypes for managing reasoning effort levels - Added
getAvailableThinkingEffort()function to determine supported thinking effort levels based on model capabilities - Added
parseThinkingEffort(),parseThinkingLevel(), andparseThinkingMode()functions for parsing thinking configuration strings - Added
THINKING_LEVELS,ALL_THINKING_LEVELS, andALL_THINKING_MODESconstants for iterating over available thinking options - Added
THINKING_MODE_DESCRIPTIONSandTHINKING_MODE_LABELSfor displaying thinking modes in user interfaces - Added
formatThinking()function to format thinking modes as compact display labels
Changed
- Refactored thinking level handling to distinguish between
ThinkingEffort(provider-level, no "off") andThinkingLevel(user-facing, includes "off") - Updated
ThinkingBudgetstype to useThinkingEffortinstead ofThinkingLevelfor more precise token budget configuration - Improved reasoning option handling to explicitly support "off" value for disabling reasoning across all providers
- Simplified thinking effort mapping logic by centralizing provider-specific clamping behavior
[13.7.8] - 2026-03-04
Added
- Added ZenMux provider support with mixed API routing: Anthropic-owned models discovered from
https://zenmux.ai/api/v1/modelsnow use the Anthropic transport (https://zenmux.ai/api/anthropic), while other ZenMux models use the OpenAI-compatible transport.
[13.7.7] - 2026-03-04
Changed
- Modified response ID normalization to preserve existing item ID prefixes when truncating oversized IDs
- Updated tool call ID normalization to use
fc_prefix for generated item IDs instead ofitem_prefix
Fixed
- Fixed handling of reasoning item IDs to remain untouched during response normalization while function call IDs are properly normalized
[13.7.2] - 2026-03-04
Added
- Added support for Kagi API key authentication via
login kagicommand - Added Kagi to the list of available OAuth providers
Fixed
- MCP tool schemas with
$ref/$defsare now dereferenced before being sent to LLM providers, fixing dangling references that left models without type definitions - Ajv schema validation no longer emits
console.warn()for non-standard format keywords (e.g."uint") from MCP servers, preventing TUI corruption - Tool schema compilation is now cached per schema identity, eliminating redundant recompilation on every tool call
[13.6.0] - 2026-03-03
Added
- Added Anthropic Foundry gateway mode controlled by
CLAUDE_CODE_USE_FOUNDRY, with support forFOUNDRY_BASE_URL,ANTHROPIC_FOUNDRY_API_KEY,ANTHROPIC_CUSTOM_HEADERS, and optional mTLS material (CLAUDE_CODE_CLIENT_CERT,CLAUDE_CODE_CLIENT_KEY,NODE_EXTRA_CA_CERTS) - Added LM Studio provider support with OpenAI-compatible model discovery and OAuth login.
- Added support for
LM_STUDIO_API_KEYandLM_STUDIO_BASE_URLenvironment variables for authentication and custom host configuration.
Changed
- Anthropic key resolution now prefers
ANTHROPIC_FOUNDRY_API_KEYoverANTHROPIC_OAUTH_TOKENandANTHROPIC_API_KEYwhen Foundry mode is enabled - Anthropic auth base-URL fallback now prefers
FOUNDRY_BASE_URLwhenCLAUDE_CODE_USE_FOUNDRYis enabled
[13.5.8] - 2026-03-02
Fixed
- Fixed schema compatibility issue where patternProperties in tool parameters caused failures when converting to legacy Antigravity format
[13.5.5] - 2026-03-01
Changed
- Anthropic Claude system-block cloaking now leaves the agent identity block uncached and applies
cache_control: { type: "ephemeral" }to injected user system blocks without forcingttl: "1h"
Fixed
- Anthropic request payload construction now enforces a maximum of 4
cache_controlbreakpoints (tools/system/messages priority order) before dispatch - Anthropic cache-control normalization now removes later
ttl: "1h"entries when a default/5m block has already appeared earlier in evaluation order
[13.5.3] - 2026-03-01
Fixed
- Fixed tool argument coercion to handle malformed JSON with trailing wrapper braces by parsing leading JSON containers
[13.4.0] - 2026-03-01
Breaking Changes
- Removed
TInputgeneric parameter fromToolResultMessageinterface and removed$normativeproperty
Added
hasUnrepresentableStrictObjectMap()pre-flight check intryEnforceStrictSchema: schemas withpatternPropertiesor schema-valuedadditionalPropertiesnow degrade gracefully to non-strict mode instead of throwing during enforcementgenerateClaudeCloakingUserId()generates structured user IDs for Anthropic OAuth metadata (user_{hex64}_account_{uuid}_session_{uuid})isClaudeCloakingUserId()validates whether a string matches the cloaking user-ID formatmapStainlessOs()andmapStainlessArch()mapprocess.platform/process.archto Stainless header values; X-Stainless-Os and X-Stainless-Arch inclaudeCodeHeadersare now runtime-computedbuildClaudeCodeTlsFetchOptions()attaches SNI and default TLS ciphers for directapi.anthropic.comconnectionscreateClaudeBillingHeader()generates thex-anthropic-billing-headerblock (SHA-256 payload fingerprint + random build hash)buildAnthropicSystemBlocks()now injects a billing header block and the Claude Agent SDK identity block withephemeral1h cache-control whenincludeClaudeCodeInstructionis setresolveAnthropicMetadataUserId()auto-generates a cloaking user ID for OAuth requests whenmetadata.user_idis absent or invalidAnthropicOAuthFlowis now exported for direct use- OAuth callback server timeout extended from 2 min to 5 min
parseGeminiCliCredentials()parses Google Cloud credential JSON with support for legacy ({token,projectId}), alias (project_id/refresh/expires), and enriched formatsshouldRefreshGeminiCliCredentials()and proactive token refresh before requests for both Gemini CLI and Antigravity providers (60s pre-expiry buffer)normalizeAntigravityTools()convertsparametersJsonSchema→parametersin function declarations for Antigravity compatibilityANTIGRAVITY_SYSTEM_INSTRUCTIONis now exported for use by search and other consumersANTIGRAVITY_LOAD_CODE_ASSIST_METADATAconstant exported from OAuth module withANTIGRAVITYideType- Antigravity project onboarding:
onboardProjectWithRetries()provisions a new project viaonboardUserLRO whenloadCodeAssistreturns no existing project (up to 5 attempts, 2s interval) getOAuthApiKeynow includesrefreshToken,expiresAt,email, andaccountIdin the Gemini/Antigravity JSON credential payload to enable proactive refresh- Antigravity model discovery now tries the production daily endpoint first, with sandbox as fallback
ANTIGRAVITY_DISCOVERY_DENYLISTfilters low-quality/internal models from discovery results
Changed
- Replaced
sanitizeSurrogates()utility with nativeString.prototype.toWellFormed()for handling unpaired Unicode surrogates across all providers - Extended
ANTHROPIC_OAUTH_BETAconstant in the OpenAI-compat Anthropic route withinterleaved-thinking-2025-05-14,context-management-2025-06-27, andprompt-caching-scope-2026-01-05beta flags claudeCodeVersionbumped to2.1.63;claudeCodeSystemInstructionupdated to identify as Claude Agent SDKclaudeCodeHeaders: removedX-Stainless-Helper-Method, updated package version to0.74.0, runtime version tov24.3.0applyClaudeToolPrefix/stripClaudeToolPrefixnow accept an optional prefix override and skip Anthropic built-in tool names (web_search,code_execution,text_editor,computer)- Accept-Encoding header updated to
gzip, deflate, br, zstd - Non-Anthropic base URLs now receive
Authorization: Bearerregardless of OAuth status - Prompt-caching logic now skips applying breakpoints when any block already carries
cache_control, instead of stripping then re-applying fine-grained-tool-streaming-2025-05-14removed from default beta set- Anthropic OAuth token URL changed from
platform.claude.comtoapi.anthropic.com - Anthropic OAuth scopes reduced to
org:create_api_key user:profile user:inference - OAuth code exchange now strips URL fragment from callback code, using the fragment as state override when present
- Claude usage headers aligned: user-agent updated to
claude-cli/2.1.63 (external, cli), anthropic-beta extended with full beta set - Antigravity session ID format changed to signed decimal (negative int63 derived from SHA-256 of first user message, or random bounded int63)
- Antigravity
requestIdnow usesagent-{uuid}format; non-Antigravity requests no longer include requestId/userAgent/requestType in the payload ANTIGRAVITY_DAILY_ENDPOINTcorrected todaily-cloudcode-pa.googleapis.com; sandbox endpoint kept as fallback only- Antigravity discovery: removed
recommended/agentModelSortsfilter; now includes all non-internal, non-denylisted models - Antigravity discovery no longer sends
projectin the request body - Gemini/Antigravity OAuth flows no longer use PKCE (code_challenge removed)
- Antigravity
loadCodeAssistmetadata ideType changed fromIDE_UNSPECIFIEDtoANTIGRAVITY - Antigravity
discoverProjectnow uses a single canonical production endpoint; falls back to project onboarding instead of a hardcoded default project ID VALIDATEDtool calling config applied to Antigravity requests with Claude modelsmaxOutputTokensremoved from Antigravity generation config for non-Claude models- System instruction injection for Antigravity scoped to Claude and
gemini-3-pro-highmodels only
Removed
- Removed
sanitizeSurrogates()utility function; use nativeString.prototype.toWellFormed()instead
[13.3.14] - 2026-02-28
Added
- Exported schema utilities from new
./utils/schemamodule, consolidating JSON Schema handling across providers - Added
CredentialRankingStrategyinterface for providers to implement usage-based credential selection - Added
claudeRankingStrategyfor Anthropic OAuth credentials to enable smart multi-account selection based on usage windows - Added
codexRankingStrategyfor OpenAI Codex OAuth credentials with priority boost for fresh 5-hour window starts - Added
adaptSchemaForStrict()helper for unified OpenAI strict schema enforcement across providers - Added schema equality and merging utilities:
areJsonValuesEqual(),mergeCompatibleEnumSchemas(),mergePropertySchemas() - Added Cloud Code Assist schema normalization:
copySchemaWithout(),stripResidualCombiners(),prepareSchemaForCCA() - Added
sanitizeSchemaForGoogle()andsanitizeSchemaForCCA()for provider-specific schema sanitization - Added
StringEnum()helper for creating string enum schemas compatible with Google and other providers - Added
enforceStrictSchema()andsanitizeSchemaForStrictMode()for OpenAI strict mode schema validation - Added package exports for
./utils/schemaand./utils/schema/*subpaths - Added
validateSchemaCompatibility()to statically audit a JSON Schema against provider-specific rules (openai-strict,google,cloud-code-assist-claude) and return structured violations - Added
validateStrictSchemaEnforcement()to verify the strict-fail-open contract: enforced schemas pass strict validation, failed schemas return the original object identity - Added
COMBINATOR_KEYS(anyOf,allOf,oneOf) andCCA_UNSUPPORTED_SCHEMA_FIELDSas exported constants infields.tsto eliminate duplication across modules - Added
tryEnforceStrictSchemaresult cache (WeakMap) to avoid redundant sanitize + enforce work for the same schema object - Added comprehensive schema normalization test suite (
schema-normalization.test.ts) covering strict mode, Google, and Cloud Code Assist normalization paths - Added schema compatibility validation test suite (
schema-compatibility.test.ts) covering all three provider targets
Changed
- Moved schema utilities from
./utils/typebox-helpersto new./utils/schemamodule with expanded functionality - Refactored OpenAI provider tool conversion to use unified
adaptSchemaForStrict()helper across codex, completions, and responses - Updated
AuthStorageto support generic credential ranking viaCredentialRankingStrategyinstead of Codex-only logic - Moved Google schema sanitization functions from
google-shared.tsto./utils/schemamodule - Changed export path:
./utils/typebox-helpers→./utils/schemain main index sanitizeSchemaForGoogle()/sanitizeSchemaForCCA()now accept a parameterizedunsupportedFieldsset internally, enabling code reuse between the two sanitizerscopySchemaWithout()rewritten using object-rest destructuring for clarity
Fixed
- Fixed cycle detection:
WeakSetguards added to all recursive schema traversals (sanitizeSchemaForStrictMode,enforceStrictSchema,normalizeSchemaForCCA,normalizeNullablePropertiesForCloudCodeAssist,stripResidualCombiners,sanitizeSchemaImpl,hasResidualCloudCodeAssistIncompatibilities) — circular schemas no longer cause infinite loops or stack overflows - Fixed
hasResidualCloudCodeAssistIncompatibilities: cycle detection now returnsfalse(nottrue) for already-visited nodes, eliminating false positives that forced the CCA fallback schema on valid recursive inputs - Fixed
stripResidualCombinersto iterate to a fixpoint rather than making a single pass, ensuring chained combiner reductions (where one reduction enables another) are fully resolved - Fixed
mergeObjectCombinerVariantsrequired-field computation: the flattened object now takes the intersection of all variants'requiredarrays (unioned with own-level required properties that exist in the merged schema), preventing required fields from being silently dropped or over-included - Fixed
mergeCompatibleEnumSchemasto use deep structural equality (areJsonValuesEqual) instead ofObject.iswhen deduplicating object-valued enum members - Fixed
sanitizeSchemaForGoogleconst-to-enum deduplication to use deep equality instead of reference equality - Fixed
sanitizeSchemaForGoogletype inference foranyOf/oneOf-flattened const enums: type is now derived from all variants (must agree), falling back to inference from enum values; mixed null/non-null infers the non-null type and setsnullable - Fixed
sanitizeSchemaForGooglerecursion to spread options when descending (previously onlyinsideProperties,normalizeTypeArrayToNullable,stripNullableKeywordwere forwarded; new fieldsunsupportedFieldsandseenwere silently dropped) - Fixed
sanitizeSchemaForGooglearray-valuedtypefiltering to exclude non-string entries before processing - Removed incorrect
additionalProperties: falsestripping fromsanitizeSchemaForGoogle(the field is valid in Google schemas whenfalse) - Fixed
sanitizeSchemaForStrictModeto strip thenullablekeyword and expand it intoanyOf: [schema, {type: "null"}]in the output, matching what OpenAI strict mode actually expects - Fixed
sanitizeSchemaForStrictModeto infertype: "array"whenitemsis present buttypeis absent - Fixed
sanitizeSchemaForStrictModeto infer a scalartypefrom uniformenumvalues whentypeis not explicitly set - Fixed
sanitizeSchemaForStrictModeconst-to-enum merge to use deep equality, preventing duplicate enum entries whenconstandenumboth exist with the same value - Fixed
enforceStrictSchemato dropadditionalPropertiesunconditionally (previously only object-valuedadditionalPropertieswas recursed into; non-object values were passed through, violating strict schema requirements) - Fixed
enforceStrictSchemato recurse into$defsanddefinitionsblocks so referenced sub-schemas are also made strict-compliant - Fixed
enforceStrictSchemato handle tuple-styleitemsarrays (previously only single-schemaitemsobjects were recursed) - Fixed
enforceStrictSchemadouble-wrapping: optional properties already expressed asanyOf: [..., {type: "null"}]are not wrapped again - Fixed
enforceStrictSchemaArray.isArraytype-narrowing fortypefield to filter non-string entries before checking for"object"
[13.3.8] - 2026-02-28
Fixed
- Fixed response body reuse error when handling 429 rate limit responses with retry logic
[13.3.7] - 2026-02-27
Added
- Added
tryEnforceStrictSchemafunction that gracefully downgrades to non-strict mode when schema enforcement fails, enabling better compatibility with malformed or circular schemas - Added
sanitizeSchemaForStrictModefunction to normalize JSON schemas by stripping non-structural keywords, convertingconsttoenum, and expanding type arrays intoanyOfvariants - Added Kilo Gateway provider support with OpenAI-compatible model discovery, OAuth
/login kilo, andKILO_API_KEYenvironment variable support (#193)
Changed
- Changed strict mode handling in OpenAI providers to use
tryEnforceStrictSchemafor safer schema enforcement with automatic fallback to non-strict mode - Enhanced
enforceStrictSchemato properly handle schemas with type arrays containingobject(e.g.,type: ["object", "null"])
Fixed
- Fixed
enforceStrictSchemato properly handle malformed object schemas with required keys but missing properties - Fixed
enforceStrictSchemato correctly process nested object schemas withinanyOf,allOf, andoneOfcombinators
[13.3.1] - 2026-02-26
Added
- Added
topP,topK,minP,presencePenalty, andrepetitionPenaltyoptions toStreamOptionsfor fine-grained control over model sampling behavior
[13.3.0] - 2026-02-26
Changed
- Allowed OAuth provider logins to supply a manual authorization code handler with a default prompt when none is provided
[13.2.0] - 2026-02-23
Added
- Added support for GitHub Copilot provider in strict mode for both openai-completions and openai-responses tool schemas
Fixed
- Fixed tool descriptions being rejected when undefined by providing empty string fallback across all providers
[12.19.1] - 2026-02-22
Added
- Exported
isProviderRetryableErrorfunction for detecting rate-limit and transient stream errors - Support for retrying malformed JSON stream-envelope parse errors from Anthropic-compatible proxy endpoints
Changed
- Expanded retry detection to include JSON parse errors (unterminated strings, unexpected end of input) in addition to rate-limit errors
[12.19.0] - 2026-02-22
Added
- Added GitLab Duo provider with support for Claude, GPT-5, and other models via GitLab AI Gateway
- Added OAuth authentication for GitLab Duo with automatic token refresh and direct access caching
- Added 16 new GitLab Duo models including Claude Opus/Sonnet/Haiku variants and GPT-5 series models
- Added
isOAuthoption to Anthropic provider to force OAuth bearer auth mode for proxy tokens - Added
streamGitLabDuofunction to route requests through GitLab AI Gateway with direct access tokens - Added
getGitLabDuoModelsfunction to retrieve available GitLab Duo model configurations - Added
clearGitLabDuoDirectAccessCachefunction to manually clear cached direct access tokens
Changed
- Enhanced
getModelMapping()to support both GitLab Duo alias IDs (e.g.,duo-chat-gpt-5-codex) and canonical model IDs (e.g.,gpt-5-codex) for improved model resolution flexibility - Migrated
AuthCredentialStoreandAuthStorageinto@oh-my-pi/pi-aias shared credential primitives for downstream packages - Moved Anthropic auth helpers (
findAnthropicAuth,isOAuthToken,buildAnthropicSearchHeaders,buildAnthropicUrl) into shared AI utilities for reuse across providers - Replaced
CliAuthStoragewithAuthCredentialStorefor improved credential management with multiple credentials per provider - Updated models.json pricing for Claude 3.5 Sonnet (input: 0.23→0.45, output: 3→2.2, added cache read: 0.225) and Claude 3 Opus (input: 0.3→0.95)
- Moved
mapAnthropicToolChoicefunction from gitlab-duo provider to stream module for broader reusability - Enhanced HTTP status code extraction to handle string-formatted status codes in error objects
Removed
- Removed
CliAuthStorageclass in favor of newAuthCredentialStorewith enhanced functionality
[12.17.2] - 2026-02-21
Added
- Exported
getAntigravityUserAgent()function for constructing Antigravity User-Agent headers
Changed
- Updated default Antigravity version from 1.15.8 to 1.18.3
- Unified User-Agent header generation across Antigravity API calls to use centralized
getAntigravityUserAgent()function
[12.17.1] - 2026-02-21
Added
- Added new export paths for provider models via
./provider-modelsand./provider-models/* - Added new export paths for Cursor and OpenAI Codex providers via
./providers/cursor/gen/*and./providers/openai-codex/* - Added new export paths for usage utilities via
./usage/* - Added new export paths for discovery and OAuth utilities via
./utils/discoveryand./utils/oauthwith subpath exports
Changed
- Simplified main export path to use wildcard pattern
./src/*.tsfor broader module access - Updated
models.jsonexport to include TypeScript declaration file at./src/models.json.d.ts - Reorganized package.json field ordering for improved readability
[12.17.0] - 2026-02-21
Fixed
- Cursor provider: bind
execHandlerswhen passing handler methods to the exec protocol so handlers receive correctthiscontext (fixes "undefined is not an object (evaluating 'this.options')" when using exec tools such as web search with Cursor)
[12.16.0] - 2026-02-21
Added
- Exported
readModelCacheandwriteModelCachefunctions for direct SQLite-backed model cache access - Added
<turn_aborted>guidance marker as synthetic user message when assistant messages are aborted or errored, informing the model that tools may have partially executed - Added support for Sonnet 4.6 models in adaptive thinking detection
Changed
- Updated model cache schema version to support improved global model fallback resolution
- Improved GitHub Copilot model resolution to prefer provider-specific model definitions over global references when context window is larger, ensuring optimal model capabilities
- Migrated model cache from per-provider JSON files to unified SQLite database (models.db) for atomic cross-process access
- Renamed
cachePathoption tocacheDbPathin ModelManagerOptions to reflect database-backed storage - Improved non-authoritative cache handling with 5-minute retry backoff instead of retrying on every startup
- Modified handling of aborted/errored assistant messages to preserve tool call structure instead of converting to text summaries, with synthetic 'aborted' tool results injected
- Updated tool call tracking to use status map (Resolved/Aborted) instead of separate sets for better handling of duplicate and aborted tool results
[12.15.0] - 2026-02-20
Fixed
- Improved error messages for OAuth token refresh failures by including detailed error information from the provider
- Separated rate limit and usage limit error handling to provide distinct user-friendly messages for ChatGPT rate limits vs subscription usage limits
Changed
- Increased SDK retry attempts to 5 for OpenAI, Azure OpenAI, and Anthropic clients (was SDK default of 2)
- Changed 429 retry strategy for OpenAI Codex and Google Gemini CLI to use a 5-minute time budget when the server provides a retry delay, instead of a fixed attempt cap
[12.14.0] - 2026-02-19
Added
- Added
gemini-3.1-promodel to opencode provider with text and image input support - Added
trinity-large-preview-freemodel to opencode provider - Added
google/gemini-3.1-pro-previewmodel to nanogpt provider - Added
google/gemini-3.1-pro-previewmodel to openrouter provider with text and image input support - Added
gemini-3.1-promodel to cursor provider - Added optional
intentfield toToolCallinterface for harness-level intent metadata
Changed
- Changed
big-picklemodel API fromopenai-completionstoanthropic-messages - Changed
big-picklemodel baseUrl fromhttps://opencode.ai/zen/v1tohttps://opencode.ai/zen - Changed
minimax-m2.5-freemodel API fromopenai-completionstoanthropic-messages - Changed
minimax-m2.5-freemodel baseUrl fromhttps://opencode.ai/zen/v1tohttps://opencode.ai/zen
Fixed
- Fixed tool argument validation to iteratively coerce nested JSON strings across multiple passes, enabling proper handling of deeply nested JSON-serialized objects and arrays
[12.13.0] - 2026-02-19
Added
- Added NanoGPT provider support with API-key login, dynamic model discovery from
https://nano-gpt.com/api/v1/models, and text-model filtering for catalog/runtime discovery (#111)
[12.12.3] - 2026-02-19
Fixed
- Fixed retry logic to recognize 'unable to connect' errors as transient failures
[12.11.3] - 2026-02-19
Fixed
- Fixed OpenAI Codex streaming to fail truncated responses that end without a terminal completion event, preventing partial outputs from being treated as successful completions.
- Fixed Codex websocket append fallback by resetting stale turn-state/model-etag session metadata when request shape diverges from appendable history.
[12.11.1] - 2026-02-19
Added
- Added support for Claude 4.6 Opus and Sonnet models via Cursor API
- Added support for Composer 1.5 model via Cursor API
- Added support for GPT-5.1 Codex Mini and GPT-5.1 High models via Cursor API
- Added support for GPT-5.2 and GPT-5.3 Codex variants (Fast, High, Low, Extra High) via Cursor API
- Added HTTP/2 transport support for Cursor API requests (required by Cursor API)
Changed
- Updated pricing for Claude 3.5 Sonnet model
- Updated Claude 3.5 Sonnet context window from 262,144 to 131,072 tokens
- Simplified Cursor model display names by removing '(Cursor)' suffix
- Changed Cursor API timeout from 15 seconds to 5 seconds
- Switched Cursor API transport from HTTP/1.1 to HTTP/2
[12.11.0] - 2026-02-19
Added
- Added
priorityfield to Model interface for provider-assigned model prioritization - Added
CatalogDiscoveryConfiginterface to standardize catalog discovery configuration across providers - Added type guards
isCatalogDescriptor()andallowsUnauthenticatedCatalogDiscovery()for safer descriptor handling - Added
DEFAULT_MODEL_PER_PROVIDERexport from descriptors module for centralized default model management - Support for 11 new AI providers: Cloudflare AI Gateway, Hugging Face Inference, LiteLLM, Moonshot, NVIDIA, Ollama, Qianfan, Qwen Portal, Together, Venice, vLLM, and Xiaomi MiMo
- Login flows for new providers with API key validation and OAuth token support
- Extended
KnownProvidertype to include all newly supported providers - API key environment variable mappings for all new providers in service provider map
- Model discovery and configuration for Cloudflare AI Gateway, Hugging Face, LiteLLM, Moonshot, NVIDIA, Ollama, Qianfan, Qwen Portal, Together, Venice, vLLM, and Xiaomi MiMo
Changed
- Refactored OAuth credential retrieval to simplify storage lifecycle management in model generation script
- Parallelized special model discovery sources (Antigravity, Codex) for improved generation performance
- Reorganized model JSON structure to place
contextWindowandmaxTokensbeforecompatfield for consistency - Added
priorityfield to OpenAI Codex models for provider-assigned model prioritization - Refactored provider descriptors to use helper functions (
descriptor,catalog,catalogDescriptor) for reduced code duplication - Refactored models.dev provider descriptors to use helper functions (
simpleModelsDevDescriptor,openAiCompletionsDescriptor,anthropicMessagesDescriptor) for improved maintainability - Unified provider descriptors into single source of truth in
descriptors.tsfor both runtime model discovery and catalog generation, improving maintainability - Refactored model generation script to use declarative
CatalogProviderDescriptorinterface instead of separate descriptor types, reducing code duplication - Reorganized models.dev provider descriptors into logical groups (Bedrock, Core, Coding Plans, Specialized) for better code organization
- Simplified API resolution for OpenCode and GitHub Copilot providers using rule-based matching instead of inline conditionals
- Refactored model generation script to use declarative provider descriptors instead of inline provider-specific logic, improving maintainability and reducing code duplication
- Extracted model post-processing policies (cache pricing corrections, context window normalization) into dedicated
model-policies.tsmodule for better testability and clarity - Removed static bundled models for Ollama and vLLM from
models.jsonto rely on dynamic discovery instead, reducing static catalog size - Updated
OAuthProvidertype to include new provider identifiers - Expanded model registry (models.json) with thousands of new model entries across all new providers
- Modified environment variable resolution to use
$pickenvfor providers with multiple possible env var names - Updated README documentation to list all newly supported providers and their authentication requirements
[12.10.1] - 2026-02-18
- Added Synthetic provider
- Added API-key login helpers for Synthetic and Cerebras providers
[12.10.0] - 2026-02-18
Breaking Changes
- Renamed public API functions:
getModel()→getBundledModel(),getModels()→getBundledModels(),getProviders()→getBundledProviders()
Added
- Exported
ModelManagerAPI for runtime-aware model resolution with dynamic endpoint discovery - Exported provider-specific model manager configuration helpers for Google, OpenAI-compatible, Codex, and Cursor providers
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints
- Added
createModelManager()function to manage bundled and dynamically discovered models with configurable refresh strategies - Added support for on-disk model caching with TTL-based invalidation
- Added
resolveProviderModels()function for runtime model resolution across multiple providers - Added EU cross-region inference variants for Claude Haiku 3.5 on Bedrock
- Added Claude Sonnet 4.6 and Claude Sonnet 4.6 Thinking models to Antigravity provider
- Added GLM-5 Free model via OpenCode provider
- Added GLM-4.7-FlashX model via ZAI provider
- Added MiniMax-M2.5-highspeed model across multiple providers (minimax-code, minimax-code-cn, minimax, minimax-cn)
- Added Claude Sonnet 4.6 model to OpenRouter provider
- Added Qwen 3.5 Plus model to Vercel AI Gateway provider
- Added Claude Sonnet 4.6 model to Vercel AI Gateway provider
Changed
- Renamed
getModel()togetBundledModel()to clarify it returns compile-time bundled models only - Renamed
getModels()togetBundledModels()for consistency - Renamed
getProviders()togetBundledProviders()for consistency - Refactored model generation script to use modular discovery functions instead of monolithic provider-specific logic
- Updated models.json with new model entries and pricing updates across multiple providers
- Updated pricing for deepseek/deepseek-v3 model on OpenRouter
- Updated maxTokens from 65536 to 4096 for deepseek/deepseek-v3 on OpenRouter
- Updated pricing and maxTokens for mistralai/mistral-large-2411 on OpenRouter
- Updated pricing for qwen/qwen-max on Together AI
- Updated pricing for qwen/qwen-vl-plus on Together AI
- Updated pricing for qwen/qwen-plus on Together AI
- Updated pricing for qwen/qwen-turbo on Together AI
- Expanded EU cross-region inference variant support to all Claude models on Bedrock (previously limited to Haiku, Sonnet, and Opus 4.5)
[12.8.0] - 2026-02-16
Added
- Added
contextPromotionTargetmodel property to specify preferred fallback model when context promotion is triggered - Added automatic context promotion target assignment for Spark models to their base model equivalents
- Added support for Brave search provider with BRAVE_API_KEY environment variable
Changed
- Updated Qwen model context window and max token limits for improved accuracy
[12.7.0] - 2026-02-16
Added
- Added DeepSeek-V3.2 model support via Amazon Bedrock
- Added GLM-5 model support via OpenCode
- Added MiniMax M2.5 model support via OpenCode
Changed
- Updated GLM-4.5, GLM-4.5-Air, GLM-4.5-Flash, GLM-4.5V, GLM-4.6, GLM-4.6V, GLM-4.7, GLM-4.7-Flash, and GLM-5 models to use anthropic-messages API instead of openai-completions
- Updated GLM models base URL from https://api.z.ai/api/coding/paas/v4 to https://api.z.ai/api/anthropic
- Updated pricing for multiple models including Mistral, Moonshot, and Qwen variants
- Updated context window and max tokens for several models to reflect accurate specifications
Removed
- Removed compat field with supportsDeveloperRole and thinkingFormat properties from GLM models
[12.6.0] - 2026-02-16
Added
- Added source-scoped custom API and OAuth provider registration helpers for extension-defined providers.
Changed
- Expanded
Apityping to allow extension-defined API identifiers while preserving built-in API exhaustiveness checks.
Fixed
- Fixed custom API registration to reject built-in API identifiers and prevent accidental provider overrides.
[12.2.0] - 2026-02-13
Added
- Added automatic retry logic for WebSocket stream closures before response completion, with configurable retry budget to improve reliability on flaky connections
- Added
providerSessionStateoption to enable provider-scoped mutable state persistence across agent turns - Added WebSocket retry logic with configurable retry budget and delay via
PI_CODEX_WEBSOCKET_RETRY_BUDGETandPI_CODEX_WEBSOCKET_RETRY_DELAY_MSenvironment variables - Added WebSocket idle timeout detection via
PI_CODEX_WEBSOCKET_IDLE_TIMEOUT_MSenvironment variable to fail stalled connections - Added WebSocket v2 beta header support via
PI_CODEX_WEBSOCKET_V2environment variable for newer OpenAI API versions - Added WebSocket handshake header capture to extract and replay session metadata (turn state, models etag, reasoning flags) across SSE fallback requests
- Added
preferWebsocketsoption to enable WebSocket transport for OpenAI Codex responses when supported - Added
prewarmOpenAICodexResponses()function to establish and reuse WebSocket connections across multiple requests - Added
getOpenAICodexTransportDetails()function to inspect transport layer details including WebSocket status and fallback information - Added
getProviderDetails()function to retrieve formatted provider configuration and transport information - Added automatic fallback from WebSocket to SSE when connection fails, with transparent retry logic
- Added session state management to reuse WebSocket connections and enable request appending across turns
- Added support for x-codex-turn-state header to maintain conversation state across SSE requests
Changed
- Changed WebSocket session state storage from global maps to provider-scoped session state for multi-agent isolation
- Changed WebSocket connection initialization to accept idle timeout configuration and handshake header callbacks
- Changed WebSocket error handling to use standardized transport error messages with
Codex websocket transport errorprefix - Changed WebSocket retry behavior to retry transient failures before activating sticky fallback, improving reliability on flaky connections
- Changed OpenAI Codex model configuration to prefer WebSocket transport by default with
preferWebsockets: true - Changed header handling to use appropriate OpenAI-Beta header values for WebSocket vs SSE transports
- Perplexity OAuth token refresh now uses JWT expiry extraction instead of Socket.IO RPC, improving reliability when server is unreachable
- Removed Socket.IO client implementation for Perplexity token refresh; tokens are now validated using embedded JWT expiry claims
Removed
- Removed
refreshPerplexityTokenexport; token refresh is now handled internally via JWT expiry detection
Fixed
- Fixed WebSocket stream retry logic to properly handle mid-stream connection closures and retry before falling back to SSE transport
- Fixed
preferWebsocketsoption handling to correctly respect explicitfalsevalues when determining transport preference - Fixed WebSocket append state not being reset after aborted requests, preventing stale state from affecting subsequent turns
- Fixed WebSocket append state not being reset after stream errors, preventing failed append attempts from blocking future requests
- Fixed Codex model context window metadata to use 272000 input tokens (instead of 400000 total budget) for non-Spark Codex variants
[12.0.0] - 2026-02-12
Added
- Added GPT-5.3 Codex Spark model with 128K context window and extended reasoning capabilities
- Added MiniMax M2.5 and M2.5 Lightning models via OpenAI-compatible API (minimax-code provider)
- Added MiniMax M2.5 and M2.5 Lightning models via OpenAI-compatible API (minimax-code-cn provider for China region)
- Added MiniMax M2.5 and M2.5 Lightning models via Anthropic API (minimax and minimax-cn providers)
- Added Llama 3.1 8B model via Cerebras API
- Added MiniMax M2.5 model via OpenRouter
- Added MiniMax M2.5 model via Vercel AI Gateway
- Added MiniMax M2.5 Free model via OpenCode
- Added Qwen3 VL 32B Instruct multimodal model via OpenRouter
Changed
- Updated Z.ai GLM-5 pricing and context window configuration on OpenRouter
- Updated Qwen3 Max Thinking max tokens from 32768 to 65536 on OpenRouter
- Updated OpenAI GPT-5 Image Mini pricing on OpenRouter
- Updated OpenAI GPT-5 Pro pricing and context window on OpenRouter
- Updated OpenAI o4-mini pricing and context window on OpenRouter
- Updated Claude Opus 4.5 Thinking model name formatting (removed parentheses)
- Updated Claude Opus 4.6 Thinking model name formatting (removed parentheses)
- Updated Claude Sonnet 4.5 Thinking model name formatting (removed parentheses)
- Updated Gemini 2.5 Flash Thinking model name formatting (removed parentheses)
- Updated Gemini 3 Pro High and Low model name formatting (removed parentheses)
- Updated GPT-OSS 120B Medium model name formatting (removed parentheses) and context window to 131072
Removed
- Removed GLM-5 model from Z.ai provider
- Removed Trinity Large Preview Free model from OpenCode provider
- Removed MiniMax M2.1 Free model from OpenCode provider
- Removed deprecated Anthropic model entries:
claude-3-5-haiku-latest,claude-3-5-haiku-20241022,claude-3-7-sonnet-20250219,claude-3-7-sonnet-latest,claude-3-opus-20240229,claude-3-sonnet-20240229(#33)
Fixed
- Added deprecation filter in model generation script to prevent re-adding deprecated Anthropic models (#33)
[11.14.1] - 2026-02-12
Added
- Added prompt-caching-scope-2026-01-05 beta feature support
Changed
- Updated Claude Code version header to 2.1.39
- Updated runtime version header to v24.13.1 and package version to 0.73.0
- Increased request timeout from 60s to 600s
- Reordered Accept-Encoding header values for compression preference
- Updated OAuth authorization and token endpoints to use platform.claude.com
- Expanded OAuth scopes to include user:sessions:claude_code and user:mcp_servers
Removed
- Removed claude-code-20250219 beta feature from default models
- Removed fine-grained-tool-streaming-2025-05-14 beta feature
[11.13.1] - 2026-02-12
Added
- Added Perplexity (Pro/Max) OAuth login support via native macOS app extraction or email OTP authentication
- Added
loginPerplexityandrefreshPerplexityTokenfunctions for Perplexity account integration - Added Socket.IO v4 client implementation for authenticated WebSocket communication with Perplexity API
[11.12.0] - 2026-02-11
Changed
- Increased maximum retry attempts for Codex requests from 2 to 5 to improve reliability on transient failures
Fixed
- Fixed tool result content handling in Anthropic provider to provide fallback error message when content is empty
- Improved retry delay calculation to parse delay values from error response bodies (e.g., 'Please try again in 225ms')
[11.11.0] - 2026-02-10
Breaking Changes
- Replaced
./models.generatedexport with./models.json- update imports fromimport { MODELS } from './models.generated'toimport MODELS from './models.json' with { type: 'json' }
Added
- Added TypeScript type declarations for
models.jsonto enable proper type inference when importing the JSON file
Changed
- Updated available models in google-antigravity provider with new model variants and updated context window/token limits
- Simplified type signatures for
getModel()andgetModels()functions for improved usability - Changed models export from TypeScript module to JSON format for improved performance and reduced bundle size
- Updated
@anthropic-ai/sdkdependency from ^0.72.1 to ^0.74.0
[11.10.0] - 2026-02-10
Added
- Added support for Kimi K2, K2 Turbo Preview, and K2.5 models with reasoning capabilities
Fixed
- Fixed Claude Opus 4.6 context window to 200K across all providers (was incorrectly set to 1M)
- Fixed Claude Sonnet 4 context window to 200K across multiple providers (was incorrectly set to 1M)
[11.8.0] - 2026-02-10
Added
- Added
automodel alias for OpenRouter with automatic model routing - Added
openrouter/aurora-alphamodel with reasoning capabilities - Added
qwen/qwen3-max-thinkingmodel with extended context window support - Added support for
parametersJsonSchemain Google Gemini tool definitions for improved JSON Schema compatibility
Changed
- Updated Claude Sonnet 4 and 4.5 context window from 1M to 200K tokens to reflect actual limits
- Updated Claude Opus 4.6 context window to 200K tokens across providers
- Changed default
reasoningSummaryfor OpenAI Codex fromundefinedtoauto - Updated Qwen model pricing and context window specifications across multiple variants
- Modified Google Gemini CLI system instruction to use compact format
- Changed tool parameter handling for Claude models on Google Cloud Code Assist to use legacy
parametersfield for API translation
Removed
- Removed
glm-4.7-freemodel from OpenCode provider - Removed
qwen3-codermodel from OpenCode provider - Removed
ai21/jamba-mini-1.7model from OpenRouter - Removed
stepfun-ai/step3model from OpenRouter - Removed duplicate test suite for Google Antigravity Provider with
gemini-3-pro-high
Fixed
- Fixed Amazon Bedrock HTTP/1.1 handler import to use direct import instead of dynamic import
- Fixed Qwen model context window and pricing inconsistencies across OpenRouter
- Fixed cache read pricing for multiple Qwen models
- Fixed OpenAI Codex reasoning effort clamping for
gpt-5.3-codexmodel
[11.7.1] - 2026-02-07
Added
- Added Claude Opus 4.6 Thinking model for Antigravity provider
- Added Gemini 2.5 Flash, Gemini 2.5 Flash Thinking, and Gemini 2.5 Pro models for Antigravity provider
- Added Pony Alpha model via OpenRouter
Changed
- Updated Antigravity models to use free tier pricing (0 cost) across all models
- Changed Antigravity model fetching to dynamically load from API when credentials are available, with hardcoded fallback models
- Updated Claude Opus 4.6 context window from 200,000 to 1,000,000 tokens across Bedrock regions
- Updated Claude Opus 4.6 cache pricing from 1.5/18.75 to 0.5/6.25 for EU and US regions
- Updated Antigravity model pricing to free tier (0 cost) for Claude Opus 4.5 Thinking, Claude Sonnet 4.5 Thinking, Gemini 3 Flash, Gemini 3 Pro variants, and GPT-OSS 120B Medium
- Updated GPT-OSS 120B Medium reasoning capability from false to true
- Updated Gemini 3 Flash max tokens from 65,535 to 65,536
- Updated Claude Opus 4.5 Thinking display name formatting to include parentheses
- Updated various model pricing and context window parameters across OpenRouter and other providers
- Removed Claude Opus 4.6 20260205 model from Anthropic provider
Fixed
- Fixed Claude Opus 4.6 model ID format by removing version suffix (:0) in Bedrock configurations
- Fixed Llama 3.1 70B Instruct pricing and context window parameters
- Fixed Mistral model pricing and cache read costs
- Fixed DeepSeek and other model pricing inconsistencies
- Fixed Qwen model pricing and token limits
- Fixed GLM model pricing and context window specifications
[11.6.0] - 2026-02-07
Added
- Added Bedrock cache retention support with
PI_CACHE_RETENTIONenv var and per-requestcacheRetentionoption - Added adaptive thinking support for Bedrock Opus 4.6+ models
- Added
AWS_BEDROCK_SKIP_AUTHenv var to support unauthenticated Bedrock proxies - Added
AWS_BEDROCK_FORCE_HTTP1env var to force HTTP/1.1 for custom Bedrock endpoints - Re-exported
Static,TSchema, andTypefrom@sinclair/typebox
Fixed
- Fixed OpenAI Responses storage disabled by default (
store: false) - Fixed reasoning effort clamping for gpt-5.3 Codex models (minimal -> low)
- Fixed Bedrock
supportsPromptCachingto also check model cost fields
[11.5.1] - 2026-02-07
Fixed
- Fixed schema normalization to handle array-valued
typefields by converting them to a single type with nullable flag for Google provider compatibility
[11.3.0] - 2026-02-06
Added
- Added
cacheRetentionoption to control prompt cache retention preference ('none', 'short', 'long') across providers - Added
maxRetryDelayMsoption to cap server-requested retry delays and fail fast when delays exceed the limit - Added
effortoption for Anthropic Opus 4.6+ models to control adaptive thinking effort levels ('low', 'medium', 'high', 'max') - Added support for Anthropic Opus 4.6+ adaptive thinking mode that lets Claude decide when and how much to think
- Added
PI_AI_ANTIGRAVITY_VERSIONenvironment variable to customize Antigravity sandbox endpoint version - Exported
convertAnthropicMessagesfunction for converting message formats to Anthropic API - Automatic fallback for Anthropic assistant-prefill requests: appends synthetic user "Continue." message when conversation ends with assistant turn to maintain API compatibility
Changed
- Changed
supportsXhigh()to include GPT-5.1 Codex Max and broaden Anthropic support to all Anthropic Messages API models with budget-based thinking capability - Changed Anthropic thinking mode to use adaptive thinking for Opus 4.6+ models instead of budget-based thinking
- Changed
supportsXhigh()to support GPT-5.2/5.3 and Anthropic Opus 4.6+ models with adaptive thinking - Changed prompt caching to respect
cacheRetentionoption and support TTL configuration for Anthropic - Changed OpenAI tool definitions to conditionally include
strictfield only when provider supports it - Changed Qwen model support to use
enable_thinkingboolean parameter instead of OpenAI-style reasoning_effort
Fixed
- Fixed indentation and formatting in
convertAnthropicMessagesfunction - Fixed handling of conversations ending with assistant messages on Anthropic-routed models that reject assistant prefill requests
[11.2.3] - 2026-02-05
Added
- Added Claude Opus 4.6 model support across multiple providers (Anthropic, Amazon Bedrock, GitHub Copilot, OpenRouter, OpenCode, Vercel AI Gateway)
- Added GPT-5.3 Codex model support for OpenAI
- Added
readSseJsonutility import for improved SSE stream handling in Google Gemini CLI provider
Changed
- Updated Google Gemini CLI provider to use
readSseJsonutility for cleaner SSE stream parsing - Updated pricing for Llama 3.1 405B model on Vercel AI Gateway (cache read rate adjusted)
- Updated Llama 3.1 405B context window and max tokens on Vercel AI Gateway (256000 for both)
Removed
- Removed Kimi K2, Kimi K2 Turbo Preview, and Kimi K2.5 models
- Removed Deep Cogito Cogito V2 Preview models from OpenRouter
[11.0.0] - 2026-02-05
Changed
- Replaced direct
process.envaccess withgetEnv()utility from@oh-my-pi/pi-utilsfor consistent environment variable handling across all providers - Updated environment variable names from
OMP_*prefix toPI_*prefix for consistency (e.g.,OMP_CODING_AGENT_DIR→PI_CODING_AGENT_DIR)
Removed
- Removed automatic environment variable migration from
PI_*toOMP_*prefixes viamigrate-env.tsmodule
[10.5.0] - 2026-02-04
Changed
- Updated @anthropic-ai/sdk to ^0.72.1
- Updated @aws-sdk/client-bedrock-runtime to ^3.982.0
- Updated @google/genai to ^1.39.0
- Updated @smithy/node-http-handler to ^4.4.9
- Updated openai to ^6.17.0
- Updated @types/node to ^25.2.0
Removed
- Removed proxy-agent dependency
- Removed undici dependency
[9.4.0] - 2026-01-31
Added
- Added
getEnv()function to retrieve environment variables from process.env, cwd/.env, or ~/.env - Added support for reading .env files from home directory and current working directory
- Added support for
exaandperplexityas known providers ingetEnvApiKey()
Changed
- Changed
getEnvApiKey()to check process.env, cwd/.env, and ~/.env files in order of precedence - Refactored provider API key resolution to use a declarative service provider map
[9.2.2] - 2026-01-31
Added
- Added OpenCode Zen provider with API key authentication for accessing multiple AI models
- Added 4 new free models via OpenCode: glm-4.7-free, kimi-k2.5-free, minimax-m2.1-free, trinity-large-preview-free
- Added glm-4.7-flash model via Zai provider
- Added Kimi Code provider with OpenAI and Anthropic API format support
- Added prompt cache retention support with PI_CACHE_RETENTION env var
- Added overflow patterns for Bedrock, MiniMax, Kimi; reclassified 429 as rate limiting
- Added profile endpoint integration to resolve user emails with 24-hour caching
- Added automatic token refresh for expired Kimi OAuth credentials
- Added Kimi Code OAuth handler with device authorization flow
- Added Kimi Code usage provider with quota caching
- Added 4 new Kimi Code models (kimi-for-coding, kimi-k2, kimi-k2-turbo-preview, kimi-k2.5)
- Added Kimi Code provider integration with OAuth and token management
- Added tool-choice utility for mapping unified ToolChoice to provider-specific formats
- Added ToolChoice type for controlling tool selection (auto, none, any, required, function)
Changed
- Updated Kimi K2.5 cache read pricing from 0.1 to 0.08
- Updated MiniMax M2 pricing: input 0.6→0.6, output 3→3, cache read 0.1→0.09999999999999999
- Updated OpenRouter DeepSeek V3.1 pricing and max tokens: input 0.6→0.5, output 3→2.8, maxTokens 262144→4096
- Updated OpenRouter DeepSeek R1 pricing and max tokens: input 0.06→0.049999999999999996, output 0.24→0.19999999999999998, maxTokens 262144→4096
- Updated Anthropic Claude 3.5 Sonnet max tokens from 256000 to 65536 on OpenRouter
- Updated Vercel AI Gateway Claude 3.5 Sonnet cache read pricing from 0.125 to 0.13
- Updated Vercel AI Gateway Claude 3.5 Sonnet New cache read pricing from 0.125 to 0.13
- Updated Vercel AI Gateway GPT-5.2 cache read pricing from 0.175 to 0.18 and display name to 'GPT 5.2'
- Updated Zai GLM-4.6 cache read pricing from 0.024999999999999998 to 0.03
- Updated Zai Qwen QwQ max tokens from 66000 to 16384
- Added delta event batching and throttling (50ms, 20 updates/sec max) to AssistantMessageEventStream
- Updated MiniMax-M2 pricing: input 1.2→0.6, output 1.2→3, cacheRead 0.6→0.1
Removed
- Removed OpenRouter google/gemini-2.0-flash-exp:free model
- Removed Vercel AI Gateway stealth/sonoma-dusk-alpha and stealth/sonoma-sky-alpha models
Fixed
- Fixed rate limit issues with Kimi models by always sending max_tokens
- Added handling for sensitive stop reason from Anthropic API safety filters
- Added optional chaining for safer JSON schema property access in Anthropic provider
[8.6.0] - 2026-01-27
Changed
- Replaced JSON5 dependency with Bun.JSON5 parsing
Fixed
- Filtered empty user text blocks for OpenAI-compatible completions and normalized Kimi reasoning_content for OpenRouter tool-call messages
[8.4.0] - 2026-01-25
Added
- Added Azure OpenAI Responses provider with deployment mapping and resource-based base URL support
Changed
- Added OpenRouter routing preferences for OpenAI-compatible completions
Fixed
- Defaulted Google tool call arguments to empty objects when providers omit args
- Guarded Responses/Codex streaming deltas against missing content parts and handled arguments.done events
[8.2.1] - 2026-01-24
Fixed
- Fixed handling of streaming function call arguments in OpenAI responses to properly parse arguments when sent via
response.function_call_arguments.doneevents
[8.2.0] - 2026-01-24
Changed
- Migrated node module imports from named to namespace imports across all packages for consistency with project guidelines
[8.0.0] - 2026-01-23
Fixed
- Fixed OpenAI Responses API 400 error "function_call without required reasoning item" when switching between models (same provider, different model). The fix omits the
idfield for function_calls from different models to avoid triggering OpenAI's reasoning/function_call pairing validation - Fixed 400 errors when reading multiple images via GitHub Copilot's Claude models. Claude requires tool_use -> tool_result adjacency with no user messages interleaved. Images from consecutive tool results are now batched into a single user message
[7.0.0] - 2026-01-21
Added
- Added usage tracking system with normalized schema for provider quota/limit endpoints
- Added Claude usage provider for 5-hour and 7-day quota windows
- Added GitHub Copilot usage provider for chat, completions, and premium requests
- Added Google Antigravity usage provider for model quota tracking
- Added Google Gemini CLI usage provider for tier-based quota monitoring
- Added OpenAI Codex usage provider for primary and secondary rate limit windows
- Added ZAI usage provider for token and request quota tracking
Changed
- Updated Claude usage provider to extract account identifiers from response headers
- Updated GitHub Copilot usage provider to include account identifiers in usage reports
- Updated Google Gemini CLI usage provider to handle missing reset time gracefully
Fixed
- Fixed GitHub Copilot usage provider to simplify token handling and improve reliability
- Fixed GitHub Copilot usage provider to properly resolve account identifiers for OAuth credentials
- Fixed API validation errors when sending empty user messages (resume with
.) across all providers: - Google Cloud Code Assist (google-shared.ts)
- OpenAI Responses API (openai-responses.ts)
- OpenAI Codex Responses API (openai-codex-responses.ts)
- Cursor (cursor.ts)
- Amazon Bedrock (amazon-bedrock.ts)
- Clamped OpenAI Codex reasoning effort "minimal" to "low" for gpt-5.2 models to avoid API errors
- Fixed GitHub Copilot usage fallback to internal quota endpoints when billing usage is unavailable
- Fixed GitHub Copilot usage metadata to include account identifiers for report dedupe
- Fixed Anthropic usage metadata extraction to include account identifiers when provided by the usage endpoint
- Fixed Gemini CLI usage windows to consistently label quota windows for display suppression
[6.9.69] - 2026-01-21
Added
- Added duration and time-to-first-token (ttft) metrics to all AI provider responses
- Added performance tracking for streaming responses across all providers
[6.9.0] - 2026-01-21
Removed
- Removed openai-codex provider exports from main package index
- Removed openai-codex prompt utilities and moved them inline
- Removed vitest configuration file
[6.8.4] - 2026-01-21
Changed
- Updated prompt caching strategy to follow Anthropic's recommended hierarchy
- Fixed token usage tracking to properly handle cumulative output tokens from message_delta events
- Improved message validation to filter out empty or invalid content blocks
- Increased OAuth callback timeout from 120 seconds to 120,000 milliseconds
[6.8.3] - 2026-01-21
Added
- Added
headersoption to all providers for custom request headers - Added
onPayloadhook to observe provider request payloads before sending - Added
strictResponsesPairingoption for Azure OpenAI Responses API compatibility - Added
originatoroption tologinOpenAICodexfor custom OAuth flow identification - Added per-request
headersandonPayloadhooks toStreamOptions - Added
originatoroption tologinOpenAICodex
Fixed
- Fixed tool call ID normalization for OpenAI Responses API cross-provider handoffs
- Skipped errored or aborted assistant messages during cross-provider transforms
- Detected AWS ECS/IRSA credentials for Bedrock authentication checks
- Detected AWS ECS/IRSA credentials for Bedrock authentication checks
- Normalized Responses API tool call IDs during handoffs and refreshed handoff tests
- Enforced strict tool call/result pairing for Azure OpenAI Responses API
- Skipped errored or aborted assistant messages during cross-provider transforms
Security
- Enhanced AWS credential detection to support ECS task roles and IRSA web identity tokens
[6.8.2] - 2026-01-21
Fixed
- Improved error handling for aborted requests in Google Gemini CLI provider
- Enhanced OAuth callback flow to handle manual input errors gracefully
- Fixed login cancellation handling in GitHub Copilot OAuth flow
- Removed fallback manual input from OpenAI Codex OAuth flow
Security
- Hardened database file permissions to prevent credential leakage
- Set secure directory permissions (0o700) for credential storage
[6.8.0] - 2026-01-20
Added
- Added
logoutcommand to CLI for OAuth provider logout - Added
statuscommand to show logged-in providers and token expiry - Added persistent credential storage using SQLite database
- Added OAuth callback server with automatic port fallback
- Added HTML callback page with success/error states
- Added support for Cursor OAuth provider
Changed
- Updated Promise.withResolvers usage for better compatibility
- Replaced custom sleep implementations with Bun.sleep and abortableSleep
- Simplified SSE stream parsing using readLines utility
- Updated test framework from vitest to bun:test
- Replaced temp directory creation with createTempDirSync utility
- Changed credential storage from auth.json to ~/.omp/agent/agent.db
- Changed CLI command examples from npx to bunx
- Refactored OAuth flows to use common callback server base class
- Updated OAuth provider interfaces to use controller pattern
Fixed
- Fixed OAuth callback handling with improved error states
- Fixed token refresh for all OAuth providers
[6.7.670] - 2026-01-19
Changed
- Updated Claude Code compatibility headers and version
- Improved OAuth token handling with proper state generation
- Enhanced cache control for tool and user message blocks
- Simplified tool name prefixing for OAuth traffic
- Updated PKCE verifier generation for better security
[5.7.67] - 2026-01-18
Fixed
- Added error handling for unknown OAuth providers
[5.6.77] - 2026-01-18
Fixed
- Prevented duplicate tool results for errored or aborted messages when results already exist
[5.6.7] - 2026-01-18
Added
- Added automatic retry logic for OpenAI Codex responses with configurable delay and max retries
- Added tool call ID sanitization for Amazon Bedrock to ensure valid characters
- Added tool argument validation that coerces JSON-encoded strings for expected non-string types
Changed
- Updated environment variable prefix from PI_ to OMP_ for better consistency
- Added automatic migration for legacy PI_ environment variables to OMP_ equivalents
- Adjusted Bedrock Claude thinking budgets to reserve output tokens when maxTokens is too low
Fixed
- Fixed orphaned tool call handling to ensure proper tool_use/tool_result pairing for all assistant messages
- Fixed message transformation to insert synthetic tool results for errored/aborted assistant messages with tool calls
- Fixed tool prefix handling in Claude provider to use case-insensitive comparison
- Fixed Gemini 3 model handling to treat unsigned tool calls as context-only with anti-mimicry context
- Fixed message transformation to filter out empty error messages from conversation history
- Fixed OpenAI completions provider compatibility detection to use provider metadata
- Fixed OpenAI completions provider to avoid using developer role for opencode provider
- Fixed orphaned tool call handling to skip synthetic results for errored assistant messages
[5.5.0] - 2026-01-18
Changed
- Updated User-Agent header from 'opencode' to 'pi' for OpenAI Codex requests
- Simplified Codex system prompt instructions
- Removed bridge text override from Codex system prompt builder
[5.3.0] - 2026-01-15
Changed
- Replaced detailed Codex system instructions with simplified pi assistant instructions
- Updated internal documentation references to use pi-internal:// protocol
[5.1.0] - 2026-01-14
Added
- Added Amazon Bedrock provider with
bedrock-converse-streamAPI for Claude models via AWS - Added MiniMax provider with OpenAI-compatible API
- Added EU cross-region inference model variants for Claude models on Bedrock
Fixed
- Fixed Gemini CLI provider retries with proper error handling, retry delays from headers, and empty stream retry logic
- Fixed numbered list items showing "1." for all items when code blocks break list continuity (via
startproperty)
[5.0.0] - 2026-01-12
Added
- Added support for
xhighthinking level inthinkingBudgetsconfiguration
Changed
- Changed Anthropic thinking token budgets: minimal (1024→3072), low (2048→6144), medium (8192→12288), high (16384→24576)
- Changed Google thinking token budgets: minimal (1024), low (2048→4096), medium (8192), high (16384), xhigh (24575)
- Changed
supportsXhigh()to return true for all Anthropic models
[4.6.0] - 2026-01-12
Fixed
- Fixed incorrect classification of thought signatures in Google Gemini responses—thought signatures are now correctly treated as metadata rather than thinking content indicators
- Fixed thought signature handling in Google Gemini CLI and Vertex AI streaming to properly preserve signatures across text deltas
- Fixed Google schema sanitization stripping property names that match schema keywords (e.g., "pattern", "format") from tool definitions
[4.4.9] - 2026-01-12
Fixed
- Fixed Google provider schema sanitization to strip additional unsupported JSON Schema fields (patternProperties, additionalProperties, min/max constraints, pattern, format)
[4.4.8] - 2026-01-12
Fixed
- Fixed Google provider schema sanitization to properly collapse
anyOf/oneOfwith const values into enum arrays - Fixed const-to-enum conversion to infer type from the const value when type is not specified
[4.4.6] - 2026-01-11
Fixed
- Fixed tool parameter schema sanitization to only apply Google-specific transformations for Gemini models, preserving original schemas for other model types
[4.4.5] - 2026-01-11
Changed
- Exported
sanitizeSchemaForGoogleutility function for external use
Fixed
- Fixed Google provider schema sanitization to strip additional unsupported JSON Schema fields ($schema, $ref, $defs, format, examples, and others)
- Fixed Google provider to ignore
additionalProperties: falsewhich is unsupported by the API
[4.4.4] - 2026-01-11
Fixed
- Fixed Cursor todo updates to bridge update_todos tool calls to the local todo_write tool
[4.3.0] - 2026-01-11
Added
- Added debug log filtering and display script for Cursor JSONL logs with follow mode and coalescing support
- Added protobuf definition extractor script to reconstruct .proto files from bundled JavaScript
- Added conversation state caching to persist context across multiple Cursor API requests in the same session
- Added shell streaming support for real-time stdout/stderr output during command execution
- Added JSON5 parsing for MCP tool arguments with Python-style boolean and None value normalization
- Added Cursor provider with support for Claude, GPT, and Gemini models via Cursor's agent API
- Added OAuth authentication flow for Cursor including login, token refresh, and expiry detection
- Added
cursor-agentAPI type with streaming support and tool execution handlers - Added Cursor model definitions including Claude 4.5, GPT-5.x, Gemini 3, and Grok variants
- Added model generation script to automatically fetch and update AI model definitions from models.dev and OpenRouter APIs
Changed
- Changed Cursor debug logging to use structured JSONL format with automatic MCP argument decoding
- Changed MCP tool argument decoding to use protobuf Value schema for improved type handling
- Changed tool advertisement to filter Cursor native tools (bash, read, write, delete, ls, grep, lsp) instead of only exposing mcp_ prefixed tools
Fixed
- Fixed Cursor conversation history serialization so subagents retain task context and can call complete
[4.2.1] - 2026-01-11
Changed
- Updated
reasoningSummaryoption to accept only"auto","concise","detailed", ornull(removed"off"and"on"values) - Changed default
reasoningSummaryfrom"auto"to"detailed" - OpenAI Codex: switched to bundled system prompt matching opencode, changed originator to "opencode", simplified prompt handling
Fixed
- Fixed Cloud Code Assist tool schema conversion to avoid unsupported
constfields
[4.0.0] - 2026-01-10
Added
- Added
betasoption inAnthropicOptionsfor passing custom Anthropic beta feature flags - OpenCode Zen provider support with 26 models (Claude, GPT, Gemini, Grok, Kimi, GLM, Qwen, etc.). Set
OPENCODE_API_KEYenv var to use. thinkingBudgetsoption inSimpleStreamOptionsfor customizing token budgets per thinking level on token-based providerssessionIdoption inStreamOptionsfor providers that support session-based caching. OpenAI Codex provider uses this to setprompt_cache_keyand routing headers.supportsUsageInStreamingcompatibility flag for OpenAI-compatible providers that rejectstream_options: { include_usage: true }. Defaults totrue. Set tofalsein model config for providers like gatewayz.ai.GOOGLE_APPLICATION_CREDENTIALSenv var support for Vertex AI credential detection (standard for CI/production)- Exported OpenAI Codex utilities:
CacheMetadata,getCodexInstructions,getModelFamily,ModelFamily,buildCodexPiBridge,buildCodexSystemPrompt,CodexSystemPrompt - Headless OAuth support for all callback-server providers (Google Gemini CLI, Antigravity, OpenAI Codex): paste redirect URL when browser callback is unreachable
- Cancellable GitHub Copilot device code polling via AbortSignal
- Improved error messages for OpenRouter providers by including raw metadata from upstream errors
Changed
- Changed Anthropic provider to include Claude Code system instruction for all API key types, not just OAuth tokens (except Haiku models)
- Changed Anthropic OAuth tool naming to use
proxy_prefix instead of mapping to Claude Code tool names, avoiding potential name collisions - Changed Anthropic provider to include Claude Code headers for all requests, not just OAuth tokens
- Anthropic provider now maps tool names to Claude Code's exact tool names (Read, Write, Edit, Bash, Grep, Glob) instead of using prefixed names
- OpenAI Completions provider now disables strict mode on tools to allow optional parameters without null unions
Fixed
- Fixed Anthropic OAuth code parsing to accept full redirect URLs in addition to raw authorization codes
- Fixed Anthropic token refresh to preserve existing refresh token when server doesn't return a new one
- Fixed thinking mode being enabled when tool_choice forces a specific tool, which is unsupported
- Fixed max_tokens being too low when thinking budget is set, now auto-adjusts to model's maxTokens
- Google Cloud Code Assist OAuth for paid subscriptions: properly handles long-running operations for project provisioning, supports
GOOGLE_CLOUD_PROJECT/GOOGLE_CLOUD_PROJECT_IDenv vars for paid tiers os.homedir()calls at module load time; now resolved lazily when needed- OpenAI Responses tool strict flag to use a boolean for LM Studio compatibility
- Gemini CLI abort handling: detect native
AbortErrorin retry catch block, cancel SSE reader when abort signal fires - Antigravity provider 429 errors by aligning request payload with CLIProxyAPI v6.6.89
- Thinking block handling for cross-model conversations: thinking blocks are now converted to plain text when switching models
- OpenAI Codex context window from 400,000 to 272,000 tokens to match Codex CLI defaults
- Codex SSE error events to surface message, code, and status
- Context overflow detection for
context_length_exceedederror codes - Codex provider now always includes
reasoning.encrypted_contenteven when customincludeoptions are passed - Codex requests now omit the
reasoningfield entirely when thinking is off - Crash when pasting text with trailing whitespace exceeding terminal width
[3.37.1] - 2026-01-10
Added
- Added automatic type coercion for tool arguments when LLMs return JSON-encoded strings instead of native types (numbers, booleans, arrays, objects)
Changed
- Changed tool argument validation to attempt JSON parsing and type coercion before rejecting mismatched types
- Changed validation error messages to include both original and normalized arguments when coercion was attempted
[3.37.0] - 2026-01-10
Changed
- Enabled type coercion in JSON schema validation to automatically convert compatible types
[3.35.0] - 2026-01-09
Added
- Enhanced error messages to include retry-after timing information from API rate limit headers
[3.20.0] - 2026-01-06
Added
- Added support for kwaipilot/kat-coder-pro model via OpenRouter
- Added OpenAI Codex responses provider with OAuth login support for ChatGPT Plus/Pro accounts
- Added Google Vertex AI provider (Gemini via Vertex) with Application Default Credentials support
Changed
- Updated model specifications including context windows, max tokens, and pricing for multiple OpenRouter models
Removed
- Removed alibaba/tongyi-deepresearch-30b-a3b:free model from OpenRouter
- Removed nousresearch/hermes-4-405b model from OpenRouter
- Removed tngtech/tng-r1t-chimera:free model from OpenRouter
[3.15.0] - 2026-01-05
Changed
- Made
isErrorfield optional inToolResultMessageinterface, defaulting to non-error state
[3.5.1337] - 2026-01-03
Added
- Added localhost URL detection for OpenAI-compatible provider auto-configuration
[1.337.1] - 2026-01-02
Changed
- Forked to @oh-my-pi scope with unified versioning across all packages
Fixed
- Gemini CLI rate limit handling: Added automatic retry with server-provided delay for 429 errors
[1.337.0] - 2026-01-02
Initial release under @oh-my-pi scope. See previous releases at badlogic/pi-mono.
[0.50.1] - 2026-01-26
Fixed
- Fixed OpenCode Zen model generation to exclude deprecated models (#970 by @DanielTatarkin)
[0.50.0] - 2026-01-26
Added
- Added OpenRouter provider routing support for custom models via
openRouterRoutingcompat field (#859 by @v01dpr1mr0s3) - Added
azure-openai-responsesprovider support for Azure OpenAI Responses API. (#890 by @markusylisiurunen) - Added HTTP proxy environment variable support for API requests (#942 by @haoqixu)
- Added
createAssistantMessageEventStream()factory function for use in extensions. - Added
resetApiProviders()to clear and re-register built-in API providers.
Changed
- Refactored API streaming dispatch to use an API registry with provider-owned
streamSimplemapping. - Moved environment API key resolution to
env-api-keys.tsand re-exported it from the package entrypoint. - Azure OpenAI Responses provider now uses base URL configuration with deployment-aware model mapping and no longer includes service tier handling.
Fixed
- Fixed Bun runtime detection for dynamic imports in browser-compatible modules (stream.ts, openai-codex-responses.ts, openai-codex.ts) (#922 by @dannote)
- Fixed streaming functions to use
model.apiinstead of hardcoded API types - Fixed Google providers to default tool call arguments to an empty object when omitted
- Fixed OpenAI Responses streaming to handle
arguments.doneevents on OpenAI-compatible endpoints (#917 by @williballenthin) - Fixed OpenAI Codex Responses tool strictness handling after the shared responses refactor
- Fixed Azure OpenAI Responses streaming to guard deltas before content parts and correct metadata and handoff gating
- Fixed OpenAI completions tool-result image batching after consecutive tool results (#902 by @terrorobe)
[0.49.3] - 2026-01-22
Added
- Added
headersoption toStreamOptionsfor custom HTTP headers in API requests. Supported by all providers except Amazon Bedrock (which uses AWS SDK auth). Headers are merged with provider defaults andmodel.headers, withoptions.headerstaking precedence. - Added
originatoroption tologinOpenAICodex()for custom OAuth client identification - Browser compatibility for pi-ai: replaced top-level Node.js imports with dynamic imports for browser environments (#873)
Fixed
- Fixed OpenAI Responses API 400 error "function_call without required reasoning item" when switching between models (same provider, different model). The fix omits the
idfield for function_calls from different models to avoid triggering OpenAI's reasoning/function_call pairing validation (#886)
[0.49.2] - 2026-01-19
Added
- Added AWS credential detection for ECS/Kubernetes environments:
AWS_CONTAINER_CREDENTIALS_RELATIVE_URI,AWS_CONTAINER_CREDENTIALS_FULL_URI,AWS_WEB_IDENTITY_TOKEN_FILE(#848)
Fixed
- Fixed OpenAI Responses 400 error "reasoning without following item" by skipping errored/aborted assistant messages entirely in transform-messages.ts (#838)
Removed
- Removed
strictResponsesPairingcompat option (no longer needed after the transform-messages fix)
[0.49.1] - 2026-01-18
Added
- Added
OpenAIResponsesCompatinterface withstrictResponsesPairingoption for Azure OpenAI Responses API, which requires strict reasoning/message pairing in history replay (#768 by @nicobako)
Changed
- Split
OpenAICompatintoOpenAICompletionsCompatandOpenAIResponsesCompatfor type-safe API-specific compat settings
Fixed
- Fixed tool call ID normalization for cross-provider handoffs (e.g., Codex to Antigravity Claude) (#821)
[0.49.0] - 2026-01-17
Changed
- OpenAI Codex responses now use the context system prompt directly in the instructions field.
Fixed
- Fixed orphaned tool results after errored assistant messages causing Codex API errors. When an assistant message has
stopReason: "error", its tool calls are now excluded from pending tool tracking, preventing synthetic tool results from being generated for calls that will be dropped by provider-specific converters. (#812) - Fixed Bedrock Claude max_tokens handling to always exceed thinking budget tokens, preventing compaction failures. (#797 by @pjtf93)
- Fixed Claude Code tool name normalization to match the Claude Code tool list case-insensitively and remove invalid mappings.
[0.48.0] - 2026-01-16
Fixed
- Fixed OpenAI-compatible provider feature detection to use
model.providerin addition to URL, allowing custom base URLs (e.g., proxies) to work correctly with provider-specific settings (#774) - Fixed Gemini 3 context loss when switching from providers without thought signatures: unsigned tool calls are now converted to text with anti-mimicry notes instead of being skipped
- Fixed string numbers in tool arguments not being coerced to numbers during validation (#786 by @dannote)
- Fixed Bedrock tool call IDs to use only alphanumeric characters, avoiding API errors from invalid characters (#781 by @pjtf93)
- Fixed empty error assistant messages (from 429/500 errors) breaking the tool_use to tool_result chain by filtering them in
transformMessages
[0.47.0] - 2026-01-16
Fixed
- Fixed OpenCode provider's
/v1endpoint to usesystemrole instead ofdeveloperrole, fixing400 Incorrect role informationerror for models usingopenai-completionsAPI (#755 by @melihmucuk) - Added retry logic to OpenAI Codex provider for transient errors (429, 5xx, connection failures). Uses exponential backoff with up to 3 retries. (#733)
[0.46.0] - 2026-01-15
Added
- Added MiniMax China (
minimax-cn) provider support (#725 by @tallshort) - Added
gpt-5.2-codexmodels for GitHub Copilot and OpenCode Zen providers (#734 by @aadishv)
Fixed
- Avoid unsigned Gemini 3 tool calls (#741 by @roshanasingh4)
- Fixed signature support for non-Anthropic models in Amazon Bedrock provider (#727 by @unexge)
[0.45.7] - 2026-01-13
Fixed
- Fixed OpenAI Responses timeout option handling (#706 by @markusylisiurunen)
- Fixed Bedrock tool call conversion to apply message transforms (#707 by @pjtf93)
[0.45.6] - 2026-01-13
Fixed
- Export
parseStreamingJsonfrom main package for tsx dev mode compatibility
[0.45.4] - 2026-01-13
Added
- Added Vercel AI Gateway provider with model discovery and
AI_GATEWAY_API_KEYenv support (#689 by @timolins)
Fixed
- Fixed z.ai thinking/reasoning: z.ai uses
thinking: { type: "enabled" }instead of OpenAI'sreasoning_effort. AddedthinkingFormatcompat flag to handle this. (#688)
[0.45.0] - 2026-01-13
Added
- MiniMax provider support with M2 and M2.1 models via Anthropic-compatible API (#656 by @dannote)
- Add Amazon Bedrock provider with prompt caching for Claude models (experimental, tested with Anthropic Claude models only) (#494 by @unexge)
- Added
serviceTieroption for OpenAI Responses requests (#672 by @markusylisiurunen) - Anthropic caching on OpenRouter: Interactions with Anthropic models via OpenRouter now set a 5-minute cache point using Anthropic-style
cache_controlbreakpoints on the last assistant or user message. (#584 by @nathyong) - Google Gemini CLI provider improvements: Added Antigravity endpoint fallback (tries daily sandbox then prod when
baseUrlis unset), header-based retry delay parsing (Retry-After,x-ratelimit-reset,x-ratelimit-reset-after), stablesessionIdderivation from first user message for cache affinity, empty SSE stream retry with backoff, andanthropic-betaheader for Claude thinking models (#670 by @kim0)
[0.43.0] - 2026-01-11
Fixed
- Fixed Google provider thinking detection:
isThinkingPart()now only checksthought === true, notthoughtSignature. Per Google docs,thoughtSignatureis for context replay and can appear on any part type. Also removedidfield fromfunctionCall/functionResponse(rejected by Vertex AI and Cloud Code Assist), and addedtextSignatureround-trip for multi-turn reasoning context. (#631 by @theBucky)
[0.42.3] - 2026-01-10
Changed
- OpenAI Codex: switched to bundled system prompt matching opencode, changed originator to "pi", simplified prompt handling
[0.42.2] - 2026-01-10
Added
- Added
GOOGLE_APPLICATION_CREDENTIALSenv var support for Vertex AI credential detection (standard for CI/production). - Added
supportsUsageInStreamingcompatibility flag for OpenAI-compatible providers that rejectstream_options: { include_usage: true }. Defaults totrue. Set tofalsein model config for providers like gatewayz.ai. (#596 by @XesGaDeus) - Improved Google model pricing info (#588 by @aadishv)
Fixed
- Fixed
os.homedir()calls at module load time; now resolved lazily when needed. - Fixed OpenAI Responses tool strict flag to use a boolean for LM Studio compatibility (#598 by @gnattu)
- Fixed Google Cloud Code Assist OAuth for paid subscriptions: properly handles long-running operations for project provisioning, supports
GOOGLE_CLOUD_PROJECT/GOOGLE_CLOUD_PROJECT_IDenv vars for paid tiers, and handles VPC-SC affected users (#582 by @cmf)
[0.42.0] - 2026-01-09
Added
- Added OpenCode Zen provider support with 26 models (Claude, GPT, Gemini, Grok, Kimi, GLM, Qwen, etc.). Set
OPENCODE_API_KEYenv var to use.
[0.39.0] - 2026-01-08
Fixed
- Fixed Gemini CLI abort handling: detect native
AbortErrorin retry catch block, cancel SSE reader when abort signal fires (#568 by @tmustier) - Fixed Antigravity provider 429 errors by aligning request payload with CLIProxyAPI v6.6.89: inject Antigravity system instruction with
role: "user", setrequestType: "agent", and useantigravityuserAgent. Added bridge prompt to override Antigravity behavior (identity, paths, web dev guidelines) with Pi defaults. (#571 by @ben-vargas) - Fixed thinking block handling for cross-model conversations: thinking blocks are now converted to plain text (no
<thinking>tags) when switching models. Previously,<thinking>tags caused models to mimic the pattern and output literal tags. Also fixed empty thinking blocks causing API errors. (#561)
[0.38.0] - 2026-01-08
Added
thinkingBudgetsoption inSimpleStreamOptionsfor customizing token budgets per thinking level on token-based providers (#529 by @melihmucuk)
Breaking Changes
- Removed OpenAI Codex model aliases (
gpt-5,gpt-5-mini,gpt-5-nano,codex-mini-latest,gpt-5-codex,gpt-5.1-codex,gpt-5.1-chat-latest). Use canonical model IDs:gpt-5.1,gpt-5.1-codex-max,gpt-5.1-codex-mini,gpt-5.2,gpt-5.2-codex. (#536 by @ghoulr)
Fixed
- Fixed OpenAI Codex context window from 400,000 to 272,000 tokens to match Codex CLI defaults and prevent 400 errors. (#536 by @ghoulr)
- Fixed Codex SSE error events to surface message, code, and status. (#551 by @tmustier)
- Fixed context overflow detection for
context_length_exceedederror codes.
[0.37.6] - 2026-01-06
Added
- Exported OpenAI Codex utilities:
CacheMetadata,getCodexInstructions,getModelFamily,ModelFamily,buildCodexPiBridge,buildCodexSystemPrompt,CodexSystemPrompt(#510 by @mitsuhiko)
[0.37.3] - 2026-01-06
Added
sessionIdoption inStreamOptionsfor providers that support session-based caching. OpenAI Codex provider uses this to setprompt_cache_keyand routing headers.
[0.37.2] - 2026-01-05
Fixed
- Codex provider now always includes
reasoning.encrypted_contenteven when customincludeoptions are passed (#484 by @kim0)
[0.37.0] - 2026-01-05
Breaking Changes
- OpenAI Codex models no longer have per-thinking-level variants (e.g.,
gpt-5.2-codex-high). Use the base model ID and set thinking level separately. The Codex provider clamps reasoning effort to what each model supports internally. (initial implementation by @ben-vargas in #472)
Added
- Headless OAuth support for all callback-server providers (Google Gemini CLI, Antigravity, OpenAI Codex): paste redirect URL when browser callback is unreachable (#428 by @ben-vargas, #468 by @crcatala)
- Cancellable GitHub Copilot device code polling via AbortSignal
Fixed
- Codex requests now omit the
reasoningfield entirely when thinking is off, letting the backend use its default instead of forcing a value. (#472)
[0.36.0] - 2026-01-05
Added
- OpenAI Codex OAuth provider with Responses API streaming support:
openai-codex-responsesstreaming provider with SSE parsing, tool-call handling, usage/cost tracking, and PKCE OAuth flow (#451 by @kim0)
Fixed
- Vertex AI dummy value for
getEnvApiKey(): Returns"<authenticated>"when Application Default Credentials are configured (~/.config/gcloud/application_default_credentials.jsonexists) and bothGOOGLE_CLOUD_PROJECT(orGCLOUD_PROJECT) andGOOGLE_CLOUD_LOCATIONare set. This allowsstreamSimple()to work with Vertex AI without explicitapiKeyoption. The ADC credentials file existence check is cached per-process to avoid repeated filesystem access.
[0.32.3] - 2026-01-03
Fixed
- Google Vertex AI models no longer appear in available models list without explicit authentication. Previously,
getEnvApiKey()returned a dummy value forgoogle-vertex, causing models to show up even when Google Cloud ADC was not configured.
[0.32.0] - 2026-01-03
Added
- Vertex AI provider with ADC (Application Default Credentials) support. Authenticate with
gcloud auth application-default login, setGOOGLE_CLOUD_PROJECTandGOOGLE_CLOUD_LOCATION, and access Gemini models via Vertex AI. (#300 by @default-anton)
Fixed
- Gemini CLI rate limit handling: Added automatic retry with server-provided delay for 429 errors. Parses delay from error messages like "Your quota will reset after 39s" and waits accordingly. Falls back to exponential backoff for other transient errors. (#370)
[0.31.0] - 2026-01-02
Breaking Changes
- Agent API moved: All agent functionality (
agentLoop,agentLoopContinue,AgentContext,AgentEvent,AgentTool,AgentToolResult, etc.) has moved to@oh-my-pi/pi-agent-core. Import from that package instead of@oh-my-pi/pi-ai.
Added
GoogleThinkingLeveltype: Exported type that mirrors Google'sThinkingLevelenum values ("THINKING_LEVEL_UNSPECIFIED" | "MINIMAL" | "LOW" | "MEDIUM" | "HIGH"). Allows configuring Gemini thinking levels without importing from@google/genai.ANTHROPIC_OAUTH_TOKENenv var: Now checked beforeANTHROPIC_API_KEYingetEnvApiKey(), allowing OAuth tokens to take precedence.event-stream.jsexport:AssistantMessageEventStreamutility now exported from package index.
Changed
- OAuth uses Web Crypto API: PKCE generation and OAuth flows now use Web Crypto API (
crypto.subtle) instead of Node.jscryptomodule. This improves browser compatibility while still working in Node.js 20+. - Deterministic model generation:
generate-models.tsnow sorts providers and models alphabetically for consistent output across runs. (#332 by @mrexodia)
Fixed
- OpenAI completions empty content blocks: Empty text or thinking blocks in assistant messages are now filtered out before sending to the OpenAI completions API, preventing validation errors. (#344 by @default-anton)
- zAi provider API mapping: Fixed zAi models to use
openai-completionsAPI with correct base URL (https://api.z.ai/api/coding/paas/v4) instead of incorrect Anthropic API mapping. (#344, #358 by @default-anton)
[0.28.0] - 2025-12-25
Breaking Changes
- OAuth storage removed (#296): All storage functions (
loadOAuthCredentials,saveOAuthCredentials,setOAuthStorage, etc.) removed. Callers are responsible for storing credentials. - OAuth login functions:
loginAnthropic,loginGitHubCopilot,loginGeminiCli,loginAntigravitynow returnOAuthCredentialsinstead of saving to disk. - refreshOAuthToken: Now takes
(provider, credentials)and returns newOAuthCredentialsinstead of saving. - getOAuthApiKey: Now takes
(provider, credentials)and returns{ newCredentials, apiKey }or null. - OAuthCredentials type: No longer includes
type: "oauth"discriminator. Callers add discriminator when storing. - setApiKey, resolveApiKey: Removed. Callers must manage their own API key storage/resolution.
- getApiKey: Renamed to
getEnvApiKey. Only checks environment variables for known providers.
[0.27.7] - 2025-12-24
Fixed
- Thinking tag leakage: Fixed Claude mimicking literal
</thinking>tags in responses. Unsigned thinking blocks (from aborted streams) are now converted to plain text without<thinking>tags. The TUI still displays them as thinking blocks. (#302 by @nicobailon)
[0.25.1] - 2025-12-21
Added
- xhigh thinking level support: Added
supportsXhigh()function to check if a model supports xhigh reasoning level. Also clamps xhigh to high for OpenAI models that don't support it. (#236 by @theBucky)
Fixed
- Gemini multimodal tool results: Fixed images in tool results causing flaky/broken responses with Gemini models. For Gemini 3, images are now nested inside
functionResponse.partsper the docs. For older models (which don't support multimodal function responses), images are sent in a separate user message. - Queued message steering: When
getQueuedMessagesis provided, the agent loop now checks for queued user messages after each tool call and skips remaining tool calls in the current assistant message when a queued message arrives (emitting error tool results). - Double API version path in Google provider URL: Fixed Gemini API calls returning 404 after baseUrl support was added. The SDK was appending its default apiVersion to baseUrl which already included the version path. (#251 by @shellfyred)
- Anthropic SDK retries disabled: Re-enabled SDK-level retries (default 2) for transient HTTP failures. (#252)
[0.23.5] - 2025-12-19
Added
- Gemini 3 Flash thinking support: Extended thinking level support for Gemini 3 Flash models (MINIMAL, LOW, MEDIUM, HIGH) to match Pro models' capabilities. (#212 by @markusylisiurunen)
- GitHub Copilot thinking models: Added thinking support for additional Copilot models (o3-mini, o1-mini, o1-preview). (#234 by @aadishv)
Fixed
- Gemini tool result format: Fixed tool result format for Gemini 3 Flash Preview which strictly requires
{ output: value }for success and{ error: value }for errors. Previous format using{ result, isError }was rejected by newer Gemini models. Also improved type safety by removingas anycasts. (#213, #220) - Google baseUrl configuration: Google provider now respects
baseUrlconfiguration for custom endpoints or API proxies. (#216, #221 by @theBucky) - GitHub Copilot vision requests: Added
Copilot-Vision-Requestheader when sending images to GitHub Copilot models. (#222) - GitHub Copilot X-Initiator header: Fixed X-Initiator logic to check last message role instead of any message in history. This ensures proper billing when users send follow-up messages. (#209)
[0.22.3] - 2025-12-16
Added
- Image limits test suite: Added comprehensive tests for provider-specific image limitations (max images, max size, max dimensions). Discovered actual limits: Anthropic (100 images, 5MB, 8000px), OpenAI (500 images, ≥25MB), Gemini (~2500 images, ≥40MB), Mistral (8 images, ~15MB), OpenRouter (~40 images context-limited, ~15MB). (#120)
- Tool result streaming: Added
tool_execution_updateevent and optionalonUpdatecallback toAgentTool.execute()for streaming tool output during execution. Tools can now emit partial results (e.g., bash stdout) that are forwarded to subscribers. (#44) - X-Initiator header for GitHub Copilot: Added X-Initiator header handling for GitHub Copilot provider to ensure correct call accounting (agent calls are not deducted from quota). Sets initiator based on last message role. (#200 by @kim0)
Changed
- Normalized tool_execution_end result:
tool_execution_endevent now always containsAgentToolResult(no longerAgentToolResult | string). Errors are wrapped in the standard result format.
Fixed
- Reasoning disabled by default: When
reasoningoption is not specified, thinking is now explicitly disabled for all providers. Previously, some providers like Gemini with "dynamic thinking" would use their default (thinking ON), causing unexpected token usage. This was the original intended behavior. (#180 by @markusylisiurunen)
[0.22.2] - 2025-12-15
Added
- Interleaved thinking for Anthropic: Added
interleavedThinkingoption toAnthropicOptions. When enabled, Claude 4 models can think between tool calls and reason after receiving tool results. Enabled by default (no extra token cost, just unlocks the capability). SetinterleavedThinking: falseto disable.
[0.22.1] - 2025-12-15
Dedicated to Peter's shoulder (@steipete)
Added
- Interleaved thinking for Anthropic: Enabled interleaved thinking in the Anthropic provider, allowing Claude models to output thinking blocks interspersed with text responses.
[0.22.0] - 2025-12-15
Added
- GitHub Copilot provider: Added
github-copilotas a known provider with models sourced from models.dev. Includes Claude, GPT, Gemini, Grok, and other models available through GitHub Copilot. (#191 by @cau1k)
Fixed
- GitHub Copilot gpt-5 models: Fixed API selection for gpt-5 models to use
openai-responsesinstead ofopenai-completions(gpt-5 models are not accessible via completions endpoint) - GitHub Copilot cross-model context handoff: Fixed context handoff failing when switching between GitHub Copilot models using different APIs (e.g., gpt-5 to claude-sonnet-4). Tool call IDs from OpenAI Responses API were incompatible with other models. (#198)
- Gemini 3 Pro thinking levels: Thinking level configuration now works correctly for Gemini 3 Pro models. Previously all levels mapped to -1 (minimal thinking). Now LOW/MEDIUM/HIGH properly control test-time computation. (#176 by @markusylisiurunen)
[0.18.2] - 2025-12-11
Changed
- Anthropic SDK retries disabled: Set
maxRetries: 0on Anthropic client to allow application-level retry handling. The SDK's built-in retries were interfering with coding-agent's retry logic. (#157)
[0.18.1] - 2025-12-10
Added
- Mistral provider: Added support for Mistral AI models via the OpenAI-compatible API. Includes automatic handling of Mistral-specific requirements (tool call ID format). Set
MISTRAL_API_KEYenvironment variable to use.
Fixed
- Fixed Mistral 400 errors after aborted assistant messages by skipping empty assistant messages (no content, no tool calls) (#165)
- Removed synthetic assistant bridge message after tool results for Mistral (no longer required as of Dec 2025) (#165)
- Fixed bug where
ANTHROPIC_API_KEYenvironment variable was deleted globally after first OAuth token usage, causing subsequent prompts to fail (#164)
[0.17.0] - 2025-12-09
Added
agentLoopContinuefunction: Continue an agent loop from existing context without adding a new user message. Validates that the last message isuserortoolResult. Useful for retry after context overflow or resuming from manually-added tool results.- Added
validateToolCall(tools, toolCall)helper that finds the tool by name and validates arguments. - OpenAI compatibility overrides: Added
compatfield toModelforopenai-completionsAPI, allowing explicit configuration of provider quirks (supportsStore,supportsDeveloperRole,supportsReasoningEffort,maxTokensField). Falls back to URL-based detection if not set. Useful for LiteLLM, custom proxies, and other non-standard endpoints. (#133, thanks @fink-andreas for the initial idea and PR) - xhigh reasoning level: Added
xhightoReasoningEfforttype for OpenAI codex-max models. For non-OpenAI providers (Anthropic, Google),xhighis automatically mapped tohigh. (#143)
Breaking Changes
- Removed provider-level tool argument validation. Validation now happens in
agentLoopviaexecuteToolCalls, allowing models to retry on validation errors. For manual tool execution, usevalidateToolCall(tools, toolCall)orvalidateToolArguments(tool, toolCall).
Changed
- Updated SDK versions: OpenAI SDK 5.21.0 → 6.10.0, Anthropic SDK 0.61.0 → 0.71.2, Google GenAI SDK 1.30.0 → 1.31.0
[0.13.0] - 2025-12-06
Breaking Changes
- Added
totalTokensfield toUsagetype: All code that constructsUsageobjects must now include thetotalTokensfield. This field represents the total tokens processed by the LLM (input + output + cache). For OpenAI and Google, this uses native API values (total_tokens,totalTokenCount). For Anthropic, it's computed asinput + output + cacheRead + cacheWrite.
[0.12.10] - 2025-12-04
Added
- Added
gpt-5.1-codex-maxmodel support
Fixed
- OpenAI Token Counting: Fixed
usage.inputto exclude cached tokens for OpenAI providers. Previously,inputincluded cached tokens, causing double-counting when calculating total context size viainput + cacheRead. Nowinputrepresents non-cached input tokens across all providers, makinginput + output + cacheRead + cacheWritethe correct formula for total context size. - Fixed Claude Opus 4.5 cache pricing (was 3x too expensive)
- Corrected cache_read: $1.50 → $0.50 per MTok
- Corrected cache_write: $18.75 → $6.25 per MTok
- Added manual override in
scripts/generate-models.tsuntil upstream fix is merged - Submitted PR to models.dev: https://github.com/sst/models.dev/pull/439
[0.9.4] - 2025-11-26
Initial release with multi-provider LLM support.