- Normalized result-bearing Codex image items on terminal output events and emitted standard image content.
- Preserved result-bearing image calls during full Responses history replay despite stale provider status.
- Added stream and replay regressions for the Codex path.
Fixes#7445
The OpenAI service tier could only be chosen through the `tier.openai`
setting, or for a resumed session through whatever tier that session
recorded. Wanting flex or priority for a single run meant editing
settings and putting them back afterwards, while `bench` already took a
`--service-tier` flag that the session CLI did not offer.
Add `--service-tier` to the root command. The flag wins over the
configured setting and over a resumed session's recorded tier, leaves
the Anthropic and Google entries untouched, and records the resulting
map so a later resume keeps it. `none` removes the OpenAI entry, which
omits `service_tier` from the request.
Signed-off-by: Christian Stewart <christian@aperture.us>
- Stored replay-sanitized Codex response items as the append baseline.
- Disabled append state for responses without replayable output and covered oversized call IDs.
Fixes#7279
Routed Codex prompt cache identity through the shared retention-aware resolver while preserving transport session identity.
Covered explicit and environment-derived opt-outs, option precedence, direct body construction, and transport headers.
Fixes#7219
- Reserve the `code_mode_tool_names` metadata key to prevent caller-supplied client extra collisions.
- Preserve the `encrypted_function_args` plaintext-collaboration marker on replayed function calls.
Passed provider-specific PI_PROXY settings and standard HTTPS/ALL proxy variables to Bun WebSocket connections while preserving NO_PROXY bypasses.
Fixes#5384
Gave each Codex SSE transport attempt an independent pre-response watchdog while reserving the retry-loop signal for caller cancellation.
Added provider-level coverage for retry recovery and explicit abort behavior.
Fixes#5329
Codex #handleOutputItemDone unconditionally rebuilt block.thinking from
item.summary, clobbering already-streamed reasoning with "" when the
done event carries no summary (GPT-5.x Codex renders <!-- --> in TUI).
The generic processResponsesStream had the same clobber, and the codex
handler also dropped response.reasoning_text.delta events entirely.
Added shared finalizeReasoningThinking(): prefer done-item summary,
then reasoning_text content, then the streamed accumulation; wired
response.reasoning_text.delta into the codex processor.
Adopted from PR #4935 minus unrelated prompt churn.
Fixes#4918
Added Codex turn diagnostics that pair WebSocket/SSE request shape with raw provider usage and displayed usage buckets, including delta-chain state for previous_response_id requests.
Covered the compact/resume WebSocket delta path where raw usage lacks orchestration fields but reports a persistent uncached suffix.
Refs #4707
`isCodexStalePreviousResponseError` short-circuited for `CodexProviderStreamError`
by comparing `error.code` only to `previous_response_not_found`, so a proxy code
such as `codex_previous_response_stale` never reached the message-based
fallback. WebSocket continuations that hit a stale upstream response anchor
surfaced the terminal error to the user instead of retrying with full context.
Codex WebSocket continuations now treat both the OpenAI-standard
`previous_response_not_found` and the proxy `codex_previous_response_stale`
code as the same recovery class, and every `Error` — not just plain ones —
falls through to the existing `previous[ _]?response` / `expired|stale|...`
message check.
Fixes#4624
Made Codex account ids optional for openai-codex-responses custom providers, omitting chatgpt-account-id when opaque API keys cannot provide a ChatGPT account claim.
Added SSE and websocket regression coverage for custom Codex proxy API keys.
Fixes#4526
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.
Fixes#4469
- Added textVerbosity option to OpenAIResponsesOptions and SimpleStreamOptions.
- Updated stream mapping logic to propagate verbosity setting to the API request body.
- Implemented endpoint validation to ensure verbosity settings are only applied to official OpenAI endpoints.
- Added comprehensive test coverage for request payload inspection and stream event handling.
- Default the reasoning context to all_turns in the OpenAI Codex request transformer.
- Increase the default text verbosity for Codex requests to medium.
- Set a detailed reasoning summary for AI streams by default.
- Remove outdated testing logic for low verbosity defaults.
- Migrated configuration constants to use dynamic environment variable lookups.
- Removed unused helper functions and associated internal logic.
- Deleted comprehensive test suite for decommissioned WebSocket transport mechanisms.
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
- Centralized transient transport error patterns to enable consistent reuse across error handling modules.
- Refactored `isOpaqueStatusBody` for improved accessibility in classification logic.
- Updated retryable error detection to include the consolidated transport pattern and additional provider-specific error criteria.
- Implement ID-based guards for Codex WebSocket frames to reject stale or unauthorized interleaved frames from previous turns.
- Check sequence numbers within Codex responses to detect and handle out-of-order frame delivery.
- Restrict tool result consumption to occurrences appearing after the associated tool call, preventing the reuse of orphaned results from earlier conversation turns.
The Codex Responses backend (ChatGPT-subscription `/backend-api/codex/responses`)
returns 400 `{"detail":"Unsupported parameter: temperature"}` for every sampling
control, but `buildTransformedCodexRequestBody` forwarded `temperature`, `top_p`,
`top_k`, `min_p`, `presence_penalty`, and `repetition_penalty` straight through
whenever the caller's `StreamOptions` carried them — so any non-default sampling
setting failed every turn.
- Provider drops the full sampling set (matching codex-rs, which sends none).
- `RequestBody` no longer types fields it never populates for codex.
- Auth-gateway defensive strip in `buildStreamOptions` and the pi-native handler
widened from `{temperature, topP}` to the same set plus
`stopSequences`/`frequencyPenalty`.
- Regression test pins the captured body to omit every key even when the caller
sets all of them.
Fixes#3117
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
The keyed maps only see items whose `output_item.added` carries `item.id`
or `output_index`. A fully keyless add never reached either map, and the
unkeyed fallback was scanning those maps instead of the actual current
item — so `function_call_arguments.delta` / `output_item.done` for a
keyless tool call landed on null and the stored block kept `{}`. Mixed
streams also picked an older `output_index` entry before a later
id-only current item for the same reason.
`CodexStreamRuntime.currentEntry` now always points at the most recently
added `output_item.added` (whether or not it has keys). `openItemForEvent`
returns `currentEntry` when both `item_id` and `output_index` are absent,
and `closeCodexOpenItem` clears `currentEntry` (alongside the legacy
`currentItem` / `currentBlock` mirrors) when its item closes. The keyed
maps are unchanged so deliberate drop-on-mismatch for keyed events still
holds.
Regression tests cover the fully keyless function-call stream and the
mixed id-only vs `output_index`-only ordering. Existing reasoning/message
flow keeps singleton semantics through the same fallback.
Fixes#2619
Codex Responses function/custom tool call items can omit `id` while still
carrying `output_index` on `output_item.added` and `output_item.done`.
The first pass only keyed open items by `item.id`, so idless done events
could not find the stored block and the authoritative final arguments were
not copied into `output.content`.
Track open items by both `item.id` and `output_index`. Event lookup now
uses `item_id` first, then `output_index`, and only falls back to the
singleton-current path when neither key is present. Closing an item removes
both keys, so stale keyed deltas remain dropped instead of leaking into a
sibling.
Added a regression covering idless function and custom tool calls finalized
only by `output_item.done`, including out-of-order completion and per-call
`contentIndex`.
Fixes#2619
The Codex Responses stream runtime tracked a singleton
`currentItem`/`currentBlock`. With more than one tool call open
concurrently every `response.function_call_arguments.delta` was
appended to whichever item was added most recently, and the next
`response.output_item.done` for the earlier call overwrote the
sibling's stored arguments. On the agent loop's `task` tool this
surfaced as `tasks: Invalid input: expected array, received undefined`.
Open items are now tracked in a `Map<string, CodexOpenItem>` keyed by
`item.id`, each entry carrying its own `block` and `contentIndex`.
`response.function_call_arguments.{delta,done}`,
`response.custom_tool_call_input.{delta,done}`, and
`response.output_item.done` route through `openItemForEvent` and
operate on the matching entry's block; the legacy singleton-current
fallback only kicks in when an event omits `item_id`. A delta whose
keyed item already closed is dropped instead of leaking into a
sibling, and `toolcall_delta` / `toolcall_end` stream events emit the
right `contentIndex` for each call. Recovery sites converge on a new
`resetCodexStreamAccumulators` helper that clears the open-items map
in lockstep with `currentItem`/`currentBlock`/`nativeOutputItems`.
Regression test (`packages/ai/test/openai-codex-stream.test.ts`)
interleaves two function-call argument streams plus a stale
post-close delta and asserts per-call argument integrity and
per-call stream `contentIndex`.
Fixes#2619
- Removed stateful SSE chain controls from Codex provider options.
- Changed SSE transport to send full request bodies without `previous_response_id` deltas.
- Updated stale-chain recovery to treat `unsupported` as stale and recover via websocket.
- Scoped `x-codex-turn-state` to in-turn follow-ups and cleared turn state otherwise.
- Added `responsesLite`, `reasoningContext`, `clientMetadata`, and `onModerationMetadata` options.
- Changed Lite request transformation to default `reasoning.context` to `all_turns`.
- Changed request shaping to strip image `detail` fields and force `parallel_tool_calls` off in Lite mode.
- Fixed streaming behavior by excluding `client_metadata` from websocket append diffs and swallowing moderation metadata errors.
- Mapped Codex `end_turn:false` terminal events to `pause_turn` stop details in response stream parsing.
- Updated `agent-loop` to re-sample `pause_turn` turns, reset on tool calls, and cap continuations at 8.
- Added coverage for pause-turn mapping and continuation-capping in agent and AI stream tests.
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.