- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Added utility functions to strip descriptions from JSON schemas and tool definitions for optimized token output.
- Integrated `pruneToolDescriptions` configuration across agent loops and sessions to enable optional schema pruning.
- Updated agent context and snapshot logic to propagate pruning settings and maintain fingerprinting integrity.
- Verified schema structural integrity and removal of annotation descriptions through new unit tests.
Treat empty strings on optional tool arguments as omitted before schema validation so MCP calls do not fail pattern or type checks for model-filled placeholders.
Fixes#2981
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
- Prevents absolute `AbortSignal.timeout` from cutting off active stream bodies by introducing a clearable pre-response timer.
- Clears the watchdog timer for Bedrock, Gemini, Ollama, and Codex providers once headers are received.
- Restores regular caller abort signaling capability on the combined fetch stream.
- Adds comprehensive fake-timer tests to cover the timeout arming and clearing lifecycle.
- Added native support for parsing, normalized validation, and serialization of ArkType schemas throughout the agent pipeline.
- Implemented pruning of unconstrained union branches and normalization helpers to handle unrepresentable strict-mode branches.
- Patched ArkType's schema package to preserve declared object key order during serialization.
- Integrated ArkType schemas into `agent-loop` and migrated coding agent tool parameters to ArkType format.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Added Gemini thinking-loop detection helpers for near-duplicate and verbatim output checks.
- Wrapped `stream`, `streamPiNative`, and `streamSimple` dispatches with the loop guard.
- Emitted retryable empty-content loop errors and stopped completion events on loop hits.
- Added `enableGeminiThinkingLoopGuard` options for OpenAI compatibility with Gemini defaults and overrides.
- Enabled recursive parsing in `tryParseJsonForTypes` to coerce double-encoded JSON strings.
- Updated `coerceArgsFromIssues` to parse object/array JSON strings before singleton-array fallback.
- Prevented malformed container strings from being wrapped into arrays so validation returns array errors directly.
Rewrote unsupported lookaround patternProperties keys to a supported catch-all pattern instead of dropping their value schemas, preserving dynamic-key tool arguments when additionalProperties is false.
Extended sanitizer and Codex conversion regression coverage for the closed dynamic-key case.
Fixes#2784
Converted schema nodes emptied by OpenAI Responses lookaround stripping to boolean true so pattern-only nodes keep the existing empty-schema semantics.
Extended sanitizer and Codex regression coverage for pattern-only property and propertyNames schemas.
Fixes#2784
Removed JSON Schema pattern values containing regex lookaround from OpenAI Responses/Codex tool schemas so incompatible MCP tools do not poison the request.
Added schema-normalization and Codex conversion regression coverage for Figma-style fileKey patterns.
Fixes#2784
Recognized the empty Ollama finish_reason:length guidance as context overflow so coding-agent promotion and compaction recovery still run after surfacing the actionable error.
Added overflow utility coverage for the Ollama prompt-filled-context wording.
One MCP tool whose input schema can't be emitted as a valid strict tool schema
for the active provider made the whole request 400, so the assistant couldn't
respond at all (#2652). `convertTools` now validates each tool's emitted
parameter schema for enum/const-vs-type contradictions that pass structural
JSON-Schema validation but the provider rejects (a non-null enum on a
type:"null" node; an enum on an array node), and drops just the offending tool
with a `logger.warn` naming the tool + schema path, keeping the rest of the
request valid.
- New `findStrictToolSchemaViolation` in utils/schema: a semantic enum/const-vs-
type checker. The existing `isValidJsonSchema` is structural-only and accepts
these contradictions, which is exactly why they reach the provider.
- Tests: the three reported shapes (nullable-enum, enum-on-array, anyOf/const),
nested-path reporting, valid combinations incl. nullable unions, and the
convertTools quarantine (bad tool dropped, others survive).
- Dialect scanners for DeepSeek, Gemini, Gemma, GLM, Kimi, Pi, and Qwen3 now emit thinkingEnd on final chunks when a thinking state is still open.
- Markup healing was changed to return event streams with synthesized tool-call events removed, and OpenAI/Ollama streaming now uses this path when tool_calls are already structured.
- OpenAI completions streaming now suppresses healed thinking output when explicit reasoning content is present, with new tests covering unterminated and duplicate-thinking scenarios.
- Renamed ToolCallSyntax type to Dialect and Grammar interface to DialectDefinition across all packages.
- Moved grammar directory to dialect and updated all import paths in agent, ai, catalog, and coding-agent packages.
- Added renderTranscript and renderThinking methods to DialectDefinition, enabling native dialect-aware conversation serialization.
- Consolidated rendering utilities into new dialect/rendering.ts with shared helpers for ChatML, legacy text, and dialect-specific formatting.
- Updated conversation serialization in agent and coding-agent to use dialect.renderTranscript() for native turn envelope rendering.
- Added `jsonSchemaToTypeScript` and `renderToolInventory` to generate tool blocks with TypeScript signatures.
- Added `examples` and `TSchema` fields to dump-tool metadata and passed them through prompt rendering.
- Changed Harmony invocation rendering to omit `<|constrain|>json` markers in tool call payloads.
- Added compact native tool list-mode inventory rendering with full `# Tool:` output elsewhere.
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
Bun's fetch enforces a hard ~300s pre-response timeout that the caller's AbortSignal cannot lengthen. Every streaming provider's first-event/idle/SDK watchdog was silently capped by it, so cold large-context streams (multi-hundred-K prompts against slow-prefill backends) died at exactly 300s with TimeoutError.
- Added FetchWithRetryOptions.timeout (forwarded to fetch) so fetchWithRetry callers can pass Bun's timeout: false / numeric override directly.
- Passed timeout: false in openai-http, openai-codex-responses, amazon-bedrock, google-gemini-cli, ollama where each provider already constructs the fetch init.
- Added timeout to AnthropicFetchOptions and threaded timeout: false through buildAnthropicClientOptions's fetchOptions; anthropic-client already spreads fetchOptions into every fetch call.
- Allowed compat.streamIdleTimeoutMs: 0 in models.yml so the documented per-model disable knob matches the env-var escape hatch.
Verified with an out-of-tree smoke that stalls a local server 305s before responding: pre-fix the streaming provider died at ~300003ms with the Bun TimeoutError; post-fix the request completes (SUCCESS elapsed=305006ms). The smoke is discarded per directive.
Fixes#2422
- Replaced Azure and OpenAI provider calls with `postOpenAIStream` flow.
- Fixed stream error handling by retaining status, headers, and body on failures.
- Fixed stream parsing by handling raw JSON SSE frames and `[DONE]` events.
- Added `OpenAIHttpError` with parsed messages and timeout-aware retry behavior.
- Added remote-compaction tests for timeout, abort, and 500-fallback behavior.
- Added explicit request-debug path helpers in `packages/ai/src/utils/request-debug.ts` for one-shot dumps.
- Added `/debug dump-next-request`, `/debug dump-request`, and `/debug next-request` subcommands to arm next-request dumps.
- Changed `debug` handling in `packages/coding-agent/src/slash-commands/builtin-registry.ts` to execute args instead of always opening selector.
- Fixed explicit request-debug mode to resolve `~`/relative paths and create parent directories before logging.
- Fixed one-shot request-debug mode to consume its target after one call and overwrite existing response logs.
- Added a terminal-aware async iterator that enforces a post-completion grace window, then ends iteration and optionally aborts the request.
- Wrapped OpenAI completions streaming with terminal detection of finish-reason and usage payloads so iteration stops once a response is logically complete.
- Stopped OpenAI Responses stream consumption at terminal response events instead of waiting for connection close timeouts.
- Tracked whether strict tool schemas were actually applied and used that flag when deciding strict-to-nonstrict retries.
- Recorded strict-tool failures in session state so later requests skip strict mode retries and avoid extra error rounds.
- Hardened strict schema normalization by flattening nested pure anyOf unions and expanded schema error matching for additional invalid-schema rejections.
Only flatten optional anyOf schemas when the union wrapper has no sibling constraints that would exclude null.
Added a strict-schema regression for constrained anyOf properties so null remains an outer branch.
Flattened optional union tool schemas during strict enforcement so OpenRouter DeepSeek V4 no longer receives nested anyOf branches without a type.
Added schema-level and OpenRouter DeepSeek payload regressions for optional string-or-array tool parameters.
Fixes#2270
- Added catalog-level host/model predicates and compat resolvers.
- Extended compatibility types and model schema with timeout and replay flags.
- Replaced provider-specific heuristics with shared resolver-based checks.
- Updated host/identity and resolver tests to validate the new behavior.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
in-band Gemini errors, promptFeedback blocks, and missing finishReason no longer report success; toolUse override stops masking SAFETY/MALFORMED finishes; schema normalization keeps DAG-shared subtrees while detecting true cycles; Google/AWS shared credential resolution detached from first caller's signal and bounded by own timeout; Bedrock keeps toolConfig under toolChoice none; eventstream cancels body on abnormal exit.
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
- Derived descriptors, default-model map, env keys, login list, and refresh dispatch from one ProviderDefinition per provider.
- Disabled OpenAI Codex stream obfuscation and interrupted whitespace-only tool-call argument deltas.
- Derived auth-broker callback ports and paste-code login set from the registry.
- Halved input, output, and cache-read costs for one model.
- Adjusted pricing and lowered maxTokens for two other models.
- Removed extra blank lines in api-key-validation.
- Validated keys against the `/v1/messages` endpoint with base URL normalization.
- Added `anthropic-messages` strategy to API key login config.
- Emitted `message_stop` in stream timeout test fixtures.
- Scoped Anthropic provider session state by base URL and model so strict-tools/fast-mode fallback state no longer leaks across unrelated endpoints or models, and fast-mode re-arming clears all matching scoped entries.
- Dropped stale strict fallback error messages after successful retries and only enabled adaptive thinking display for models that advertise support, avoiding unsupported-model 400 errors.
- Updated idle stream iteration to close the upstream iterator on non-continuing exits (including consumer early break) and added coverage for upstream closure behavior.
Zhipu Coding Plan keys are formatted '<id>.<secret>' (no 'sk-' prefix),
so the OAuth login prompt's 'sk-...' placeholder misled users into
thinking their pasted key was malformed.
Also extended the SECRET_PATTERNS redaction example in
docs/skills/authoring-hooks.md with the Zhipu shape and a comment noting
the list is not exhaustive, so hook authors referencing the example do
not silently miss Zhipu keys.
Fixes#2106
- Added reconstructed SSE event emission for OpenAI, Azure, and Anthropic streams.
- Added raw SSE text to debug report bundles, including raw-sse.txt output.
- Included dropped-record metadata in raw SSE text when events were trimmed.
- Updated raw SSE and sse-debug tests for observer-based capture and safety checks.
Promoted MiniMax model/provider detection into stream markup healing so OpenCode Zen MiniMax streams use the existing thinking-tag parser instead of exposing raw tags.
Added a regression test for OpenCode Zen minimax-m3 chunks split across <think> boundaries.
Fixes#2049
Only mark Zod and JSON Schema validation issues as union-branch when their own path matches the combinator's path so nested array fields inside a tag-selected branch keep their singleton-wrap repair.
Fixes#2026