Some providers (notably Gemini) serialize array tool arguments with
flattened property paths — questions[0].id, questions[0].options[0].label —
instead of a nested questions array. The schema sees only unrecognized extra
keys and rejects the call (e.g. the ask tool).
Add a pre-validation normalization pass (alongside the existing LLM-quirk
passes) that rebuilds the nested structure. Conservative: fires only when a
key is a well-formed array-index path, preserves non-flattened siblings, and
aborts wholesale on any shape conflict so genuine schema mistakes still
surface as validation errors.
Fixes#8886
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
- Added healing mechanisms in the GLM dialect to detect and repair missing or incorrect in-band value-closer tags during stream scanning.
- Introduced argument spill recovery utilities to isolate and extract native tool arguments contaminated by leaked in-band syntax.
- Updated validation logic to perform automated argument repair and iterative coercion passes after initial validation failures.
- Added comprehensive test suites to verify argument recovery, schema coercion, and preservation of valid prose.
Resolved local JSON Schema refs before trimming enum/const string values so plain JSON Schema tools using or definitions get the same normalization path as inline enum schemas.
Added regression coverage for a referenced enum and a legacy definitions const.
Fixes#4461
Extended the pre-dispatch whitespace strip from path/URL keys to also cover the display-label keys 'title' and 'label', so tools like eval no longer surface stray trailing newlines the model emits on the title field. Content-carrying fields (content, input, code, command, body, text) still keep their trailing whitespace.
Added an eval-tool regression using the actual ArkType enum shape to prove the language field trim survives ArkType's wire schema.
Fixes#4461
Extended the whitespace normalizer to strip trailing whitespace from string values on well-known path/URL property names (path, paths, file, file_path, url, uri) so read-like tools no longer see stray trailing newlines emitted by the model. Content-carrying fields (content, input, etc.) are excluded to preserve intentional trailing newlines.
Fixes#4461
Normalized schema-matching enum and const strings before tool validation so provider-emitted trailing newlines do not reject valid tool calls.
Fixes#4461
- Added `normalizeSingleStringField` to dynamically map misplaced string inputs to required schema fields for single-argument tools.
- Integrated argument normalization into `validateToolArguments` to handle model-specific variations in JSON payloads during validation passes.
- Updated `coding-agent` streaming and rendering components to recognize `_input` as a legacy alias for `input` across various UI paths and logic flows.
- Refactored `hashlineEditParamsSchema` to strictly enforce the `input` field while maintaining support for legacy aliases via runtime coercion rather than schema definition.
- Corrected unit tests to reflect that `_input` is rejected by the strict schema but handled gracefully by the validation layer.
- Centralized transient transport error patterns to enable consistent reuse across error handling modules.
- Refactored `isOpaqueStatusBody` for improved accessibility in classification logic.
- Updated retryable error detection to include the consolidated transport pattern and additional provider-specific error criteria.
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
- Updated array normalization to unwrap double-encoded object keys within JSON-stringified arrays.
- Added a regression test to verify successful coercion of double-encoded objects inside string-encoded union types.
- Implemented recursive key normalization to unwrap object keys incorrectly serialized by LLMs (e.g., `{"\"op\"": "done"}`).
- Integrated normalization pass into `validateToolArguments` execution flow to prevent keys from being dropped by schema repairs.
- Added support for multi-layered encoding peels and protection against sibling key collisions during renaming.
- Handled edge cases where double-encoded keys are exposed only after internal JSON-string payload parsing.
- Added comprehensive testing for various scenarios, including nested objects and string-wrapped JSON payloads.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
Treat empty strings on optional tool arguments as omitted before schema validation so MCP calls do not fail pattern or type checks for model-filled placeholders.
Fixes#2981
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
- Enabled recursive parsing in `tryParseJsonForTypes` to coerce double-encoded JSON strings.
- Updated `coerceArgsFromIssues` to parse object/array JSON strings before singleton-array fallback.
- Prevented malformed container strings from being wrapped into arrays so validation returns array errors directly.
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
Only mark Zod and JSON Schema validation issues as union-branch when their own path matches the combinator's path so nested array fields inside a tag-selected branch keep their singleton-wrap repair.
Fixes#2026
Marked JSON Schema validation issues that come from a failed anyOf/oneOf branch and threaded the marker through the coercion bridge so the singleton-array wrap mirrors the Zod branch guard.
Fixes#2026
Skipped singleton array wrapping for expectations that come only from failed Zod union branches so other branch coercions can succeed first.
Fixes#2026
Wrapped non-string singleton values when an array field is expected so malformed Anthropic-compatible tool calls can validate instead of looping on errors.
Fixes#2026
Added regression coverage for a nested JSON-stringified object containing a
string-encoded array union field. The existing retry-loop normalization already
handles the case by rerunning after each container coercion; the test guards
that behavior and the loop comment now documents nested containers explicitly.
Fixes#1788
When the root `arguments` itself arrives as a JSON string (the entire
object stringified), the pre-validation `normalizeStringEncodedArrayUnions`
pass sees only the root string and no-ops — the schema accepts object,
not string-or-array, at the root. The validator then fails, the coercion
loop in `validateToolArguments` parses the root into an object via
`coerceArgsFromIssues`, but the union normalization never reruns, so any
inner `paths: '["..."]'` field stays a literal string and the #1788
search bug returns for root-string callers.
Re-ran `normalizeStringEncodedArrayUnions` inside the retry loop after
each `coerceArgsFromIssues` pass, mirroring how
`normalizeOptionalNullsForSchema` is already re-run there.
Added a regression test feeding both layers (root JSON-string +
inner JSON-string array) and asserting both unwind.
Fixes#1788
Only parsed JSON-array-shaped strings into arrays when the parsed value also
satisfies the schema's array branch. This preserves legitimate string inputs
such as `[1]` for `string | string[]` fields instead of turning them into
invalid numeric arrays.
Added a regression test for the reviewer-reported case.
Fixes#1788
Some providers (Z.AI / GLM) double-serialize array tool-call arguments
into JSON strings (`paths: '["a","b"]'` instead of `paths: ["a","b"]`).
The `search` and `gh` tools declare `paths` / `pr` as
`union(string, array<string>)` so zod accepted the JSON-encoded form
against the string branch — no type error fired, the existing JSON→array
coercion in coerceArgsFromIssues never ran, and downstream tools treated
the literal `["a","b"]` as a single path. With glob characters `[` / `]`
in the path, ripgrep parsed it as a character class: single-element
arrays silently matched nothing (e.g. `p`, `a`, `c`, … as filenames);
multi-element arrays parse-errored on the comma ("unclosed character
class"); absolute paths range-errored on out-of-order chars.
Added a pre-validation normalization pass that walks the JSON schema and,
when a node accepts BOTH `string` AND `array` (via `anyOf`/`oneOf` or
multi-type), JSON-parses array-shaped strings before zod sees them. Safe
because:
- Only fires when the schema explicitly accepts both shapes — single-
type schemas already get the existing coercion via type errors.
- Only substitutes when JSON.parse yields an Array — strings that fail
to parse or parse to non-arrays fall through unchanged.
- Strings without a leading `[` are skipped, so glob char-classes like
`paths: "[abc]"` are left alone for the tool.
Fixes#1788
- Expanded path parsing to split top-level comma, semicolon, and whitespace entries.
- Updated find/search and scope resolution to apply delimiter expansion before path-spec validation.
- Updated read tool fallback to try split path parts before raising missing-path errors.
- Coerced bare string arguments into singleton arrays for array-typed schemas.
- Migrated per-object caches (chat/tool starts, model fingerprints, validation contexts, provider indexes, render IDs) from WeakMap to Symbol-keyed properties on the objects themselves.
- Rewrote SSE debug tee as a single-pass inline parser, eliminating the body.tee() + readSseEvents re-parse pipeline.
- Refactored MockModel from a factory function + external WeakMap state into a self-contained class.
- Added FIFO memoization caches for heuristic candidate expansion and namespace suffix lookups.
- Added path-based key deletion helpers that remove keys from object and array arguments while preserving sibling structure sharing.
- Mapped Zod unrecognized_keys and JSON Schema additionalProperties violations to an unrecognized issue type and applied them in issue repair to tolerate strict input shapes.
- Updated tool argument coercion tests to confirm extra keys are stripped for strict and nested strict schemas and additionalProperties false payloads.
- JSON Schema validator now enforces propertyNames, patternProperties, dependentRequired, dependencies, if/then/else, contains, and prefixItems instead of silently accepting values that violate them; unevaluated* still permissive but warns once.
- Recursive $ref no longer short-circuits to true on revisit: cycle detection keys on (ref, value-identity) with a primitive depth cap, so nested sub-schema violations are caught.
- Meta-validator now structurally validates if/then/else/dependencies sub-schemas and accepts draft-07 dependencies as either schemas or string-array dependent keys.
- Wire-schema null normalization no longer strips null-valued unknown root fields before preserveUnknownRootFields snapshots them, so task.simple and similar callers still see disallowed null arguments.
- Replaced fromTypeBox conversion with a JSON-schema validator flow in ai tool handling and execution paths.
- Added recursive schema validation and expanded TypeBox checks for refs, enums, uniqueItems, and constraint keywords.
- Sanitized Azure/CCA tool schemas by dropping unsupported fields and rewriting oneOf tool branches as anyOf.
- Tightened argument and model-config validation, preserving unknown tool fields and adding apiKey plus compatibility flags.
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
- Preserved description at the top level when wrapping optional properties with anyOf/null.
- Substituted schema-supplied defaults when a required field arrives as null or "null", cloning to prevent cross-call mutation.
- Added coercion tests for default substitution, isolation, optional null stripping, and nested JSON deserialization.
- In `normalizeOptionalNullsForSchema`, unknown object keys were removed when their value was `null` or `"null"` and `additionalProperties` was false.
- Non-null unknown fields were left intact so malformed extra properties still triggered validation errors.
- Escaped raw control characters inside JSON string literals in `tryParseJsonForTypes` and `tryParseLeadingJsonContainer`.
- Fixed `validateToolArguments` parsing of stringified tool args containing literal newlines or tabs in string values.
- Added regression coverage for raw-control-character JSON arguments in `tool-argument-coercion.test.ts`.
- Extracted `structuredCloneJSON()` utility to centralized package for consistent deep cloning with JSON fallback.
- Consolidated duplicate cloning logic across openai-responses, openai-responses-shared, and validation modules.
- Added resilience to message cloning in extension runner with try-catch fallback for non-cloneable objects.
- Optimized context emission in extension runner by skipping message cloning when no handlers exist.
- Added optional `contentIndex` field to AssistantMessageEvent variants for type consistency.
- Added automatic cleaning of literal escape sequences (`\n`, `\t`, `\r`) in JSON parsing to handle LLM encoding confusion.
- Added support for healing JSON with trailing junk after balanced containers (e.g., `]\n</invoke>`).
- Implemented `cleanLiteralEscapes()` function to distinguish between literal backslash sequences outside JSON strings and actual escape sequences within strings.
- Added 2 test cases validating escape sequence cleaning and trailing junk recovery in tool argument coercion.
- Added automatic healing of malformed JSON with single-character bracket errors at the end of strings, improving LLM tool argument parsing robustness.
- Implemented tryHealMalformedJson() function to correct common LLM mistakes like misplaced or extra brackets via single-character edits.
- Added 3 test cases covering bracket healing scenarios: extra bracket, wrong bracket type, and deeply broken JSON rejection.
- Added `edit.blockAutoGenerated` setting to control enforcement of auto-generated file detection.
- Improved auto-generated file detection to use language-specific comment parsing instead of broad regex patterns, reducing false positives.
- Enhanced marker detection to scan only leading header comments (1024-byte limit) rather than entire file prefix for better accuracy.
- Fixed tool argument validation to properly handle string 'null' values on optional LLM tool arguments.
- Improved type safety by changing validateToolCall and validateToolArguments return types from any to ToolCall["arguments"].
- Add dereferenceJsonSchema() that inlines local $ref pointers and strips
$defs/definitions from MCP tool schemas before they reach LLM providers.
Previously, Anthropic's convertTools() extracted only properties/required,
dropping $defs and leaving dangling $ref — the LLM never saw the actual
type definitions (e.g. SourceAnchorInput enum values from nucleus).
- Silence Ajv logger (logger: false) on all three instances that use
strict: false. MCP servers may declare non-standard format keywords
(e.g. "uint") that caused console.warn() to corrupt TUI output.
- Cache compiled Ajv validators per schema object identity in validation.ts,
eliminating redundant recompilation on every tool call.
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Fixed tool argument coercion to handle malformed JSON with trailing wrapper braces by parsing leading JSON containers.
- Relaxed JSON detection heuristics to check only leading braces instead of requiring matching closing braces.
- Added test case for array strings with trailing wrapper braces from nested JSON malformation.
* fix(ai): coerce string-to-number for Optional<number> parameters
When a tool schema uses Type.Optional(Type.Number()), TypeBox generates
an anyOf:[{type:"number"},{type:"null"}] schema. AJV reports validation
failures against this as keyword:"anyOf" errors rather than
keyword:"type" errors.
coerceArgsFromErrors only processes keyword:"type" errors, so Optional
numeric fields were invisible to the coercion loop. normalizeOptionalNullsForSchema
traverses anyOf branches recursively but returned early at primitive-type
branches (type !== "object") without attempting string coercion, so the
anyOf traversal produced no change and the string was left uncoerced.
Fix: add a string-to-number coercion guard in normalizeOptionalNullsForSchema
before the early return, covering schema branches that declare
type:"number" or type:"integer". This allows the existing anyOf
traversal to coerce "1.0" -> 1.0 when the schema is Optional<number>.
Required numeric fields (Type.Number()) were unaffected as they produce
keyword:"type" errors handled by the existing path.
Reproducer: any MCP tool with an optional float parameter (e.g. tick_size:
Option<f64> in Rust/schemars) receives the value as a string from the
LLM and fails deserialization despite the coercion infrastructure being
designed to handle exactly this case.
* style: biome format validation.ts and coercion test
---------
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>