- Deleted the `ValidationVerdict` type and `evaluateOutputAgainstSchema` API from the output schema validator.
- Updated `yield.ts` to bind `buildOutputValidator`'s error directly to `schemaError` during validator setup.
- Removed the obsolete evaluator tests and adjusted validation success fixture to match the raw summary input shape.
- Unified output schema construction and validation by adding buildOutputValidator and using it in YieldTool and task executor.
- Added MAX_SCHEMA_RETRIES so YieldTool now retries schema failures three times with hints before overriding.
- Updated failure handling to use shared summarizeValidationFailure and formatters for required-field reporting.
- Added tests for output-schema-validator and YieldTool covering malformed schemas and nested-array retry edge cases.
Centralizing OAuth refresh in AuthStorage (e6893515) introduced five
follow-on bugs surfaced by an audit of the commit; this fixes all of
them and updates the tests that relied on the old refresh seam.
1. packages/ai/src/auth-storage.ts (#tryOAuthCredential):
For built-in providers the path went directly to `getOAuthApiKey`
with the (possibly still-expired) selection.credential when the
pre-refresh at line 2587 caught a transient error. `getOAuthApiKey`
then threw the "expired … must be refreshed via AuthStorage"
precondition error, which the disable classifier matched against
`/expired.*refresh/` and soft-disabled the row. A single network
blip during refresh could permanently kill a still-valid Anthropic /
OpenAI / Gemini-CLI / Copilot credential. Built-in providers now
route through the broker-aware single-flighted
`#refreshOAuthCredential` first, so transient failures surface as
network errors (5-min temp block) instead of definitive auth
failures.
2. packages/ai/src/auth-storage.ts (#fetchUsageUncached):
The usage refresh check only fired once `Date.now() >= expiresAt`,
missing the 60-second skew that `getApiKey` honors. A token
expiring inside the skew window was posted to the usage endpoint
and 401'd mid-flight, briefly hiding quota in the UI. Aligned with
`OAUTH_REFRESH_SKEW_MS`.
3. packages/coding-agent/src/web/search/index.ts (webSearchCustomTool):
The CustomTool counterpart of WebSearchTool dropped sessionId so
SDK callers that opted into `web_search` via toolNames lost
per-session credential stickiness — multi-account users saw the
provider round-robin between searches in the same session. Threads
`ctx.sessionManager.getSessionId()` through to `executeSearch`.
4. packages/coding-agent/src/web/search/providers/perplexity.ts
(findOAuthToken):
`authStorage.getApiKey("perplexity")` returns runtime/config
overrides, stored api_key credentials, OAuth bearers, and env keys.
Filtering only env keys meant a config-pinned `pplx-…` API key was
POSTed to `www.perplexity.ai/rest/sse/perplexity_ask` (the OAuth
endpoint) instead of falling through to
`api.perplexity.ai/chat/completions`, producing 401s. Switched to
`getOAuthAccess` so only true OAuth bearers reach the OAuth
branch; api_key credentials/overrides correctly fall through.
5. packages/ai/scripts/generate-models.ts:
`getOAuthApiKey` was being called directly with possibly-expired
credentials. The new contract throws on expired, the broad catch
swallowed it, and the build silently fell back to bundled models
instead of refreshing. Both helpers now route through
AuthStorage's `getApiKey` / `getOAuthAccess`, which trigger the
full broker-aware refresh pipeline.
Test updates:
- auth-storage-credential-disabled-event.test.ts,
sdk-credential-disabled-bridge.test.ts: the `failOAuthRefresh`
helper used to spy on `getOAuthApiKey` to inject invalid_grant.
With refresh now happening before that helper, the spy never fired.
Switched to spying on `refreshOAuthToken` so the simulated failure
reaches the disable classifier.
- auth-storage-rotation.test.ts: stub `refreshOAuthToken` so the test
doesn't hit a real OAuth endpoint when the seeded credential lands
inside the 60s skew window.
Combined task.enableLsp with the parent session enableLsp gate before spawning subagents, so --no-lsp remains authoritative even when subagent LSP is enabled in settings.
Fixes#1385
Added task.enableLsp (boolean, default false) and routed both regular and isolated subagent dispatch through it. Keeps subagents cheap by default while letting users opt in to LSP-aware delegation. Updated regression tests to cover the default-off, opt-in, plan-mode, and isolated paths.
Fixes#1385
Subagents now inherit the parent session's enableLsp value, so a top-level --no-lsp invocation propagates into spawned tasks instead of falling back to the executor's default of true.
Fixes#1385
Removed the hardcoded subagent LSP disable flag so executor defaults and user settings control LSP availability. Passed the effective plan-mode agent definition into both regular and isolated subagent dispatch so plan-mode tool restrictions apply consistently.
Fixes#1385
- packages/ai/test/issue-957-repro.test.ts now tests:
- refreshKimiToken applies the 5-minute server-side skew (Kimi-specific)
- AuthStorage refreshes kimi-code credentials inside its 60s skew window
- packages/ai/test/anthropic-stream-timeout.test.ts: raise the
streamFirstEventTimeoutMs from 10ms to 5000ms so slow CI scheduling
cannot fire the first-event watchdog before the mocked events arrive.
The test still exercises the (1ms) idle path it was written for.
fix(web): allow Parallel extract via PARALLEL_API_KEY env var without storage
The fetch tool and YouTube scraper previously gated the Parallel extract
branch behind `storage && findParallelApiKey(storage)`. With no
AgentStorage the env key was never consulted, so callers that ran
without a per-session storage (e.g. ReadTool sessions in unit tests, and
in practice any caller that has only an env API key) silently fell back
to raw-html / no-ytdlp paths.
- findCredential/findParallelApiKey now accept null or undefined storage
and rely solely on the env-first path when no storage is supplied.
- searchWithParallel/extractWithParallel mirror the same nullable shape.
- Drop the redundant `storage && ` guards in fetch.ts and youtube.ts;
the inner findParallelApiKey call already returns null when no
credential is available.
- Centralized OAuth access lifecycle in `AuthStorage`, returning identity metadata and new access-result types.
- Added 60-second skew and strict expiry checks, returning undefined/throws for stale or expired OAuth credentials.
- Removed provider-local token refresh flows from Gemini, Gemini CLI, Antigravity, Kimi, and related OAuth helpers.
- Migrated web-search providers from `AgentStorage` to `AuthStorage` session-aware lookup with `authStorage`/`sessionId`/`signal` flow.
- Replaced `findAnthropicAuth`/DB auth lookup with `buildAnthropicAuthConfig` and explicit base-url override/env fallback ordering.
`ref` is a JTD-reserved keyword (RFC 8927) used by the schema-reference
form, so the JTD-to-JSON-Schema converter on releases prior to 15.3.2
silently dropped it from the generated JSON Schema and required it at
the same time. Every explore-agent invocation then failed validation
with `schema_violation: files.0.ref: must not be present`.
The converter side was hardened in #1345 (shipped in 15.3.2). This
rename is defense-in-depth at the prompt level: the explore agent's
output contract no longer relies on the converter recognising a
user-named property that collides with a JTD keyword, and the field
name now matches what it actually carries.
Fixes#1379
- Added OpenAI Codex and Gemini web search provider options with updated setup/auth descriptions.
- Updated Codex OAuth flow to refresh near-expiry tokens during web_search and persist the refreshed credentials.
- Plumbed AgentStorage through search orchestrator, scrapers, and fetch paths so providers share session credentials.
- Refactored web provider and credential helpers to accept caller-provided AgentStorage and resolve keys synchronously.
Quarantined persistent session keys only while the native cancellation promise remains unsettled, so healthy cleanup restores persistent mode and stalled cleanup cannot accumulate live shell instances.
Added coverage for both stalled and settled native cleanup paths.
Fixes#1347
Queued extension-delivered user messages when deliverAs is set and waited for session_start extension message sends before prompting subagents.
Fixes#1343
Stopped marking persistent bash sessions as permanently broken when the JavaScript abort or timeout race wins.
Stopped the Rust descendant kill-wave helper once no cancellation targets remain so later commands are not swept into old cancels.
Fixes#1347
- Updated nested task-rendering tests to use parent-qualified IDs for completed child task results.
- Updated in-flight nested snapshot expectations to verify parent-aware `Parent>Subtask` labeling.
- Documented the live nested task rendering behavior in the package changelog.
- Captured `tool_execution_update` snapshots for `task` calls into in-flight progress state for live nested rendering.
- Cleared in-flight task snapshots at task start and completion to prevent stale nested progress from persisting.
- Updated progress rendering to combine completed and in-flight task details through a dedicated nested task tree view.
Raced bash execution against the JavaScript abort signal and timeout so the tool returns even when native shell cleanup does not settle.
Added regression coverage for native cleanup stalls on ESC abort and timeout.
Fixes#1347
The report_finding tool's priority is exposed as a string enum
("P0"-"P3") for ergonomics, but the reviewer agent and every
custom review agent declare priority as `type: number` in their
JTD output schema. The cast at executor.ts:1473 lied about the
runtime shape, so the auto-injected `findings[].priority` flowed
through as strings and every yield with at least one finding was
rejected with `findings.0.priority: expected number, received string`,
forcing the run into the schema_violation exit path.
Added `toReviewFinding(details)` in tools/review.ts that maps the
priority enum to its numeric ordinal via the existing PRIORITY_INFO
table and use it at the boundary in executor.ts. Render paths still
see the original `ReportFindingDetails` shape (string priority)
through normalizeReportFindings, so display formatting is unaffected.
Fixes#1350
- Extended `invalidateCredentialMatching` to accept session-scoped options and clear cached session credentials before blocking the matched credential.
- Updated the OAuth auth-error retry flow to pass `agent.sessionId` through credential invalidation.
- Added a regression test ensuring invalidating a session-sticky OAuth key rotates to the next active credential.
- Added `checkCredentials()` with result types/options for per-credential tri-state health checks.
- Added `/v1/credentials/check` endpoint via `handleCredentialsCheck` returning `{ generatedAt, credentials }`.
- Added `omp auth-gateway check` flow with provider grouping, `--json` output, and exit status 1 on failures.
- Added command examples, changelog updates, and tests for expired OAuth refresh, null/missing config, and ordering edge cases.
Loaded marketplace lspServers metadata from Claude plugin caches and embedded it for OMP marketplace installs so config-only plugins register without package code.
Fixes#1352
The JTD-to-JSON-Schema converter post-processed convertSchema's
output with normalizeMixedSchemaNode, which walked back into the
emitted JSON Schema looking for nested JTD forms. Inside a
properties block, user-defined property names whose keys happened
to collide with JTD keywords ('ref', 'elements', 'values',
'optionalProperties', 'discriminator') were misclassified as JTD
forms and re-rewritten - corrupting properties like { ref: { type:
'string' } } into { $ref: '#/$defs/[object Object]' } and breaking
the built-in explore agent's output validator with
schema_violation: files.0.ref: must not be present.
convertSchema is already fully recursive and emits pure JSON Schema,
so the post-walk is both unnecessary and unsafe. Drop it.
Fixes#1345
- Exported `normalizeTools` so `AppendOnlyContext` uses the same tool normalization as the agent loop.
- Added `BuildOptions.intentTracing` to `build()`/`reset()`/`takeSnapshot()` so intent injection is consistent and included in the prefix fingerprint.
- Improved `#computeDigest` to cover tool_calls, tool_call_id, name, and id fields to catch in-place mutations.
- Fixed `#unsubscribeAppendOnly` leak and added no-op guard in `#syncAppendOnlyContext`.
- Extracted `mergeDiscoveredModel` so discovered baseUrl takes priority over bundled entry, fixing 401s on Xiaomi tp- token-plan streams.
- User providerOverride.baseUrl still wins over both discovered and bundled values.
- Added regression tests covering all merge priority paths.
- Added Symbol-keyed sidecar on each AgentMessage to memoize estimateTokens, with a cheap content fingerprint to detect in-place mutations.
- Fixed stale cache on same-length replaceMessages, post-hoc error attachment, and branch rebuild edge cases.
- Fixed usage fetch error backoff: stamped fetchedAt on failure so the 5-min TTL also gates retries during outages.
- Extracted computeNonMessageBreakdown as shared helper to prevent drift between status-line and context panel token counts.
- Fixed /plan and /goal history preservation by snapshotting enabled state before handlePlanModeCommand/handleGoalModeCommand executes.
- Previous check read state after the call, missing cases where the handler itself toggled the mode off (e.g., confirmed exit).
- Added tests covering confirm-exit, cancel-exit, and first-activation paths.
- Raised PowerShell timeout to 8s and swallowed reap errors to prevent unhandled throws on WSL interop.
- Fixed fallback logic so arboard is skipped when no display server is present on headless WSL.
- Added test coverage for the headless WSL short-circuit path.
- Added `recoverOrphanedBackups` to promote `.jsonl..bak` files back to their primary path when the primary is missing, preventing data loss after a mid-rename crash.
- Changed backup filename from dot-prefixed to plain `..bak` so the shared `*.bak` glob can find it on both real and in-memory storage backends.
- Surfaced the original EPERM as the error `cause` and included both original and retry messages when rollback also fails.
- Replaced baked module-load value with per-call `isWebPExcluded()` so runtime env changes take effect.
- Only `"1"` and `"true"` (case-insensitive) enable exclusion; empty string and `"0"` are treated as disabled.
- Fast path now bypassed for WebP sources when exclusion is active.
- Explicit error surfaced when decode fails and WebP exclusion cannot be honored.
Anchors are formatted by read/search as LINE+HASH|TEXT, and lines may be
prefixed with marker decoration (*, >, +, -). The parser previously required
a bare LINE+HASH and rejected verbatim copy-pasted anchors with:
line N: expected a full anchor such as "119sr", ...; got "364sp|".
Loosen LID_CAPTURE_RE to allow optional leading decoration and an optional
trailing |... body on each anchor (including each side of a range).
- Added `retry.maxDelayMs` to the settings schema and interfaces, with a default cap for provider backoff delays.
- Updated session auto-retry logic to fail fast when a requested wait exceeds the cap without fallback, emitting terminal auto-retry failure state.
- Propagated retry state and failure data into task progress and rendering so children show retry/wait details and reminder prompts stop after terminal errors.
Separated model selector provider tab labels from provider ids so human-readable labels like Ollama Cloud refresh and filter the underlying ollama-cloud models.
Fixes#1153
In a compiled binary, Bun.resolveSync(spec, import.meta.dir) throws
'Cannot find module' because import.meta.dir is inside /$bunfs/root
and the virtual FS exposes no node_modules tree at runtime.
Previously this throw propagated through rewriteLegacyPiImports ->
rewriteLegacyPiImportsForRuntime -> mirrorLegacyPiFile ->
loadLegacyPiModule -> loadExtension, which swallowed it as 'Failed to
load extension' and silently dropped any plugin whose files imported
@mariozechner/pi-ai (or any @mariozechner/pi-* whose bundled
counterpart isn't reachable via resolveSync in the binary).
Fix: wrap the resolution call in rewriteLegacyPiImports in a try/catch
and return the original match on failure. rewriteBareImportsForLegacyExtension
runs immediately afterwards in every call path and already resolves bare
specifiers against the importer's real filesystem directory, so it picks
up @mariozechner/pi-ai from the plugin's installed peer deps instead.
Apply the same fallback to resolveLegacyPiSpecifier (the Bun plugin
shim's onResolve handler) for tool/hook files loaded directly via Bun's
import system rather than through loadLegacyPiModule.
Fixes#1215