Export removeSyncWithRetries from @oh-my-pi/pi-utils as a standalone
function, then migrate the highest-impact fs.rmSync call sites:
- test/helpers/temp-home-cleanup.ts: 2 fs.rmSync → removeSyncWithRetries
(affects all tests using the cleanupTempHome helper)
- test/core/apply-patch-regression.test.ts: 16 fs.rmSync → removeSyncWithRetries
(the most fs.rmSync calls of any test file)
removeSyncWithRetries retries on EBUSY/EPERM/ENOTEMPTY (40x25ms on
Windows), matching TempDir.removeSync's retry logic. On Linux/macOS
(CI) the retries are a no-op — fs.rmSync succeeds immediately.
Follow-up to #3071 review (codex + @PGupta-Git): the family rewrite alone
did not heal users running off the bundled catalog or stale SQLite cache
rows, which still routed Sonnet 4.6 thinking efforts to the 404
`claude-sonnet-4-6-thinking` wire id. `collapseEffortVariants` treats
those collapsed snapshots as authoritative and `refreshCollapsedThinking`
exits early for families without `effortBudgets` (Claude pairs), so the
new empty routing in the hand table never reached the snapshot.
- Declared the dead wire ids as `retiredMembers` on the Claude 4.6
families (`claude-sonnet-4-6-thinking` on the Sonnet family,
`claude-opus-4-6` on the Opus family). This triggers
`reconcileRetiredRouting` to rewrite every `effortRouting` entry that
targets a retired id to a live wire id (Sonnet falls back to the bare
member; Opus falls back to `-thinking`).
- Refreshed the bundled `packages/catalog/src/models.json` Sonnet 4.6
entry so fresh installs do not boot with the dangling routing — the
surgical diff matches what the generator would emit; the rest of the
catalog is left untouched to keep the bug-fix PR scoped.
- Added regression tests in `variant-collapse.test.ts` for both
reconciliation paths and a bundled-catalog test
(`issue-3067-repro.test.ts`) that pins the live-wire-id resolution end
to end through `buildModel` for every effort tier.
Fixes#3067
- Privatized the legacy `nextToolChoice` method to `#nextHardToolChoice` to ensure all tool-choice directives flow through the unified `nextToolChoiceDirective` entry point.
- Eliminated redundant dual entry points for fetching tool choices, which previously bypassed the soft pending-preview lifecycle.
- Updated test suites to consume `nextToolChoiceDirective` where appropriate to maintain consistency with internal agent-loop logic.
normalizeTools only ran toolWireSchema/zodToWireSchema/arkToWireSchema
inside the if (doInjectIntent) branch, so any tool whose intentMode
resolved to "omit" (function intent, or intent: "omit" like the
builtin eval/resolve) had its live arktype/zod schema object left in
parameters. The pruneDescriptions branch already wire-encoded
unconditionally, so the two branches disagreed; first-party providers
re-encoded with toolWireSchema themselves and never noticed, but any
downstream consumer that serialized parameters as-is ended up with an
invalid schema (arktype AST: object-valued required, domain instead of
type, no properties).
Hoist toolWireSchema(t) out of the inject-intent branch so it always
runs and inject intent on top, mirroring the pruneDescriptions shape.
Drop the now-redundant isZodSchema/isArkSchema/zodToWireSchema/
arkToWireSchema imports.
Fixes#3074
The daily Cloud Code Assist backend (`daily-cloudcode-pa`) exposes Claude 4.6
asymmetrically: `claude-sonnet-4-6` has no `-thinking` twin and
`claude-opus-4-6` has only the `-thinking` twin. The shared
`thinkingPair("claude-sonnet-4-6", …)` family (with `preserveAbsentEffortRoutes`)
kept every effort routed to `claude-sonnet-4-6-thinking` even when discovery
only returned the bare id, so any reasoning-on request 404'd with
`Requested entity was not found`. Claude on Antigravity also caps
`maxOutputTokens` at 64000, while `ANTIGRAVITY_MODEL_WIRE_PROFILES` had no
Claude entries — discovery's 65536 propagated to the wire and 400'd with
`Request contains an invalid argument`.
- Replaced the two `thinkingPair` calls for Claude 4.6 in `SHARED_CCA_FAMILIES`
with bespoke single-wire families. Sonnet collapses to the bare wire id,
Opus collapses to the `-thinking` wire id, and per-effort thinking is
carried by the request body's `thinkingBudget` on the single shared wire id.
Listing both candidate ids in `members` (priority order) keeps the collapse
correct if the backend mix ever rebalances.
- Added `claude-sonnet-4-6` and `claude-opus-4-6-thinking` entries to
`ANTIGRAVITY_MODEL_WIRE_PROFILES` capping `maxOutputTokens` at 64000.
- Made `AntigravityModelWireProfile.modelEnum` optional — Anthropic-backed
wire ids are accepted without a captured `labels.model_enum` token. The
request builder now emits the label only when the profile defines one.
- Regression tests in `variant-collapse.test.ts` (routing/wire-id resolution
for both 4.6 families across all three discovery permutations) and
`google-gemini-cli-alignment.test.ts` (request builder caps Claude
`maxOutputTokens` at 64000 and omits the unset `model_enum` label).
Fixes#3067
The #3063 fix introduced a second mutating step — `bun update <name>` —
that rewrites bun.lock before extension validation runs. Three failure
paths could still leave the rejected commit pinned in the lockfile or
active tree:
- Extension validation throwing after `bun update` had refreshed
bun.lock — rollback restored package.json and node_modules/<name>
but never touched bun.lock.
- Feature validation (`omp plugin install pkg[ghost]`) throwing
outside the rollback block entirely.
- Runtime-config save failing after a successful install with no
rollback path.
Snapshot bun.lock alongside package.json before `bun install` runs and
route every post-install step (resolution, update, package.json read,
feature validation, extension validation, runtime-config save) through
one outer catch that restores all three (package.json + bun.lock +
node_modules/<name> from snapshot). `#rollbackFailedInstall` now
tolerates an unresolved `actualName` for failures that throw before
the dep key is known.
Three regression tests in plugin-install-validation.test.ts pin the
new contract: bun.lock restoration after a git reinstall fails
validation, bun.lock removal when it didn't exist pre-install, and
rollback on an unknown feature request.
Addresses review feedback on #3069.
bun install <spec> respects the existing bun.lock pin when the spec is
unchanged and never re-resolves the remote ref, so re-running
`omp plugin install github:owner/repo` on an already-installed plugin
reported success while silently keeping the user on the original
resolved commit (1ms no-op, no network).
PluginManager.install now follows a git re-install with
`bun update <name>` to force re-resolution of the ref against the
upstream. First-time installs (no prior dep entry) skip the update —
the initial bun install already fetches HEAD. bun update failures
trigger the same rollback path as validation failures.
Fixes#3063
Forwarded peekPendingInvoker and clearPendingInvokers from the production toolSession literal so the resolve tool can dispatch staged previews and drain stale gates in real CLI sessions, not only in unit tests.
Added a regression test exercising the production wiring through AgentSession and a phantom-gate drain via a no-invoker facade.
Fixes#3061
Cleared pending-preview markers when resolve has no runnable handler so a stale gate cannot keep forcing resolve after the invoker is gone.
Added regression coverage for apply and discard draining stale pending markers.
Fixes#3061
- Introduced `withEmptyCompletionRetry` utility to manage bounded retries with exponential backoff for empty streaming responses.
- Integrated completion validation across Anthropic and OpenAI providers to ensure robust handling of intermittent empty outputs.
- Updated `omp bench` and `omp dry-balance` to correctly resolve extension-contributed providers and report content-less runs as failures.
- Added comprehensive unit and regression tests to validate retry logic, provider resolution, and failure reporting metrics.
- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
Kept observing a status-line usage fetch after the startup timeout so late successful reports still refresh the quota segment instead of being hidden behind the timeout backoff.\n\nFixes #3057
Deferred the status-line quota refresh off the render path and raced it against a short startup timeout so slow Anthropic usage lookups cannot pin interactive startup. Added regression coverage for non-synchronous refresh startup and timeout backoff.\n\nFixes #3057
- Tighten invalidation logic to only trigger when a demonstrably warm cache (previously read) goes cold.
- Ignore cold transitions following write-only turns to prevent spurious markers during initial cache warming or after natural TTL expiry.
- Update test suite to verify that consecutive cold turns following an initial write do not trigger invalidation alerts.
- Improved thinking loop detection logic by refining text normalization and treating stalls as retryable errors.
- Instrumented the agent session to recognize thinking loop markers within retryable error conditions.
- Automated the clearing of stale error banners upon successful auto-retry execution.
- Added comprehensive test coverage for chunked thinking loop errors and banner management.
- Moved authentication logic to `perplexity-auth.ts` to share logic between search providers and CLI commands.
- Updated authentication priority to prefer browser cookies over OAuth tokens during search operations.
- Modified the `token` CLI command to display active OAuth tokens when both an OAuth token and an API key are configured.
- Added comprehensive unit tests in `perplexity.test.ts` to verify authentication priority and precedence.
- Replaced usage of `ReturnType<typeof setTimeout>` and `ReturnType<typeof setInterval>` with the explicit `Timer` type across the codebase.
- Updated several type definitions and function signatures to use concrete types instead of inferred return types for improved clarity and maintainability.
- Added `ensure_workspace_dependencies` to automate `bun install` for new worktrees where dependencies are missing.
- Configured the installer to use `--frozen-lockfile` and `--ignore-scripts` to preserve security and lockfile integrity.
- Integrated the bootstrap step into the worker startup process to ensure workspace packages are resolvable.
- Downloaded config.json with tokenizer sidecars when repairing stale fastembed model caches.\n- Treated fastembed Config file missing errors as repairable initializer failures.\n- Extended cache repair coverage for the missing-config state.\n\nFixes #3054
The PR fixed the -32601 hang for the defined server->client refresh
requests but missed two real spec methods of the identical class:
workspace/inlineValue/refresh (LSP 3.17) and workspace/foldingRange/refresh.
A server emitting either still received Method not found and could stall.
Add both to the void-ack chain and extend the regression test.