- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
The previous commits patched each rebuild path individually to avoid feeding a
modifyModels hook its own output. That left the invariant implicit and the
provider-scoped path applying only a subset of hooks, which is wrong for a hook
that inspects or suppresses another provider's models.
Keep #unprojectedModels as the canonical pre-projection catalog and derive
#models from it at every mutation point, so projections are always a pure
function of the unprojected base:
- #composeUnprojectedStaticModels builds the catalog; #composeStaticModels
projects it. A scoped lookup with modifiers registered composes and projects
the whole catalog before narrowing, matching getAll() followed by a filter.
Providers without modifiers keep the cheap filtered path.
- Discovery completion, registerProvider, and runtime transport overrides
update the unprojected snapshot and reproject, instead of mutating an
already-projected array.
- Runtime metadata patches apply to the unprojected model, then reproject, so
a later registration cannot discard them.
- Provider lookup snapshots are invalidated wherever the projection changes.
Hooks no longer take a providerFilter: a modifier is a whole-catalog transform
and every rebuild now runs the full ordered set exactly once.
(cherry picked from commit e5d2e9eac7c371cc196e9b362f77d3a5d7bdf507)
registerProvider composed nextModels from the already-projected #models,
stripping only the incoming provider, then reran every stored modifier over
it. Loaders drain registrations one at a time, so the previously registered
provider's projection was fed back into its own hook — an append-style hook
compounded on each subsequent registration.
Apply only the incoming provider's hook. Every other provider's projection is
already present exactly once, and full rebuilds still go through
#composeStaticModels.
(cherry picked from commit b16db642b08cc223b062c0c556f00d46fda2520a)
Review follow-up on two defects in the original change:
- The throwing-hook fallback wrote to #lastDiscoveryWarnings, which is only
ever read to dedup a logger.warn inside #warnProviderDiscoveryFailure. No
log line was emitted, so a broken extension degraded invisibly, and the
shared key could mask a later discovery failure for the same provider. Log
via logger.warn with its own dedup map.
- #refreshRuntimeDiscoveries starts from the already-projected #models, and
the overlay merge only replaces matching provider+id pairs, so a hook's
projection-only entries survived and were fed back into it. An append-style
hook duplicated its output on every refresh. Drop each modifier provider
before the merge so it re-seeds from the unprojected overlays.
(cherry picked from commit b6f841e080d4882a08b8d713de009461b6acc6fe)
`registerProvider` applies `oauth.modifyModels` once and assigns the result
straight to `#models`, but only the pre-projection definitions are persisted
in `#runtimeModelOverlays`. Any subsequent static reload rebuilds `#models`
from those overlays and silently drops the projection.
The model selector reloads on every open (`refresh("offline")`), so an
extension provider that projects a credential-aware catalog shows its
correct models everywhere except the picker — the one place users look.
`refreshProvider()` and online discovery completion had the same hole.
Persist the hook per provider and re-apply it wherever `#models` is
recomposed, honouring the `providerFilter` used by scoped lookups. A hook
that throws now degrades to that provider's unprojected catalog instead of
failing the whole composition, so one broken extension cannot empty the
registry.
(cherry picked from commit 33b7c72f225b4253b68bb71ecb3a9186151b18b3)
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).
Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.
Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.
Fixes#4893
Co-authored-by: roboomp <omp@can.ac>
Preserved extension-registered provider and model header objects so request-time reads observe later mutations instead of registration-time snapshots.
Fixes#3725
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
The runtime override-only branch in ModelRegistry.registerProvider gated on
baseUrl/headers/apiKey/authHeader/transport and built a transportOverride that
omitted remoteCompaction, so calls like
registerProvider('openai', { remoteCompaction: { enabled: false } }) silently
left every existing model's compaction config unchanged.
Added remoteCompaction to the gate and to the override payload — the existing
mergeProviderOverride / applyProviderTransportOverride helpers already merge it
onto runtime override state and models, so wiring is enough. Added a regression
test that drives the override-only path and asserts the propagation survives
refresh and refreshProvider cycles, then is undone by clearSourceRegistrations.
Fixes#3104
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
- Added SGR mouse routing to wizard scenes for hover, click, and wheel interactions.
- Added pointer-based tab selection and panel wheel scrolling in providers tabs.
- Added splash/outro left-click handling to start and complete the wizard flow.
- Added setup-wizard tests for mouse routing, splash click entry, and local coordinates.
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
- Derived descriptors, default-model map, env keys, login list, and refresh dispatch from one ProviderDefinition per provider.
- Disabled OpenAI Codex stream obfuscation and interrupted whitespace-only tool-call argument deltas.
- Derived auth-broker callback ports and paste-code login set from the registry.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
- Documented and removed `utils/oauth` from the `ai` package entrypoint, noting it as a breaking change.
- Refactored `cli`, `auth-storage`, and `utils/oauth` to load provider modules via scoped dynamic `import()` calls.
- Removed top-level provider imports and barrel exports from `utils/oauth/index.ts`, streamlining oauth module loading.
- Consolidated OAuth symbol, type, and provider imports in coding-agent and tests to `@oh-my-pi/pi-ai/utils/oauth` modules.
- Defined `DEFAULT_LOCAL_TOKEN` locally in model-registry and removed its cross-package OAuth import usage.
- Renamed subagent completion flow from `submit_result` to `yield` across SDK tools, prompts, and docs.
- Updated executor/task handling to require and parse `yield` calls, replacing legacy submit-result extraction and state flags.
- Added `subagent-yield-reminder` and updated system prompts to require `yield` with `result.data` or `result.error`.
- Renamed hidden-tool and registration plumbing to `yield`, including discovery helpers and renderer/test surface.
Extension-registered models (via registerProvider) were stored directly in #models, which gets rebuilt from scratch on every refresh() call. The model selector triggers refresh("offline") on open, wiping all extension models from the list.
Store extension models as #runtimeModelOverlays that survive reloadStaticModels() and get merged during both #loadModels() and refreshRuntimeDiscoveries(). Persist extension API keys in a separate map restored after each clear cycle.
Remove unused buildCustomModel wrapper (sole call site refactored to use buildCustomModelOverlay + finalizeCustomModel directly).
- Added `close()` method to SessionManager and AuthStorage for proper resource cleanup and finalization of prepared statements.
- Added `initiatorOverride` option support in OpenAI and Anthropic providers for message attribution control.
- Fixed resource leaks in RpcClient timeout handling by centralizing timeout creation with unref() and adding explicit clearTimeout() calls.
- Fixed AgentSession disposal to call SessionManager's `close()` method for guaranteed resource cleanup instead of fallback flush.
- Updated all test suites to properly dispose AuthStorage instances in cleanup hooks to prevent resource leaks between tests.
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Consolidated truncation and output utilities from tools/truncate.ts and tools/output-utils.ts into session/streaming-output.ts with improved UTF-8 boundary handling.
- Renamed formatSize() to formatBytes() across codebase for consistency and clarity in byte-level formatting.
- Refactored OutputSink to use windowed byte truncation instead of full-buffer encoding, improving memory efficiency on large outputs.
- Migrated from Buffer to Uint8Array in web scrapers for better cross-platform compatibility and native browser support.
- Added getArtifactManager() lazy-initialization method to ToolSession for deferred artifact manager instantiation.
- Simplified API surface with wildcard exports from tools and session modules, reducing import complexity.
- Fixes deferred `--model` resolution to match extension-provided models before fallback
- Fixes CLI `--api-key` handling to support deferred model selection
- Adds OAuth provider support for extensions with source-scoped registration cleanup
- Adds custom API registration helpers with built-in collision checks
- Expands `Api` type to support extension-defined identifiers
- Adds tests for runtime provider registration and model selection