Commit Graph
104 Commits
Author SHA1 Message Date
can1357 74d7dfd2fa Merge PR #3354: fix: migrate coding-agent tests to removeWithRetries (@oldschoola) 2026-06-27 02:06:38 +02:00
can1357 577d2a8eb8 style: biome format/organize-imports across integrated PRs 2026-06-27 02:06:38 +02:00
can1357 0b8206a46d fix(compaction): preserve provider defaults for remote compaction 2026-06-27 01:39:34 +02:00
can1357 14cc9cba0f Merge PR #3106: fix(compaction): enable custom provider remote compaction (@roboomp) 2026-06-27 01:39:34 +02:00
oldschoola a2854ba768 fix: migrate coding-agent tests from fs.rm to removeWithRetries
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.

The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
2026-06-23 15:28:05 -07:00
roboomp b7aefe0689 fix(compaction): enabled custom remote compaction
- Added provider/model remoteCompaction metadata and models.yml propagation.\n- Routed configured OpenAI-compatible compaction endpoints for custom providers.\n- Added compactionModel as a summary-only model selector that leaves the active session model unchanged.\n\nFixes #3104
2026-06-20 07:27:09 +00:00
Alexander Kirilin eaf9248ec2 fix(catalog): merge main into MiMo efforts 2026-06-20 03:27:05 -04:00
roboomp 48d64210f9 fix(config): allowed responses image detail compat
Added the missing models.yml schema field for compat.supportsImageDetailOriginal so custom Responses-compatible proxies can opt out of snapcompact's native-resolution image hint.

Covered the CC Switch-style custom provider override and the Codex Responses wire clamp from original to auto.

Fixes #3092
2026-06-20 02:27:57 +00:00
Alexander Kirilin 057b39fc46 fix(catalog): resolve MiMo branch conflicts 2026-06-19 20:28:05 -04:00
can1357 052b6b0c61 Merge PR #3006: fix(providers): accept Bedrock inference profile ARNs (@roboomp) 2026-06-19 17:16:18 +02:00
can1357 40101bd3eb Merge PR #3043: fix(catalog): omit Ollama Cloud output caps (@wolfiesch)
# Conflicts:
#	packages/catalog/src/models.json
2026-06-19 17:15:43 +02:00
can1357 d46bd0139b refactor(coding-agent): fixed system prompt customization path
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
2026-06-19 16:16:37 +02:00
Wolfgang Schoenberger 99c90a2ef1 fix(coding-agent): normalize cached Ollama Cloud output caps 2026-06-19 03:56:02 -07:00
Wolfgang Schoenberger 409af331e7 fix(catalog): omit Ollama Cloud output caps 2026-06-19 02:51:23 -07:00
roboomp 491fbeb2bb test(providers): covered bedrock profile restore lookup
Added regression coverage for ModelRegistry.find restoring persisted Amazon Bedrock inference profile ARN models through the synthetic resolver path.

Fixes #3004
2026-06-18 22:34:38 +00:00
Alexander Kirilin 0e415efeea test(coding-agent): avoid antigravity cache alias collision 2026-06-18 16:38:22 -04:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
roboomp 65a7444e94 fix(coding-agent/model-registry): preserved transport override across openrouter rediscovery
mergeDiscoveredModel re-applied baseUrl, headers, and compat from the
provider override when an existing bundled record matched, but dropped
`transport` because the raw /v1/models payload never carries one. With
`transport: pi-native` set for an auth-gateway-routed openrouter
provider in models.yml, the next /model switch after the background
catalog refresh picked the now-transport-less entry and routed through
the default openai-completions client to ${baseUrl}/chat/completions —
a path the auth-gateway never serves, so the gateway returned
"404 No route: POST /chat/completions".

Propagate transport with the same priority as baseUrl: provider
override wins, otherwise preserve the existing (boot-time override-
applied) value, otherwise the discovered value. Cover the regression
both at the merge unit level (xiaomi-tp-discovery-merge) and at the
ModelRegistry refresh level with a mock /models fetch.

Fixes #2555
2026-06-14 09:26:06 +00:00
can1357 69051241f1 fix(ai): added finite output-token fallback when model maxTokens is unknown
- Applied a finite 64,000-token fallback for OpenAI completion requests when model maxTokens is unavailable.
- Refactored stream max-token computation to use a shared helper and preserve a finite cap when caller and model limits are null.
- Added regression coverage for null maxTokens flows and bumped the model-cache schema to invalidate entries using retired unknown-limit sentinels.
2026-06-13 15:52:06 +02:00
can1357 f0c6a54f51 fix: handled unknown model limits as null to avoid artificial token caps
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
2026-06-13 15:35:40 +02:00
can1357 e78e936fb6 feat(coding-agent): kept retired variant-id selectors resolving after catalog collapsing
- Resolved retired effort-tier variant ids in `model-resolver.ts` through the hand-table aliases (`resolveVariantAlias`, `resolveBareVariantAlias`) plus the `X-thinking` → `X` grammar (`stripThinkingVariantToken`), with exact matches always winning while a raw id is live and explicit `:effort` suffixes transferring unchanged.
- Re-keyed models.yml `modelOverrides` and rate-limit selector suppressions from raw member ids onto the collapsed model in `model-registry.ts` (`normalizeSuppressedSelector`, lazy `hasLiveModel` checks so live raw ids keep their own overrides).
- Collapsed custom/config provider model lists at registry rebuild via `collapseBuiltModelVariants`, folding config-defined `X`/`X-thinking` twins into one entry.
- Extended `model-registry.test.ts` and `model-resolver.test.ts` with effort-tier variant collapsing and alias-resolution coverage.
2026-06-12 07:37:26 +02:00
can1357 d5b12494e8 test(packages/coding-agent): updated tests for declared effort filters
- Adjusted model-registry runtime-provider test expectations for effort-map filtering.
- Updated model-registry lookup test assertion to require declared effort levels only.
2026-06-12 03:46:03 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 ae415199dc feat: added build-time compatibility in ModelSpec/buildModel pipeline
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
2026-06-10 06:20:51 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 e18b901a8d refactor(packages/coding-agent): reorganized canonical model selection
- Replaced canonical-row resolution with getCanonicalModelSelections in model lists and selector flow.
- Hydrated model selector state from registry on construction and kept cached selections during refresh.
- Preserved highlighted and cached model selection when offline refresh completed or reordered models.
- Added parity checks between getCanonicalModelSelections and resolveCanonicalModel via registry tests.
2026-06-09 20:46:08 +02:00
can1357 eb1a46baf5 feat: added injectable fetch transport across AI and coding network flows
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
2026-06-09 04:51:17 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 78700a2c77 feat(ollama): added OLLAMA_HOST and OLLAMA_CONTEXT_LENGTH support
- Used OLLAMA_HOST for implicit discovery when OLLAMA_BASE_URL is unset.
- Applied OLLAMA_CONTEXT_LENGTH override to discovered context budgeting.
2026-06-05 22:47:40 +02:00
roboomp efc046a838 fix(ai): allow per-model opt-out of max_output_tokens on the wire
Add `Model.omitMaxOutputTokens` (`models.yml` model definitions and
`modelOverrides` accept the same field). When set, the openai-responses
and openai-completions providers stop emitting `max_output_tokens` /
`max_tokens` / `max_completion_tokens` so the upstream API applies its
own default cap. Catalog `maxTokens` is still honoured for local
budgeting (compaction, context promotion); only the wire field is
suppressed.

Restores the pre-v15.8.3 behaviour for Ollama proxies fronting cloud
catalogs (GLM, Kimi, DeepSeek): OMP cannot discover their true output
limit, so users previously set `maxTokens` to the Ollama context window
to unlock the model. v15.8.3 began sending that value as
`max_output_tokens`, triggering HTTP 400 from the upstream provider.

`applyCommonResponsesSamplingParams` now takes the model object instead
of a bare provider string so the wire-suppression flag is available to
the sampling-params builder.

Fixes #1881
2026-06-04 19:18:00 +00:00
roboomp fddcee5e80 fix(ai): honored anthropic cache compat
Allowed OpenAI-compatible providers to opt into Anthropic cache_control markers via compat.cacheControlFormat while preserving OpenRouter Anthropic defaults.

Fixes #1845
2026-06-04 11:20:08 +00:00
roboomp d28bee5db0 fix(providers): loaded special provider model cache
Loaded cached startup models for special built-in providers alongside standard provider descriptors so boot-time model resolution can see cached Google Antigravity, Gemini CLI, and OpenAI Codex discoveries before refresh.\n\nFixes #1721
2026-06-02 15:55:01 +00:00
Can BölükandGitHub 112602cda3 Merge branch 'main' into farm/a22ef9df/raise-discovered-model-max-tokens 2026-06-01 18:22:11 +03:00
can1357 5cd88b7abc feat(mnemosyne): added session-scoped visibility, graph tools, and safety hardening
- Filtered recall, fact, vector, temporal, and polyphonic voices to session-owned or explicitly global memories.
- Implemented graph_query and graph_link MCP tool handlers via EpisodicGraph; wired annotations and graph into Mnemosyne's external-db path.
- Fixed restore to stage and integrity-check before replacing the live database, rolling back on failure.
- Deduped generateId with a per-process nonce to prevent batch duplicate-content collisions.
2026-05-30 15:58:26 +02:00
roboomp 3cf2b54ace fix(coding-agent): raised auto-discovered model maxTokens cap to 32K
Auto-discovered OpenAI-compatible / Ollama / llama.cpp / new-api proxy
models defaulted to maxTokens: 8192 across four discovery branches in
model-registry.ts. When models hit the 8K output cap mid-stream on
legitimate large tool calls (write/edit payloads >~5KB), providers
dropped the streaming connection and Bun surfaced it as the opaque
'socket connection was closed unexpectedly'.

Extracted DISCOVERY_DEFAULT_MAX_TOKENS = 32_768 and pointed all four
discovery sites at it. min(contextWindow, ...) still honors smaller
advertised context windows on local Ollama/llama.cpp servers.

Fixes #1528
2026-05-30 05:55:10 +00:00
can1357 625c8b1992 test(ai): updated tests for model metadata and auth token fallback
- Mocked the Vertex stream E2E test to override the home directory and clear GOOGLE_APPLICATION_CREDENTIALS so token resolution uses metadata credentials instead of local ADC files.
- Updated wafer and model-registry test expectations to match current model metadata values (Qwen3.7 Max and claude-opus-4-8).
2026-05-29 06:56:29 +02:00
roboomp d87eeaa2ac fix(providers): pruned stale synthetic models
Treat Synthetic discovery as an authoritative catalog so deprecated bundled IDs are removed from resolved model lists and cache snapshots. Validate Synthetic API keys through the models endpoint instead of a model-specific chat request.\n\nFixes #1417
2026-05-26 23:59:51 +00:00
roboomp 016adfbee9 fix(model-registry): gated bundled vertex drop on fresh authoritative cache
Threaded cache freshness/authoritativeness through #loadCachedStandardProviderModels so dropProviderModels only fires when the cached Vertex project-catalog row is both fresh and authoritative. A stale or non-authoritative snapshot (e.g. after ADC discovery failure rewrote the row with authoritative=0) now keeps the bundled Gemini fallback in place, which would otherwise be the last working catalog in API-key-only environments.

Refs #1412
2026-05-26 17:00:48 +00:00
roboomp 8fc200f6e8 fix(providers): discovered vertex project models
Added Google Vertex OpenAI-compatible model discovery with ADC auth and treated authoritative Vertex project catalogs as replacements for bundled Gemini fallbacks in the model registry.

Fixes #1412
2026-05-26 16:34:53 +00:00
roboomp ccc08821d5 fix(providers): skipped disabled discovery probes
Prevent disabled providers from being registered for implicit local discovery and from creating built-in model discovery managers. Added regression coverage for disabled local providers during model registry refresh.

Fixes #1232
2026-05-20 20:16:19 +00:00
can1357 6120adbde2 perf(coding-agent): added PI_TIMING flag to log prompt durations
- Wrapped initial and subsequent print-mode prompts with `logger.time` for timing instrumentation.
- Printed collected timings after session run when `PI_TIMING` env var is set.
2026-05-17 01:33:16 +02:00
can1357 df1c1a6ba8 feat(auth): added auth-gateway forward-proxy and broker usage/migrate endpoints
- Added `omp auth-gateway serve/token/status` — a forward-proxy injecting broker credentials for OpenAI Chat, Anthropic Messages, and OpenAI Responses wire formats.
- Added `GET /v1/usage` to auth-broker and auth-gateway; usage cache switched to 5-min per-credential TTL with jitter and last-good fallback on failure.
- Added `AuthStorage.setConfigApiKey/removeConfigApiKey/clearConfigApiKeys` so `models.yml` `apiKey` beats OAuth tokens without overriding `--api-key`.
- Added `omp auth-broker migrate --from-local` for idempotent upload of local SQLite/env credentials to the broker.
2026-05-16 23:25:10 +02:00
can1357 45fe4df39e fix(ai): corrected AI tool handling via JSON-schema validation flow
- Replaced fromTypeBox conversion with a JSON-schema validator flow in ai tool handling and execution paths.
- Added recursive schema validation and expanded TypeBox checks for refs, enums, uniqueItems, and constraint keywords.
- Sanitized Azure/CCA tool schemas by dropping unsupported fields and rewriting oneOf tool branches as anyOf.
- Tightened argument and model-config validation, preserving unknown tool fields and adding apiKey plus compatibility flags.
2026-05-15 15:16:50 +02:00
roboomp b92b6fc7a9 fix(providers): loaded cached standard model discoveries
Loaded cached standard provider discovery models into ModelRegistry at startup so retry fallback validation can resolve Ollama Cloud models that are already visible through --list-models.

Added regression coverage for cached ollama-cloud fallback selectors and fixed a readonly notices type error exposed by the focused type check.

Fixes #1052
2026-05-15 01:22:26 +00:00
Can BölükandGitHub ed88f0f90d Merge branch 'main' into fix/context-window-fallback 2026-05-14 04:48:16 +02:00
can1357 f1f6516056 refactor: reorganized exports and removed obsolete helper branches
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
2026-05-14 04:36:19 +02:00
Burke T fcaafda0fa fix(coding-agent): preserve bundled contextWindow/maxTokens when discovery returns sentinel fallbacks
When cached or freshly-discovered provider models carry UNK_CONTEXT_WINDOW
(222222) / UNK_MAX_TOKENS (8888) sentinels, #mergeResolvedModels was
replacing the bundled model wholesale — wiping out the correct values.

Switch to a field-level merge that preserves the bundled model's
contextWindow and maxTokens when the replacement only has sentinel
fallbacks. Custom models (via #mergeCustomModels) already had this
protection via ?? fallback; provider discoveries didn't.

Fixes the TUI showing 222222/8888 instead of the real context/token
limits for discovered models.
2026-05-13 08:43:06 -03:00
can1357 975941aba4 chore: remove garbage tests 2026-05-12 04:09:33 +02:00
can1357 071f15895a fix: enable ollama cloud cache metadata
fixes #937
2026-05-06 20:32:17 +02:00
can1357 9b7843821a fix(coding-agent): honor authHeader provider overrides
Fixes #929
2026-05-06 20:31:47 +02:00