Commit Graph
88 Commits
Author SHA1 Message Date
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
roboomp 65a7444e94 fix(coding-agent/model-registry): preserved transport override across openrouter rediscovery
mergeDiscoveredModel re-applied baseUrl, headers, and compat from the
provider override when an existing bundled record matched, but dropped
`transport` because the raw /v1/models payload never carries one. With
`transport: pi-native` set for an auth-gateway-routed openrouter
provider in models.yml, the next /model switch after the background
catalog refresh picked the now-transport-less entry and routed through
the default openai-completions client to ${baseUrl}/chat/completions —
a path the auth-gateway never serves, so the gateway returned
"404 No route: POST /chat/completions".

Propagate transport with the same priority as baseUrl: provider
override wins, otherwise preserve the existing (boot-time override-
applied) value, otherwise the discovered value. Cover the regression
both at the merge unit level (xiaomi-tp-discovery-merge) and at the
ModelRegistry refresh level with a mock /models fetch.

Fixes #2555
2026-06-14 09:26:06 +00:00
can1357 69051241f1 fix(ai): added finite output-token fallback when model maxTokens is unknown
- Applied a finite 64,000-token fallback for OpenAI completion requests when model maxTokens is unavailable.
- Refactored stream max-token computation to use a shared helper and preserve a finite cap when caller and model limits are null.
- Added regression coverage for null maxTokens flows and bumped the model-cache schema to invalidate entries using retired unknown-limit sentinels.
2026-06-13 15:52:06 +02:00
can1357 f0c6a54f51 fix: handled unknown model limits as null to avoid artificial token caps
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
2026-06-13 15:35:40 +02:00
can1357 e78e936fb6 feat(coding-agent): kept retired variant-id selectors resolving after catalog collapsing
- Resolved retired effort-tier variant ids in `model-resolver.ts` through the hand-table aliases (`resolveVariantAlias`, `resolveBareVariantAlias`) plus the `X-thinking` → `X` grammar (`stripThinkingVariantToken`), with exact matches always winning while a raw id is live and explicit `:effort` suffixes transferring unchanged.
- Re-keyed models.yml `modelOverrides` and rate-limit selector suppressions from raw member ids onto the collapsed model in `model-registry.ts` (`normalizeSuppressedSelector`, lazy `hasLiveModel` checks so live raw ids keep their own overrides).
- Collapsed custom/config provider model lists at registry rebuild via `collapseBuiltModelVariants`, folding config-defined `X`/`X-thinking` twins into one entry.
- Extended `model-registry.test.ts` and `model-resolver.test.ts` with effort-tier variant collapsing and alias-resolution coverage.
2026-06-12 07:37:26 +02:00
can1357 d5b12494e8 test(packages/coding-agent): updated tests for declared effort filters
- Adjusted model-registry runtime-provider test expectations for effort-map filtering.
- Updated model-registry lookup test assertion to require declared effort levels only.
2026-06-12 03:46:03 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 ae415199dc feat: added build-time compatibility in ModelSpec/buildModel pipeline
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
2026-06-10 06:20:51 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 e18b901a8d refactor(packages/coding-agent): reorganized canonical model selection
- Replaced canonical-row resolution with getCanonicalModelSelections in model lists and selector flow.
- Hydrated model selector state from registry on construction and kept cached selections during refresh.
- Preserved highlighted and cached model selection when offline refresh completed or reordered models.
- Added parity checks between getCanonicalModelSelections and resolveCanonicalModel via registry tests.
2026-06-09 20:46:08 +02:00
can1357 eb1a46baf5 feat: added injectable fetch transport across AI and coding network flows
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
2026-06-09 04:51:17 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 78700a2c77 feat(ollama): added OLLAMA_HOST and OLLAMA_CONTEXT_LENGTH support
- Used OLLAMA_HOST for implicit discovery when OLLAMA_BASE_URL is unset.
- Applied OLLAMA_CONTEXT_LENGTH override to discovered context budgeting.
2026-06-05 22:47:40 +02:00
roboomp efc046a838 fix(ai): allow per-model opt-out of max_output_tokens on the wire
Add `Model.omitMaxOutputTokens` (`models.yml` model definitions and
`modelOverrides` accept the same field). When set, the openai-responses
and openai-completions providers stop emitting `max_output_tokens` /
`max_tokens` / `max_completion_tokens` so the upstream API applies its
own default cap. Catalog `maxTokens` is still honoured for local
budgeting (compaction, context promotion); only the wire field is
suppressed.

Restores the pre-v15.8.3 behaviour for Ollama proxies fronting cloud
catalogs (GLM, Kimi, DeepSeek): OMP cannot discover their true output
limit, so users previously set `maxTokens` to the Ollama context window
to unlock the model. v15.8.3 began sending that value as
`max_output_tokens`, triggering HTTP 400 from the upstream provider.

`applyCommonResponsesSamplingParams` now takes the model object instead
of a bare provider string so the wire-suppression flag is available to
the sampling-params builder.

Fixes #1881
2026-06-04 19:18:00 +00:00
roboomp fddcee5e80 fix(ai): honored anthropic cache compat
Allowed OpenAI-compatible providers to opt into Anthropic cache_control markers via compat.cacheControlFormat while preserving OpenRouter Anthropic defaults.

Fixes #1845
2026-06-04 11:20:08 +00:00
roboomp d28bee5db0 fix(providers): loaded special provider model cache
Loaded cached startup models for special built-in providers alongside standard provider descriptors so boot-time model resolution can see cached Google Antigravity, Gemini CLI, and OpenAI Codex discoveries before refresh.\n\nFixes #1721
2026-06-02 15:55:01 +00:00
Can BölükandGitHub 112602cda3 Merge branch 'main' into farm/a22ef9df/raise-discovered-model-max-tokens 2026-06-01 18:22:11 +03:00
can1357 5cd88b7abc feat(mnemosyne): added session-scoped visibility, graph tools, and safety hardening
- Filtered recall, fact, vector, temporal, and polyphonic voices to session-owned or explicitly global memories.
- Implemented graph_query and graph_link MCP tool handlers via EpisodicGraph; wired annotations and graph into Mnemosyne's external-db path.
- Fixed restore to stage and integrity-check before replacing the live database, rolling back on failure.
- Deduped generateId with a per-process nonce to prevent batch duplicate-content collisions.
2026-05-30 15:58:26 +02:00
roboomp 3cf2b54ace fix(coding-agent): raised auto-discovered model maxTokens cap to 32K
Auto-discovered OpenAI-compatible / Ollama / llama.cpp / new-api proxy
models defaulted to maxTokens: 8192 across four discovery branches in
model-registry.ts. When models hit the 8K output cap mid-stream on
legitimate large tool calls (write/edit payloads >~5KB), providers
dropped the streaming connection and Bun surfaced it as the opaque
'socket connection was closed unexpectedly'.

Extracted DISCOVERY_DEFAULT_MAX_TOKENS = 32_768 and pointed all four
discovery sites at it. min(contextWindow, ...) still honors smaller
advertised context windows on local Ollama/llama.cpp servers.

Fixes #1528
2026-05-30 05:55:10 +00:00
can1357 625c8b1992 test(ai): updated tests for model metadata and auth token fallback
- Mocked the Vertex stream E2E test to override the home directory and clear GOOGLE_APPLICATION_CREDENTIALS so token resolution uses metadata credentials instead of local ADC files.
- Updated wafer and model-registry test expectations to match current model metadata values (Qwen3.7 Max and claude-opus-4-8).
2026-05-29 06:56:29 +02:00
roboomp d87eeaa2ac fix(providers): pruned stale synthetic models
Treat Synthetic discovery as an authoritative catalog so deprecated bundled IDs are removed from resolved model lists and cache snapshots. Validate Synthetic API keys through the models endpoint instead of a model-specific chat request.\n\nFixes #1417
2026-05-26 23:59:51 +00:00
roboomp 016adfbee9 fix(model-registry): gated bundled vertex drop on fresh authoritative cache
Threaded cache freshness/authoritativeness through #loadCachedStandardProviderModels so dropProviderModels only fires when the cached Vertex project-catalog row is both fresh and authoritative. A stale or non-authoritative snapshot (e.g. after ADC discovery failure rewrote the row with authoritative=0) now keeps the bundled Gemini fallback in place, which would otherwise be the last working catalog in API-key-only environments.

Refs #1412
2026-05-26 17:00:48 +00:00
roboomp 8fc200f6e8 fix(providers): discovered vertex project models
Added Google Vertex OpenAI-compatible model discovery with ADC auth and treated authoritative Vertex project catalogs as replacements for bundled Gemini fallbacks in the model registry.

Fixes #1412
2026-05-26 16:34:53 +00:00
roboomp ccc08821d5 fix(providers): skipped disabled discovery probes
Prevent disabled providers from being registered for implicit local discovery and from creating built-in model discovery managers. Added regression coverage for disabled local providers during model registry refresh.

Fixes #1232
2026-05-20 20:16:19 +00:00
can1357 6120adbde2 perf(coding-agent): added PI_TIMING flag to log prompt durations
- Wrapped initial and subsequent print-mode prompts with `logger.time` for timing instrumentation.
- Printed collected timings after session run when `PI_TIMING` env var is set.
2026-05-17 01:33:16 +02:00
can1357 df1c1a6ba8 feat(auth): added auth-gateway forward-proxy and broker usage/migrate endpoints
- Added `omp auth-gateway serve/token/status` — a forward-proxy injecting broker credentials for OpenAI Chat, Anthropic Messages, and OpenAI Responses wire formats.
- Added `GET /v1/usage` to auth-broker and auth-gateway; usage cache switched to 5-min per-credential TTL with jitter and last-good fallback on failure.
- Added `AuthStorage.setConfigApiKey/removeConfigApiKey/clearConfigApiKeys` so `models.yml` `apiKey` beats OAuth tokens without overriding `--api-key`.
- Added `omp auth-broker migrate --from-local` for idempotent upload of local SQLite/env credentials to the broker.
2026-05-16 23:25:10 +02:00
can1357 45fe4df39e fix(ai): corrected AI tool handling via JSON-schema validation flow
- Replaced fromTypeBox conversion with a JSON-schema validator flow in ai tool handling and execution paths.
- Added recursive schema validation and expanded TypeBox checks for refs, enums, uniqueItems, and constraint keywords.
- Sanitized Azure/CCA tool schemas by dropping unsupported fields and rewriting oneOf tool branches as anyOf.
- Tightened argument and model-config validation, preserving unknown tool fields and adding apiKey plus compatibility flags.
2026-05-15 15:16:50 +02:00
roboomp b92b6fc7a9 fix(providers): loaded cached standard model discoveries
Loaded cached standard provider discovery models into ModelRegistry at startup so retry fallback validation can resolve Ollama Cloud models that are already visible through --list-models.

Added regression coverage for cached ollama-cloud fallback selectors and fixed a readonly notices type error exposed by the focused type check.

Fixes #1052
2026-05-15 01:22:26 +00:00
Can BölükandGitHub ed88f0f90d Merge branch 'main' into fix/context-window-fallback 2026-05-14 04:48:16 +02:00
can1357 f1f6516056 refactor: reorganized exports and removed obsolete helper branches
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
2026-05-14 04:36:19 +02:00
Burke T fcaafda0fa fix(coding-agent): preserve bundled contextWindow/maxTokens when discovery returns sentinel fallbacks
When cached or freshly-discovered provider models carry UNK_CONTEXT_WINDOW
(222222) / UNK_MAX_TOKENS (8888) sentinels, #mergeResolvedModels was
replacing the bundled model wholesale — wiping out the correct values.

Switch to a field-level merge that preserves the bundled model's
contextWindow and maxTokens when the replacement only has sentinel
fallbacks. Custom models (via #mergeCustomModels) already had this
protection via ?? fallback; provider discoveries didn't.

Fixes the TUI showing 222222/8888 instead of the real context/token
limits for discovered models.
2026-05-13 08:43:06 -03:00
can1357 975941aba4 chore: remove garbage tests 2026-05-12 04:09:33 +02:00
can1357 071f15895a fix: enable ollama cloud cache metadata
fixes #937
2026-05-06 20:32:17 +02:00
can1357 9b7843821a fix(coding-agent): honor authHeader provider overrides
Fixes #929
2026-05-06 20:31:47 +02:00
Christoph Gross f908ee9496 feat(coding-agent): add disableStrictTools provider option for anthropic-messages endpoints
Exposes model.compat.disableStrictTools (already supported by the anthropic
transport since #826) via models.yml so users can configure it without code
changes.

Set disableStrictTools: true at the provider level to disable strict tool
schemas for third-party Anthropic-compatible endpoints (AWS Bedrock, Vertex
AI proxies, custom gateways) that reject the strict field.

- Add disableStrictTools to ProviderConfigSchema
- Merge { disableStrictTools: true } into provider compat override when set,
  flowing through the existing compat pipeline to model.compat.disableStrictTools
- disableStrictTools alone is sufficient for an override-only provider entry
- Update docs/models.md with field reference, Bedrock example, and proxy note
- Add tests covering provider-level propagation, built-in override, and
  overlay merge
2026-04-30 08:28:03 +02:00
can1357 bf1faf8842 test(coding-agent): drop api filter from getOpenAICompat fixture helper
The helper guarded `model.api === "openai-completions"` and returned undefined
for openai-responses models. The discoverable-custom-compat test sets
`api: "openai-responses"` on a custom model with `compat.extraBody`, so the
post-refresh assertion saw `undefined` instead of the configured proxy hint.

The OpenAICompatSchema gates user-facing custom-model compat regardless of the
underlying api wire format, so reading the field as OpenAICompat for any api
matches what the registry actually stores.
2026-04-30 05:35:00 +02:00
can1357 fed95ce524 fix(ai,coding-agent): narrow Model.compat consumers after AnthropicCompat split
Commit a190397d8 made `Model.compat` resolve to `OpenAICompat | AnthropicCompat`
under the default `TApi = any`. The widened union broke every site that treated
`compat` as openai-shaped: model-registry deep-merge, openai-completions resolved
compat, and ~20 test fixtures. This restores the assumption locally instead of
papering over it with casts.

- getBundledModel is now generic on TApi so test fixtures that spread it into
  `Model<"openai-completions">` get the narrow compat back.
- mergeCompat is generic over TBase/TOverride; the schema-driven model-registry
  override path keeps its OpenAICompat-shaped merge fields, anthropic overrides
  pass through untouched.
- OpenAICompatSchema gains the openai-only fields it was missing
  (requiresMistralToolIds, reasoningContentField, requiresReasoningContent*,
  thinkingFormat, requiresThinkingAsText, disableReasoningOnForcedToolChoice).
- resolveOpenAICompat fills in disableReasoningOnForcedToolChoice so the
  Required<OpenAICompat> shape stays satisfied.
- Anthropic tool-result block id assignment uses the proper unknown double-cast.
- isForcedToolChoice accepts unknown so it can read `params.tool_choice` whose
  type comes from the OpenAI SDK ChatCompletionToolChoiceOption (now wider than
  our local OpenAICompletionsToolChoice).
- Test fixtures and Required<OpenAICompat> literals updated for the field set.

Fixes CI red on main.
2026-04-30 05:23:20 +02:00
can1357 40971d9675 feat(coding-agent): marked custom Anthropic models as OAuth-shaped by default
- Added an `isOAuth` model flag and passed it through Anthropic stream calls to force OAuth-shaped request options.
- Extended model-registry config handling with an `auth: oauth` mode and default `isOAuth` resolution for `anthropic-messages` providers while allowing explicit `apiKey` auth to remain unset.
- Added tests covering OAuth defaults and opt-outs for Anthropic and non-Anthropic custom providers.
2026-04-29 19:07:53 +02:00
Can BölükandGitHub 6e711b9da1 Merge branch 'main' into feat/ollama-provider 2026-04-26 13:35:27 +02:00
Aidan d12d1577a8 Allow headers-only provider overrides 2026-04-24 11:24:49 -04:00
Corentin AZAISandcan1357 cc33fe5366 test(ai): add Opus 4.7 catalog and alignment coverage
- Register claude-opus-4-7 model entry in models.json
- Cover adaptive thinking/sampling payload shape in anthropic-alignment test
- Assert Opus 4.7 surfaces in ModelRegistry available models
2026-04-24 07:29:48 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 c4ea6c920f fix: restore Copilot prompt budgets and align task/model-registry expectations with opus 4.7
Three fixes to make CI green after the opus 4.7 and auto-bump landed:

1. github-copilot model mapper: prefer capabilities.limits.max_prompt_tokens
   over the root-level context_length field (which mirrors max_context_window_tokens, i.e.
   total window). Copilot's real /models response returns both for the gpt-5.x family, and
   context_length inflates contextWindow with the output budget. Also restore the bundled
   Copilot limits (claude-opus-4.6, gpt-5.2, gpt-5.4, gpt-5.4-mini, grok-code-fast-1) to
   the values the fixed mapper produces so tests that depend on truthful offline fallbacks
   pass. Update the two Copilot discovery tests whose payloads conflated context_length
   with prompt capacity.

2. coding-agent task schema: make the per-task assignment description context-mode-aware.
   The previous description unconditionally told agents that 'shared background belongs
   in context', which is wrong for independent mode where shared context is disabled.

3. coding-agent model-registry test: update the anthropic-latest canonical collapse case
   to claude-opus-4-7 since opus 4.7 is now the newest official opus in models.json.
2026-04-17 19:24:03 +02:00
Ronny Unger 03aad88db7 feat(ai): add Ollama Cloud provider with streaming, thinking, and tool support
Add ollama-chat API provider supporting:
- Streaming chat completions via Ollama Cloud (ollama.com) API
- API key authentication via OLLAMA_API_KEY env var
- Thinking/reasoning model support with configurable effort levels
- Tool calling with streaming JSON argument assembly
- Dynamic model discovery via /api/tags and /api/show metadata
- Context window detection from model info
- Login flow via pi login ollama-cloud
2026-04-15 07:43:14 +02:00
can1357 212d56bc11 feat: added strict-mode fallback for OpenAI tool calls with all_strict
- Added `toolStrictMode` support with `all_strict`/`none`/`mixed` options to OpenAI compatibility.
- Fixed OpenAI-completion strict-mode flows by capturing failed HTTP responses and retrying once as non-strict.
- Fixed completion error reporting by surfacing captured status, headers, and JSON `type`/`param`/`code` details.
- Improved strict-schema enforcement with WeakMap memoization and circular-schema detection in sanitization.
- Fixed OpenRouter provider lookup by resolving fallback model IDs for suffix and date variants in registry resolution.
- Refactored benchmark tooling and added async RPC error-window tracking for scheduled run execution.
2026-04-13 15:46:06 +02:00
can1357 5277e44139 feat(coding-agent): added canonical aliases for model role resolution
- Added canonical model equivalence types, cache helpers, and registry APIs for provider variant lookup.
- Changed model resolution to apply canonical ID overrides/excludes with provider order before fallback matching.
- Added canonical and provider model views in list-models and selector UI with canonical sorting/persistence.
- Updated role/model persistence to store selectors while runtime now resolves concrete canonical-backed provider models.
2026-04-11 08:28:50 +02:00
inprealphaandGitHub 4a4032dbea Merge branch 'main' into feat/opencode-oauth-copilot 2026-04-10 20:41:36 +05:30
can1357 c0ba48cb68 fix(coding-agent): normalize cached Ollama discovery transport 2026-04-09 18:57:52 +02:00
can1357 69c915e9d8 fix(ai): preserve Copilot enterprise metadata in peeked credentials 2026-04-09 18:57:51 +02:00
Abir Biswasandcan1357 bafa3b8a06 fix: github.com enterprise routing and structured Copilot OAuth credentials 2026-04-09 18:47:08 +02:00