bf75d80836
Thin OpenAI-compatible proxies that omit context_length / max_model_len on /v1/models made every discovered model fall back to DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a bundled model with a much larger intrinsic window. discoverProxyModels and discoverLiteLLMModels already resolve ids against the bundled reference index; discoverOpenAIModelsList (which also backs lm-studio discovery) now does the same. Behavior: - Build the reference index once outside the loop and resolve each item via resolveModelReference(). - contextWindow precedence keeps provider-reported values authoritative: item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow ?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW. - maxTokens uses reference?.maxTokens when available, otherwise the api-specific discovery default, capped at contextWindow so a bundled ref for a larger sibling can never over-request output tokens. - name / reasoning / thinking / input inherit from the reference; native lm-studio metadata still wins for input modality. - Provider-specific baseUrl, headers, and local-unknown cost stay local. - OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole / supportsReasoningEffort all false) to match the proxy sibling. Also updated two pre-existing regression tests that used deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional" ids to exercise the default-fallback branch. Those model names have since been added to the bundled catalog, so the tests were renamed to vllm-lab-fork-* ids that unambiguously miss the reference index while preserving each test's original default-fallback intent. Fixes #3983
@oh-my-pi/pi-coding-agent
Core implementation package for the omp coding agent in the oh-my-pi monorepo.
For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:
Package-specific references:
Memory backends
The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):
off(default) — no memory subsystem runs.local— existing rollout-summarisation pipeline; writesmemory_summary.mdand consolidated artifacts under the agent dir.hindsight— talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposesretain,recall, andreflect.
Hindsight quickstart
- Run a Hindsight server (Cloud or
docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest). - Set
memory.backend = "hindsight"andhindsight.apiUrl = "http://localhost:8888"(or your Cloud URL). - Optional environment overrides (env wins over settings):
HINDSIGHT_API_URL,HINDSIGHT_API_TOKEN— connectionHINDSIGHT_BANK_ID,HINDSIGHT_DYNAMIC_BANK_ID,HINDSIGHT_AGENT_NAME— bank addressingHINDSIGHT_AUTO_RECALL,HINDSIGHT_AUTO_RETAIN,HINDSIGHT_RETAIN_MODE— lifecycleHINDSIGHT_RECALL_BUDGET,HINDSIGHT_RECALL_MAX_TOKENS— recall sizingHINDSIGHT_BANK_MISSION,HINDSIGHT_DEBUG
Switching backends mid-session is honoured on the next system-prompt rebuild and the next /memory slash command. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch.