f6ca76728b
Wafer (https://wafer.ai) exposes a single OpenAI-compatible endpoint (`https://pass.wafer.ai/v1`) for two SKUs whose entitlement differs server-side, so we model them as two parallel providers — mirroring the firepass/fireworks split so a user with both subscriptions can switch without re-pasting: - `wafer-pass` — flat-rate. `/v1/models` is filtered to entries whose `wafer.tier === "pass_included"`. - `wafer-serverless` — pay-as-you-go superset of Pass. Both issue `wfr_…` keys. `/login wafer-pass` and `/login wafer-serverless` paste-and-validate via `/v1/models`. `WAFER_PASS_API_KEY` and `WAFER_SERVERLESS_API_KEY` are wired through `getEnvApiKey`. Bundled catalog: - `wafer-pass`: GLM-5.1, Qwen3.5-397B-A17B. - `wafer-serverless`: GLM-5.1, Qwen3.5-397B-A17B, Kimi-K2.6, Qwen3.6-35B-A3B. Dynamic discovery via `/v1/models` overlays additional models at runtime and folds the `wafer` envelope (tier, capabilities, cents/M pricing) into the canonical `Model<"openai-completions">` shape. GLM-family entries carry the zai-style thinking compat (`thinkingFormat: "zai"`, `reasoningContentField: "reasoning_content"`) so reasoning tokens land in the right field. Cents-per-million → dollars-per-million via /100. Tests (`packages/ai/test/wafer.test.ts`, 5 cases): bundled catalog contract for both providers and wire-id pass-through (case-sensitive, no rewrite — `GLM-5.1` must round-trip verbatim or upstream 404s). Optional `packages/ai/test/wafer.live.ts` exercises a real round-trip against `pass.wafer.ai` when `WAFER_PASS_API_KEY` is set.
@oh-my-pi/pi-coding-agent
Core implementation package for the omp coding agent in the oh-my-pi monorepo.
For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:
Package-specific references:
- CHANGELOG
- MCP configuration guide
- MCP runtime lifecycle
- MCP server/tool authoring
- DEVELOPMENT
- RenderMermaid guide
Memory backends
The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):
off(default) — no memory subsystem runs.local— existing rollout-summarisation pipeline; writesmemory_summary.mdand consolidated artifacts under the agent dir.hindsight— talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposesretain,recall, andreflect.
Hindsight quickstart
- Run a Hindsight server (Cloud or
docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest). - Set
memory.backend = "hindsight"andhindsight.apiUrl = "http://localhost:8888"(or your Cloud URL). - Optional environment overrides (env wins over settings):
HINDSIGHT_API_URL,HINDSIGHT_API_TOKEN— connectionHINDSIGHT_BANK_ID,HINDSIGHT_DYNAMIC_BANK_ID,HINDSIGHT_AGENT_NAME— bank addressingHINDSIGHT_AUTO_RECALL,HINDSIGHT_AUTO_RETAIN,HINDSIGHT_RETAIN_MODE— lifecycleHINDSIGHT_RECALL_BUDGET,HINDSIGHT_RECALL_MAX_TOKENS— recall sizingHINDSIGHT_BANK_MISSION,HINDSIGHT_DEBUG
Switching backends mid-session is honoured on the next system-prompt rebuild and the next /memory slash command. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch.