Files
oh-my-pi/packages/coding-agent
bench-local f6ca76728b feat(ai): add Wafer Pass and Wafer Serverless providers
Wafer (https://wafer.ai) exposes a single OpenAI-compatible endpoint
(`https://pass.wafer.ai/v1`) for two SKUs whose entitlement differs
server-side, so we model them as two parallel providers — mirroring the
firepass/fireworks split so a user with both subscriptions can switch
without re-pasting:

- `wafer-pass` — flat-rate. `/v1/models` is filtered to entries whose
  `wafer.tier === "pass_included"`.
- `wafer-serverless` — pay-as-you-go superset of Pass.

Both issue `wfr_…` keys. `/login wafer-pass` and `/login wafer-serverless`
paste-and-validate via `/v1/models`. `WAFER_PASS_API_KEY` and
`WAFER_SERVERLESS_API_KEY` are wired through `getEnvApiKey`.

Bundled catalog:
- `wafer-pass`: GLM-5.1, Qwen3.5-397B-A17B.
- `wafer-serverless`: GLM-5.1, Qwen3.5-397B-A17B, Kimi-K2.6, Qwen3.6-35B-A3B.

Dynamic discovery via `/v1/models` overlays additional models at runtime
and folds the `wafer` envelope (tier, capabilities, cents/M pricing) into
the canonical `Model<"openai-completions">` shape. GLM-family entries
carry the zai-style thinking compat (`thinkingFormat: "zai"`,
`reasoningContentField: "reasoning_content"`) so reasoning tokens land in
the right field. Cents-per-million → dollars-per-million via /100.

Tests (`packages/ai/test/wafer.test.ts`, 5 cases): bundled catalog
contract for both providers and wire-id pass-through (case-sensitive,
no rewrite — `GLM-5.1` must round-trip verbatim or upstream 404s).
Optional `packages/ai/test/wafer.live.ts` exercises a real round-trip
against `pass.wafer.ai` when `WAFER_PASS_API_KEY` is set.
2026-05-27 20:45:33 -07:00
..
2026-05-27 19:01:33 +02:00
2026-05-09 03:11:18 +02:00

@oh-my-pi/pi-coding-agent

Core implementation package for the omp coding agent in the oh-my-pi monorepo.

For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:

Package-specific references:

Memory backends

The agent supports three mutually-exclusive memory backends, selected via the memory.backend setting (Settings → Memory tab, or ~/.omp/config.yml):

  • off (default) — no memory subsystem runs.
  • local — existing rollout-summarisation pipeline; writes memory_summary.md and consolidated artifacts under the agent dir.
  • hindsight — talks to a Hindsight server (Cloud or self-hosted Docker), retains transcripts every Nth user turn, recalls memories on the first turn of a session, and exposes retain, recall, and reflect.

Hindsight quickstart

  1. Run a Hindsight server (Cloud or docker run -p 8888:8888 ghcr.io/vectorize-io/hindsight:latest).
  2. Set memory.backend = "hindsight" and hindsight.apiUrl = "http://localhost:8888" (or your Cloud URL).
  3. Optional environment overrides (env wins over settings):
    • HINDSIGHT_API_URL, HINDSIGHT_API_TOKEN — connection
    • HINDSIGHT_BANK_ID, HINDSIGHT_DYNAMIC_BANK_ID, HINDSIGHT_AGENT_NAME — bank addressing
    • HINDSIGHT_AUTO_RECALL, HINDSIGHT_AUTO_RETAIN, HINDSIGHT_RETAIN_MODE — lifecycle
    • HINDSIGHT_RECALL_BUDGET, HINDSIGHT_RECALL_MAX_TOKENS — recall sizing
    • HINDSIGHT_BANK_MISSION, HINDSIGHT_DEBUG

Switching backends mid-session is honoured on the next system-prompt rebuild and the next /memory slash command. Existing users with memories.enabled = true|false are migrated to memory.backend = "local"|"off" exactly once on first launch.