- Implemented a unified benchmark normalization layer to handle metrics and artifacts across harbor, edit, and snapcompact benchmarks. - Integrated the TypeScript edit benchmark directly into the manager, migrating logic and deprecating the standalone package. - Updated the API and database schema to support standardized benchmark configurations, metrics, and trace-based reporting. - Enhanced the UI to visualize comparative performance metrics, including pass rate, cost, and latency deltas for benchmark runs.
@oh-my-pi/harbor-manager
One manager for repository benchmarks. Harbor, TypeScript edit, and SnapCompact runs use the same experiment → run → trace model, SQLite store, REST/SSE API, and dashboard. Benchmark-native artifacts remain on disk; adapters normalize their live progress, scores, token usage, costs, and traces.
# Dashboard + API on :4700; launch every benchmark from the same “new run” form
bun run serve --port 4700
How Harbor runs execute
- Local omp, not npm. By default the runner bind-mounts the repo
read-only into each task container (
--install source) and runs omp straight frompackages/coding-agent/src/cli.ts— TS edits apply to the next trial with no rebuild. A cached linuxnode_modulestree (built once per lockfile change insideoven/bun, stored in<jobs-dir>/_bench/_deps/) shadows the host's darwin one, and a linuxbunbinary is mounted at/opt/omp/bin— so trial setup needs zero outbound network. Alternatives:--install local(pack a tarball per run) or--binary(prebuiltdist/omp-linux-*self-contained binaries). - Auth never enters containers. A generated
models.ymlroutes providerbaseUrls at the host pm2 auth-gateway; the gateway resolves credentials host-side. - Harbor owns trials. The runner/serve layer polls each trial's
result.jsonfor progress, spend, and outcomes.
Server
-
GET /— experiments, runs, normalized traces, and a launch form for every benchmark. -
GET /api/experiments— experiment summaries across all benchmark types. -
GET /api/runs— uniform run rows with benchmark, score, progress, spend, and tokens. -
POST /api/runs— launch through a benchmark adapter. Body:{ "benchmark": "edit", "model": "anthropic/claude-opus-4-8", "tasks": 20, "concurrency": 4, "attempts": 2, "jobName": "edit-baseline", "role": "baseline", "goal": "compare edit strategies" }benchmarkisharbor,edit, orsnapcompact. Harbor usesdataset,include,timeoutMultiplier, andslide; edit usesincludeas task IDs; SnapCompact usesconditionsand treatstasksas the passage limit. -
GET /api/runs/:name—{ run, traces }(syncs native artifacts on read). -
DELETE /api/runs/:name— cancel a manager-launched run. -
GET /api/runs/:name/traces/:trace[?raw=1]— normalized or native trace. -
GET /api/events— SSE stream of run-list snapshots (sent on change).
State lives in <jobs-dir>/_manager/harbor-manager.sqlite; the filesystem
stays the source of truth and historical CLI runs are auto-discovered.
Harbor runner options (excerpt)
| Option | Default | Notes |
|---|---|---|
-m, --model <provider/model> |
anthropic/claude-sonnet-4-6 |
Repeatable |
-l, --tasks <N> |
20 |
Max tasks |
-n, --concurrency <N> |
4 |
Concurrent trials |
-k, --attempts <N> |
1 |
Attempts per task (pass@k) |
-d, --dataset <name> |
terminal-bench@2.0 |
Any Harbor dataset id |
-i/-x, --include/--exclude <glob> |
— | Task filters (repeatable) |
--timeout-multiplier <x> |
— | Scales task agent/verifier timeouts |
--agent-arg <arg> |
— | Extra arg forwarded verbatim to the in-container omp CLI (repeatable) |
--env <KEY[=VALUE]> |
— | Forward env into the omp container (repeatable); KEY alone forwards the host value |
--binary <path> |
— | Prebuilt omp binary (repeat for arm64+x64) |
--install <source|local|published> |
source |
source = repo bind-mount, local = tarball pack, published = npm @oh-my-pi/pi-coding-agent |
--gateway-url <url> |
http://host.docker.internal:4000 |
|
--no-gateway |
off | Pass host provider keys into containers instead |
-o, --jobs-dir <path> |
<repo>/runs/harbor |
Shared with the server |
--dry-run |
off | Print the harbor command + models.yml and exit |
Outputs
<jobs-dir>/<jobName>/— Harbor trial dirs (result.jsonper trial).<jobs-dir>/_bench/<jobName>/report.md— markdown summary table.<jobs-dir>/_bench/<jobName>/harbor.log— full Harbor output.<jobs-dir>/_manager/logs/<jobName>.log— runner output for API-launched runs.
Caveats
- Network policy. On Harbor's local Docker backend only public registries work; task containers reach models via the host gateway.
--install sourcereflects local TS changes with no rebuild, but Rust natives load from the in-treepackages/natives/native/pi_natives.linux-*.nodeprebuilds — rebuild those when Rust changes (the loader skips the version sentinel for workspace loads, so a stale.noderuns silently).- Source mode is single-arch. The deps tree matches the docker daemon's
native arch; trials on emulated images (e.g. x64 tasks on an arm64 host)
fail setup with an arch-mismatch error — use
--binaryfor those. - The repo is visible (read-only) inside task containers in source mode; fine for curated benchmarks, but don't point it at untrusted tasks.
--install localreflects local TS changes (inlined intodist/cli.js), but not uncommitted Rust natives — rebuildpackages/nativesper target first (the version sentinel must match).