- Added fetch.pruneTags=false to git wrapper to prevent tag deletion during concurrent maintenance.
- Implemented atomic tag creation and push with retry logic to defend against tags being pruned by background git maintenance processes.
- Changed push strategy to use --atomic flag and retry up to 3 times on refspec mismatch errors.
- Added embedded addon tarball output in `embed-native.ts` using `embedded-addons.<platform>.tar.gz` artifacts.
- Added metadata-rich addon types with `size`, `filePath`, and optional `archive` fields.
- Updated extraction to prioritize archive unpacking and skip cached `.node` files when sizes match.
- Added `extractEmbeddedAddonArchive()` with archive parsing and safe validation of archive entry names and kinds.
- Adjusted release profile to disable line-table generation and strip symbols in `Cargo.toml`.
- Added CI-only ELF validation for forbidden sections and regression coverage for issue-823 archive extraction.
- Rejected bare `A` anchors; single-line ranges must now be spelled `A A`.
- Added a descriptive error for single-number headers to guide model output.
- Updated grammar, tokenizer, prompt docs, and tests to reflect the change.
- Added a `vault.enabled` setting and `isVaultEnabled` guard, and vault resolve, write, and path resolution now threw a disabled error when the feature was off.
- Improved CLI handling by parsing active vault path output and treating `Error:` lines from stdout/stderr as command failures.
- Updated tests to validate the disabled gate, cached active-vault path resolution, and CLI error surfacing on successful exit codes.
- Handled "incomplete" stop reasons in session recovery and auto-compaction workflows.
- Dropped the prior assistant turn before attempting recovery on incomplete-length stops.
- Expanded auto-compaction reason types and triggers to include "incomplete".
- Updated internal URLs parsing internals, export order, tests, and Obsidian URI prompt docs.
- Added vault:// URL parsing, typed variants, and path resolution with vault-root validation.
- Added VaultProtocolHandler with fs and Obsidian CLI-backed resolve/read/write/list support plus caching.
- Added vault scheme integration in router, path utils, and plan-mode guard using resolveVaultUrlToPath.
- Documented vault:// read/edit and `?op`-scoped URI formats in system prompts when Obsidian is available.
- Secured vault:// operations by rejecting traversal, absolute, and symlink-escape path cases.
- Fixed response.incomplete recovery by dropping truncated turns and promoting context.
- Added internal tests for vault protocol parsing, caching, CLI behavior, and invalid-path defenses.
- Replaced the `keepaliveWhile` Promise wrapper with a new `EventLoopKeepalive` class that registers and disposes an interval timer through `Symbol.dispose`.
- Updated `Agent` to instantiate `EventLoopKeepalive` via `using` during prompt execution instead of manually managing an interval.
- Wrapped interactive mode's await path with the new helper and removed redundant `keepaliveWhile` usage from the CLI entrypoint.
The previous commit accidentally dropped the try/catch in #emit
because the local agent.ts was based on an older main that lacked
the listener-isolation code. Restore it to match main.
- Fix typo "stauts" → "status" in stream.test.ts comment
- Update frontmatter.ts link to packages/utils/src/ (moved from coding-agent)
- Update notebook.ts link to src/edit/ (moved from src/tools/)
- Mark removed bash-normalize.ts in docs and update description
- Add CHANGELOG.md for swarm-extension and utils packages
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Addresses codex review on PR #1468 (P2): the bundled catalog was
under-reporting Serverless costs and (in the original commit, since
fixed) zeroing billable models.
Two issues compounded:
1. The `*_cents_per_million` field on `/v1/models` is *not* literal
cents. Cross-referencing every published model on wafer.ai's
Serverless rate card against the API envelope yields an exact 1.25×
ratio (120 → $1.50, 48 → $0.60, 88 → $1.10, 15 → $0.19, …). The
mapper now converts via `value × 125 / 10000` (multiply-first so the
result is a finite dyadic for every observed value — `12 × 0.0125`
produced `0.15000000000000002`, the integer-first form yields exactly
`0.15`).
2. The Pass SKU is a flat-rate subscription with no per-token charge.
Convention in this repo is to seed plan providers at `cost: 0`
(see `kimi-code`, `firepass`, `alibaba-coding-plan`). The mapper now
short-circuits Pass with all-zero cost regardless of envelope.
Updated all 7 wafer-serverless bundled entries to retail rates and
zeroed the 2 wafer-pass entries.
Tests:
- Existing assertion on Kimi-K2.6's cost updated to the retail rate
(1.1 / 4.8 / 0.1125).
- New mapper-level test fans the same upstream record through both
`waferPassModelManagerOptions` and `waferServerlessModelManagerOptions`
and asserts Pass → `{ 0, 0, 0, 0 }`, Serverless → retail conversion
(120/360/12 → 1.5/4.5/0.15). Locks the convention against future
regressions.
7 pass, 0 fail, 66 expect() calls. `bun check` clean.
The prior `mapWaferModel` pinned `thinkingFormat: "zai"` on every reasoning
model. That's the native shape for Z.AI/GLM and Moonshot Kimi, but the
wrong wire for DeepSeek and Alibaba Qwen — and Wafer passes the body
through to the partner backend rather than normalizing.
- DeepSeek V4 reasoning uses `reasoning_effort` ("openai" format) and
requires `reasoning_content` replay on assistant tool-call turns.
Sending `thinking: { type: "enabled" }` is either ignored or 400'd
depending on the backend; in either case it skips engaging thinking.
- Alibaba Qwen models use top-level `enable_thinking: boolean`
("qwen" format), not zai's binary thinking object.
The mapper now reads `wafer.provider` from the `/v1/models` envelope and
maps:
- `zai` / `moonshotai` → `thinkingFormat: "zai"`
- `qwen` → `thinkingFormat: "qwen"`
- `deepseek` and unknown → omit; `detectOpenAICompat` picks the right
default from the id at request time (deepseek-* → "openai" effort
plus reasoning_content replay + tool-choice/reasoning guards).
Bundled catalog updated to match:
- `GLM-5.1` (Pass + Serverless) and `Kimi-K2.6` keep `thinkingFormat: "zai"`
explicitly — auto-detect would mis-pick "openai" for these because
the Wafer baseUrl doesn't match the api.z.ai / api.moonshot.ai URL
patterns in `detectOpenAICompat`.
- `qwen3.7-max` drops the explicit override; `isQwen` autodetect
catches the lowercase id and picks "qwen".
- `deepseek-v4-flash` / `deepseek-v4-pro` drop the explicit override;
`isDeepseekFamily` autodetect activates `reasoning_effort` mapping
(minimal..high → "high", xhigh → "max") plus the deepseek-specific
invariants (`requiresReasoningContentForToolCalls`,
`disableReasoningOnToolChoice`).
Tests:
- `wafer.test.ts` extended with an explicit `thinkingFormat` assertion
on every reasoning bundled entry (locks the regression).
- New mapper-level test `Wafer dynamic discovery mapper` feeds a
synthetic `/v1/models` response with one entry per upstream
(zai, moonshotai, qwen, deepseek, future-provider, plus a
non-reasoning entry) and asserts the per-upstream `thinkingFormat`
branch — so a future regen against the live `/v1/models` can't
silently regress this.
6 pass, 0 fail, 62 expect() calls.
The earlier bundled catalog was missing three Serverless-only models and
had two entries whose metadata didn't match what `/v1/models` returns.
Cross-referenced against the live https://pass.wafer.ai/v1/models response
(public, unauthenticated) and rebuilt the wafer-serverless block from it.
Added:
- `qwen3.7-max` — 256k ctx, reasoning, $5/$15/$0.50 per M.
Canonical id is lowercase; round-trips verbatim on the wire.
- `deepseek-v4-flash` — 1M ctx, reasoning, $0.14/$0.28/$0.01 per M.
- `deepseek-v4-pro` — 1M ctx, reasoning, $1.74/$3.48/$0.02 per M.
Fixed:
- `Kimi-K2.6` cost was 0; live pricing is 88/384/9 cents per M
→ $0.88/$3.84/$0.09.
- `Qwen3.6-35B-A3B` was bundled with 32k context and text-only;
live `/v1/models` reports 256k and `vision: true`.
All three new reasoning models carry the same zai-style thinking compat
as the existing GLM/Kimi entries (`thinkingFormat: "zai"`,
`reasoningContentField: "reasoning_content"`) — Wafer normalizes reasoning
output to `reasoning_content` regardless of upstream provider.
Tests extended to cover the new entries plus the corrected Kimi pricing
and Qwen3.6 context/vision (5 cases, 47 expect() calls, all passing).
EventLoopKeepalive is not exported from yield.ts on main.
Use setInterval + unref() directly, matching the pattern in
keepaliveWhile(). Biome import order also fixed (type imports
before value imports within the same group).
Biome organizeImports sorts type imports before value imports within
the same relative-path group. Move EventLoopKeepalive import after
all type imports.
Wafer (https://wafer.ai) exposes a single OpenAI-compatible endpoint
(`https://pass.wafer.ai/v1`) for two SKUs whose entitlement differs
server-side, so we model them as two parallel providers — mirroring the
firepass/fireworks split so a user with both subscriptions can switch
without re-pasting:
- `wafer-pass` — flat-rate. `/v1/models` is filtered to entries whose
`wafer.tier === "pass_included"`.
- `wafer-serverless` — pay-as-you-go superset of Pass.
Both issue `wfr_…` keys. `/login wafer-pass` and `/login wafer-serverless`
paste-and-validate via `/v1/models`. `WAFER_PASS_API_KEY` and
`WAFER_SERVERLESS_API_KEY` are wired through `getEnvApiKey`.
Bundled catalog:
- `wafer-pass`: GLM-5.1, Qwen3.5-397B-A17B.
- `wafer-serverless`: GLM-5.1, Qwen3.5-397B-A17B, Kimi-K2.6, Qwen3.6-35B-A3B.
Dynamic discovery via `/v1/models` overlays additional models at runtime
and folds the `wafer` envelope (tier, capabilities, cents/M pricing) into
the canonical `Model<"openai-completions">` shape. GLM-family entries
carry the zai-style thinking compat (`thinkingFormat: "zai"`,
`reasoningContentField: "reasoning_content"`) so reasoning tokens land in
the right field. Cents-per-million → dollars-per-million via /100.
Tests (`packages/ai/test/wafer.test.ts`, 5 cases): bundled catalog
contract for both providers and wire-id pass-through (case-sensitive,
no rewrite — `GLM-5.1` must round-trip verbatim or upstream 404s).
Optional `packages/ai/test/wafer.live.ts` exercises a real round-trip
against `pass.wafer.ai` when `WAFER_PASS_API_KEY` is set.
- Namespaced `AgentSession.executePython()` session IDs before invoking the Python executor.
- Added a regression test proving eval state is visible to the user shortcut path.
Bun 1.3.x event loop busy-waits when the only pending work is an
unresolved Promise. Agent.prompt() sets #runningPromise via
Promise.withResolvers() which stays unresolved during the entire
agent loop execution (LLM calls + tool iterations), causing ~100%
CPU even when the process is idle.
PR #1419 added keepaliveWhile() to getUserInput() in main.ts, but
session.prompt() callers (interactive mode, resume, etc.) still
await the unresolved #runningPromise, bypassing the keepalive.
Install EventLoopKeepalive directly in Agent.prompt() so all
callers are covered. Dispose in the finally block after the agent
loop completes.
- Redesigned hashline patch syntax from anchor-based (`A-B:`) to hunk-header format (`@@ A..B @@`) with unified-diff compatibility.
- Removed `autoDropPureInsertDuplicates` option and simplified apply behavior to preserve duplicated boundary and context lines.
- Changed repeat operator from `^A-B` to `&A..B` and range separator from `-` to `..` for consistency with hunk-header syntax.
- Added image resizing and dimension notes to eval tool output; improved write tool hashline header sanitation for legacy formats.
- Removed 521 lines of boundary-duplicate absorption code and simplified parser to auto-convert bare body rows and unified-diff contamination.
- Classified usage-limit gateway responses as `429 rate_limit_error` in auth handling paths.
- Aligned auth-gateway and pi-native key retrieval with derived `sessionId` for `getApiKey` lookups.
- Handled usage-limit auth failures by rotating credentials with retry hints and returning undefined when none available.
- Replaced stream auth checks with retryable-upstream logic for 401 and usage-limit errors before content.
- Expanded `extractRetryHint` parsing for `~`, `sec`, `ms`, and minute/hour units.
- Added coverage for classifyGatewayError, retry-hint parsing variants, and stream-auth retry edge cases.
- Updated Anthropic prompt preparation to downscale image blocks to 2000px when a request has more than 20 images.
- Added regression tests in `anthropic-many-image-resize.test.ts` for resizing above threshold and no-op below threshold.
- Added optional `name` support to tool message schema, coercing blank values to undefined.
- Tracked assistant tool-call IDs to resolve `tool` message names from prior calls when wire names are missing.
- Updated auth-gateway classification to prefer explicit/embedded status codes over text heuristics.
- Added word-boundary matching, expanded parse/error tests, and updated changelog entries.
When the operator disables MCP client-side timeouts via timeout: 0 or OMP_MCP_TIMEOUT_MS=0, do not impose a 1s startup deadline on the optional HTTP GET SSE listener — let the listener wait as long as the server takes so server-to-client messages are not lost.
Refs #1460
Abort the optional Streamable HTTP GET SSE listener attempt after a short bounded startup window so POST-only request/response servers can finish initialization.
Fixes#1460
- Replaced 4-hex content-derived file hashes with 3-hex opaque tags minted by InMemorySnapshotStore, making tags session-bound pointers rather than content fingerprints.
- Removed lru-cache dependency; replaced LRU-bounded per-path rings with a flat 4096-slot global ring using a scrambled permutation to prevent LLM tag extrapolation.
- Made SnapshotStore required in Patcher (was optional); tag resolution now drives stale-anchor detection instead of recomputing hashes at apply time.
- Changed literal payload sigil from `|` to `+` and accepted `^A` shorthand for `^A-A`; added lenient recovery for bare bodies, lone `-` rows, and overlapping bare/concrete block pairs.
- Updated the auth-broker CLI to import the transport setter from the logger module.
- Replaced the logger.setTransports call in runServe with the dedicated setTransports helper.
- Added a `madeExecutable` result field to `WriteToolDetails` to surface executable changes.
- Implemented `maybeMarkExecutableForShebang` to chmod shebang files executable while preserving existing mode bits and swallowing chmod errors.
- Updated write flow and renderer output to return and display when a file was auto-marked executable.
- Replaced anchor shorthand syntax with explicit range format (1: -> 1-1:) and removed ^/v sigils in favor of ^A-B repeat and A-B:- delete operations.
- Added repeat edit kind to support ^A-B syntax for copying lines A through B, and inline delete syntax A-B:- for range deletions.
- Removed after_anchor cursor kind and standalone delete rows; empty anchor blocks now produce blank-line replacements instead of deletions.
- Updated parser, tokenizer, and type system to discriminate literal and repeat payloads, and refactored apply/recovery logic to expand repeat edits into individual inserts.
- Updated coding-agent test fixtures and settings documentation to reflect new hashline syntax and behavior.
Rewrote Vertex Claude rawPredict request bodies to drop the Anthropic "model" field and inject "anthropic_version": "vertex-2023-10-16", matching Vertex's required JSON shape, and asserted both invariants in the regression test.
Fixes#1456
Mapped Google Vertex Claude catalog entries to the Anthropic messages transport, rewrote placeholder rawPredict URLs with ADC auth, and covered the request URL and payload shape.
Fixes#1456
Replaced Google Vertex project discovery with the models.dev catalog so bundled model selection includes current Vertex MaaS and Gemini entries while pruning retired fallbacks.
Fixes#1456
- Narrowed the xai-oauth bundled model cast to `Model<"openai-responses">` in its regression test.
- Changed hashline stale-recovery fixtures to use `repl(...)` for both line-replacement payloads instead of `extra(pl(...))`.
A stalled Jina reader request shared the overall reader-mode AbortSignal
with the downstream trafilatura/lynx/native fallbacks. When Jina hung
until the budget timer fired, the shared signal aborted and the catch
handler's signal?.throwIfAborted() re-threw before any local fallback
ran.
- Bound Jina and Parallel extract to their own per-attempt sub-budget
(REMOTE_READER_MAX_MS, capped at 10s) so a remote stall cannot consume
the whole overall reader-mode budget.
- Catch handlers now rethrow only on real userSignal cancellation, not
on remote sub-budget or overall budget expiry.
- Wrap trafilatura/lynx in their own try/catch so a subprocess failure
or abort does not skip the in-process native renderer.
- Always attempt the native renderer last: it works on already-loaded
HTML with no network or subprocess, so even an exhausted overall
budget still yields a result.
Fixes#1449