- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
OpenAI-compatible local servers can spend longer than the generic first-event budget processing large prompts before they emit response headers or SSE frames. The OpenAI-specific idle timeout now also acts as the OpenAI-family first-event floor unless an explicit OpenAI first-event timeout is configured.
Added regression coverage for OpenAI Responses request setup so a lower generic first-event watchdog no longer undercuts PI_OPENAI_STREAM_IDLE_TIMEOUT_MS.
Fixes#1603
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
- Fixed Anthropic stream idle-timeout errors incorrectly triggering provider retries after streaming had begun.
- Fixed darwin-x64 `bun build --compile` failure by guarding `onnxruntime-node` preload behind a `process.platform === "win32"` literal for dead-code elimination.
- Added `prefill` and `stop` parameters to the tiny-model worker's `complete` message type to pin output format without biasing content.
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
- Changed tiny-device preference resolution to always default to CPU instead of platform-specific DirectML/CUDA heuristics.
- Updated settings schema, documentation, and changelog text to describe the CPU default while keeping accelerated providers behind explicit `providers.tinyModelDevice`/`PI_TINY_DEVICE` choices.
- Added persistent settings in the Providers tab for ONNX execution provider and quantization/precision, replacing env-var-only configuration.
- PI_TINY_DEVICE and PI_TINY_DTYPE env vars still override the matching setting at spawn time.
- Added tinyWorkerEnvOverlay to map settings onto worker env without clobbering explicit env vars.
- Updated docs and tests to reflect the new setting-first resolution order.
- Added `PI_TINY_DEVICE` env var to control ONNX execution provider (`gpu` default, `cpu`, `metal`, `cuda`, `dml`, `coreml`).
- Local tiny-model inference now tries accelerated GPU provider first and retries on CPU if initialization fails.
- Added `device.ts` module with normalization, preference resolution, and load-order helpers.
- Updated docs and model descriptions to drop CPU-specific language.
- Added `replace block N:` syntax in grammar/parser/tokenizer/types, emitting `kind: "block"` edits with empty-body checks.
- Added block-resolution in hashline apply/recovery, wiring `BlockResolver` to resolve edits with throw/drop unresolved behavior.
- Added native `blockRangeAt` support and exported block range/options types through pi-ast, pi-natives, and JS bindings.
- Added coding-agent resolver wiring plus internal visibility/exports updates and new parser, patcher, and setup-wizard tests.
- Added repeatable `--tag` argument parsing in `gen-npm-packages.ts` and threaded parsed tags into `generateNpmPackages` for targeted leaf publishing.
- Exported `prepareNativeCorePackage` and expanded package manifest typing to support scripted manifest rewrites used by release workflows.
- Reworked install-test smoke logic to pack the host leaf package, pack the rewritten natives core, and assert the platform leaf package resolves from the core optional dependency.
- Added Mnemosyne runtime `extractionPrompt` and `consolidationPrompt` options and wired them into resolved LLM config.
- Added fact-extraction branch to call configured completion first (temp 0), then parse facts and safely fall back.
- Added tiny local model registry features for memory/title, including keys, specs, and validation helpers.
- Added `complete` protocol messages and abort-aware client/worker paths for local title completion generation.
- Added local-models documentation for tiny/memory transformer paths, defaults, and known parser caveats.
- Added `mnemosyne.scoping` setting: `global`, `per-project`, and `per-project-tagged`.
- `per-project-tagged` writes to a project-local bank while merging global memories on recall.
- Refactored `MnemosyneSessionState` to manage scoped recall/retain targets and deduplication.
- Updated hindsight tools to route recall/retain through scoped methods.
- Introduced `@oh-my-pi/pi-mnemosyne` workspace package with BeamMemory, recall/retain, FTS, and optional fastembed ONNX embeddings.
- Wired `memory.backend = "mnemosyne"` into the coding-agent settings schema and backend resolver.
- Added `MnemosyneSessionState` for per-session auto-recall on first turn and auto-retain after completed turns.
- Documented configuration, environment variables, and operational notes in `docs/mnemosyne-memory-backend.md`.
- Added `discovery.type: proxy` that hits `GET /v1/models` and routes each model via `supported_endpoint_types` (`anthropic` → `/v1/messages`, `openai` → `/v1/chat/completions`).
- Made provider-level `api` optional when `discovery.type` is `proxy`, since wire protocol is derived per-model.
- Increased discovery fetch timeout from 250ms to 10s to accommodate remote proxies.
- Documented proxy discovery configuration in `docs/models.md`.
- Replaced bare `A B` range headers with `replace N..M:`, `delete N..M`, `insert before N:`, `insert after N:`, `insert head:`, and `insert tail:`.
- Removed `&A..B` repeat rows; insert-before/after ops now express the same intent explicitly.
- Empty replace bodies now error instead of deleting; `delete` is the canonical deletion op.
- Updated grammar, prompt, docs, tests, and all call sites to the new syntax.
- Added a `hashRecognized` field to `MismatchDetails` and `MismatchError`, defaulting it to `true` for compatibility.
- Updated stale mismatch rejection messaging to distinguish drifted hashes from session-absent hashes with explicit guidance.
- Propagated `hashRecognized: snapshot !== null` in `Patcher` and added tests for both mismatch branches.
- Prepended `¶path#TAG` hashline header to plain file, ACP-bridge, and conflict resolution write results.
- Bulk conflict resolutions emit a trailing `Snapshots:` block with one header per written file.
- Suppressed when hashline display mode is disabled or for archive/SQLite/internal-URL targets.
- Added tests covering header presence, patcher usability, and disabled-mode suppression.
- Rejected bare `A` anchors; single-line ranges must now be spelled `A A`.
- Added a descriptive error for single-number headers to guide model output.
- Updated grammar, tokenizer, prompt docs, and tests to reflect the change.
- Fix typo "stauts" → "status" in stream.test.ts comment
- Update frontmatter.ts link to packages/utils/src/ (moved from coding-agent)
- Update notebook.ts link to src/edit/ (moved from src/tools/)
- Mark removed bash-normalize.ts in docs and update description
- Add CHANGELOG.md for swarm-extension and utils packages
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Wafer (https://wafer.ai) exposes a single OpenAI-compatible endpoint
(`https://pass.wafer.ai/v1`) for two SKUs whose entitlement differs
server-side, so we model them as two parallel providers — mirroring the
firepass/fireworks split so a user with both subscriptions can switch
without re-pasting:
- `wafer-pass` — flat-rate. `/v1/models` is filtered to entries whose
`wafer.tier === "pass_included"`.
- `wafer-serverless` — pay-as-you-go superset of Pass.
Both issue `wfr_…` keys. `/login wafer-pass` and `/login wafer-serverless`
paste-and-validate via `/v1/models`. `WAFER_PASS_API_KEY` and
`WAFER_SERVERLESS_API_KEY` are wired through `getEnvApiKey`.
Bundled catalog:
- `wafer-pass`: GLM-5.1, Qwen3.5-397B-A17B.
- `wafer-serverless`: GLM-5.1, Qwen3.5-397B-A17B, Kimi-K2.6, Qwen3.6-35B-A3B.
Dynamic discovery via `/v1/models` overlays additional models at runtime
and folds the `wafer` envelope (tier, capabilities, cents/M pricing) into
the canonical `Model<"openai-completions">` shape. GLM-family entries
carry the zai-style thinking compat (`thinkingFormat: "zai"`,
`reasoningContentField: "reasoning_content"`) so reasoning tokens land in
the right field. Cents-per-million → dollars-per-million via /100.
Tests (`packages/ai/test/wafer.test.ts`, 5 cases): bundled catalog
contract for both providers and wire-id pass-through (case-sensitive,
no rewrite — `GLM-5.1` must round-trip verbatim or upstream 404s).
Optional `packages/ai/test/wafer.live.ts` exercises a real round-trip
against `pass.wafer.ai` when `WAFER_PASS_API_KEY` is set.
- Redesigned hashline patch syntax from anchor-based (`A-B:`) to hunk-header format (`@@ A..B @@`) with unified-diff compatibility.
- Removed `autoDropPureInsertDuplicates` option and simplified apply behavior to preserve duplicated boundary and context lines.
- Changed repeat operator from `^A-B` to `&A..B` and range separator from `-` to `..` for consistency with hunk-header syntax.
- Added image resizing and dimension notes to eval tool output; improved write tool hashline header sanitation for legacy formats.
- Removed 521 lines of boundary-duplicate absorption code and simplified parser to auto-convert bare body rows and unified-diff contamination.
- Replaced 4-hex content-derived file hashes with 3-hex opaque tags minted by InMemorySnapshotStore, making tags session-bound pointers rather than content fingerprints.
- Removed lru-cache dependency; replaced LRU-bounded per-path rings with a flat 4096-slot global ring using a scrambled permutation to prevent LLM tag extrapolation.
- Made SnapshotStore required in Patcher (was optional); tag resolution now drives stale-anchor detection instead of recomputing hashes at apply time.
- Changed literal payload sigil from `|` to `+` and accepted `^A` shorthand for `^A-A`; added lenient recovery for bare bodies, lone `-` rows, and overlapping bare/concrete block pairs.
- Added a dedicated @oh-my-pi/hashline package with parser, patcher, filesystem, snapshots, and release metadata.
- Migrated coding-agent hashline and stream entrypoints to @oh-my-pi/hashline and removed old hashline module exports.
- Changed multi-section hashline execution to validate section hashes and flush diagnostics only at the final commit.
- Added session fileSnapshotStore support and rewired edit/read/search/write tools to use it instead of fileReadCache.
- Changed multiline payload syntax so continuation lines must start with `+`; that prefix is stripped before writing.
- Raw unprefixed lines after an op now throw an error instead of being silently accepted as payload.
- Removed `PAYLOAD_LINE_PREFIX_DEMOTED_WARNING` and the nested-replace demotion path; inner `N:` ops inside a pending `A-B:` now raise an overlap error.
- Raw blank lines between ops are ignored; use `+` alone for an empty payload line.
- Changed identical `A-B:` duplicate ops to last-wins coalesce with a warning, fixing spurious anchor-conflict errors when models emit before/after pairs.
- Non-identical overlap shapes (different ranges, replace+delete, delete+delete) still throw.
- Added new file hash line to edit tool result output after successful apply.
- In `resolveApproval`, yolo mode now returns the user policy directly (`allow`/`prompt`/`deny`) and ignores tool `override` prompts.
- Updated approval-mode and approval unit tests to match the new behavior for critical bash patterns under yolo and auto-approve.
- Updated docs and settings metadata to describe yolo as user-policy-driven rather than override-driven.
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
- Decouple the per-tool approval gate from extension presence. ExtensionRunner
and the ExtensionToolWrapper that hosts the gate are now constructed
unconditionally in createAgentSession. Previously the runner was only built
when extensionsResult.extensions.length > 0, so the entire approval system
silently disappeared for sessions with no extensions loaded — any
tools.approvalMode: prompt|custom setting was a no-op without feedback.
Today this hole was masked by createAutoresearchExtension always being
pushed inline; the unconditional construction makes the safety invariant
explicit, and a new regression test in approval-mode.test.ts pins it.
- Extend CRITICAL_BASH_PATTERNS to cover remote-fetch-then-execute shapes
that the original `bash <(curl …)` regex missed:
- `source <(curl …)` / `. <(curl …)` (anchored at command boundary so
`find . -name foo` doesn't false-positive)
- `eval "$(curl …)"` / `eval $(curl …)` / `eval `curl …``
Also adds `chmod -R` symbolic-mode forms (`u+x`, `u+rwx,o+w …`) targeting
filesystem root, and `tee` / `tee -a` writes to /etc/{passwd,shadow,sudoers}
(the standard way to write root-owned files without redirect). Benign
forms (`source ./local.sh`, `chmod -R u+x ./build`, `tee /var/log/app.log`,
`eval "$VAR"`) are pinned negative in the test suite.
- Extend formatApprovalPrompt with payload previews for the destructive tools
that previously rendered as bare `Allow tool: <name>`: eval (language +
first cell's code), task (agent + first task's id + assignment), ast_edit
(first op's pattern / replacement / paths), browser (action + tab + url +
code), and write content (alongside path). For `task` in particular this
closes the gap that docs/approval-mode.md's "parent's approval covers the
subagent" claim was waving at — the prompt now actually shows what's being
delegated.
- Tighten isMcpToolName: drop the fallback `|| toolName.includes("__")` so
an extension tool legally named `my__feature` or `pkg__util__do` is no
longer falsely labelled `Origin: MCP server tool` in the approval prompt.
Strict `mcp__` prefix only.
- Revert the cargo-cult `{ autoApprove: true } as AgentToolContext` insertions
in agent-session-python-cleanup.test.ts and sdk-move-cwd.test.ts. The tests
create sessions without passing settings, so the wrapper falls through to
approvalMode "auto" automatically; the explicit flag was unnecessary and
the `as AgentToolContext` cast hid that autoApprove lives on
CustomToolContext, not AgentToolContext.
- Document in commands/launch.ts the dual --auto-approve declaration (oclif
Flags for --help, manual parseArgs for runtime) so a future rename catches
both call sites.
- Promote the subagent caveat in docs/approval-mode.md to a callout near the
top: anything `task` is asked to do runs unattended once the parent task
call is approved.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts → 75 pass / 0 fail
(was 57; +18 cases covering new remote-exec patterns, chmod symbolic, tee
/etc, isMcp negative, and eval/task/ast_edit/browser/write payload previews)
- bun test packages/coding-agent/test/tools/approval-mode.test.ts → 7 pass /
0 fail (was 7; +1 case asserting extensionRunner is always constructed)
- bun tsc --noEmit -p packages/coding-agent → clean
- bun x biome check . → clean
- Windows EBUSY tempdir-cleanup noise in agent-session-python-cleanup and
sdk-move-cwd is pre-existing on this branch (already documented in the
PR body) and absent on Linux CI.
- approval: user 'tool: deny' now wins over critical-pattern override
(the override only tightens allow->prompt; it must never re-arm a denied tool).
- approval: rename hindsight policy keys to match registered tool names
(recall/retain/reflect, not hindsight_recall/hindsight_retain).
- approval: head+tail truncation for bash/ssh command prompts so a
destructive suffix buried after a long benign preamble stays visible.
- task/executor: force tools.approvalMode='auto' in createSubagentSettings
so subagents (which have no UI) cannot deadlock on per-tool prompts;
the parent's approval of the task call is the authorization.
- docs/approval-mode: rewrite so every example surfaces tools.approvalMode
and explains that tools.approval is ignored outside 'custom' mode.
Re-introduces the per-tool approval system from luzidd's commit 39124f3 (which
is no longer reachable from main) and improves it before re-landing.
What's restored:
- ApprovalPolicy (allow/deny/prompt) plus DEFAULT_APPROVAL_POLICIES.
- ACTION_EXCEPTIONS registry (LSP read-only, bash critical patterns).
- getApprovalPolicy() six-level resolution order.
- ExtensionToolWrapper.execute() gate before extension handlers.
- --auto-approve / --yolo CLI flag and tools.approval.<tool> user config.
- docs/approval-mode.md user guide.
What's improved over the original:
- Replaced unchecked 'as any' casts with typed unknown narrowing helpers.
- Validate userConfig values: invalid strings, numbers, etc. fall through to
the built-in default instead of being silently honoured (typo no longer
locks a tool out or grants implicit approval).
- Expanded CRITICAL_BASH_PATTERNS: chmod -R /, chown -R /, bash <(curl ...),
writes to /etc/passwd|shadow|sudoers, shutdown/reboot/halt/init 0,
kill -9 1, nc -e / nc -c reverse shells. Pattern shapes require a
command-position boundary so 'npm run reboot-tests' and 'echo "shutdown the
queue"' don't false-positive.
- Added DEBUG_READONLY_ACTIONS exception so DAP inspection actions (threads,
stack_trace, variables, scopes, read_memory, …) auto-allow while
execution-side actions (launch, attach, continue, evaluate, write_memory,
set_breakpoint, …) still prompt.
- formatApprovalPrompt: labels mcp__<server>__<tool> calls as MCP server
tools, surfaces ssh host + command, recognises the modern § hashline header
for edit, and truncates >240-char fields so a heredoc-sized body cannot
blow out the confirmation dialog.
- Test suite grown from 40 to 57 cases — new coverage for invalid user
config, the extended critical-bash patterns, benign-keyword negatives,
debug exceptions, MCP/ssh prompt formatting, and command truncation.
Verification:
- bun test packages/coding-agent/test/tools/approval.test.ts -> 57 pass
- bun x biome check . -> clean
- bun run check:ts across all 9 workspaces -> clean
- Rejected negative values in addition to non-numeric ones, falling back to per-server config or default 30s.
- Emitted a logger warning when an invalid env value is ignored.
- Added tests covering negative and non-numeric rejection cases.
- Migrated OAuth provider authentication from standalone `pi-ai` CLI to in-process `AuthStorage.login()` flow in coding-agent.
- Made provider argument optional for `login` and `logout` commands with interactive provider picker when omitted.
- Added `list` command to enumerate registered OAuth providers with optional `--json` output format.
- Removed `pi-ai` CLI binary and `bin` entry from @oh-my-pi/ai package; library API remains unchanged.
- Updated documentation and examples to reflect new `omp auth-broker` command interface and in-process OAuth flow.
- Updated type signature to show `paths` accepts `string | string[]` instead of only arrays.
- Clarified that single string paths are wrapped into a one-element list before resolution.
- Improved prompt instructions to explicitly show both string and array usage patterns.
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.