Rolled back binary updater replacements when post-install version verification fails instead of deleting the previous working binary first.
Added a release workflow gate that downloads the published macOS arm64 asset and verifies codesign plus --version before npm publishing.
Fixes#1240
Route /force through a named Ollama tool choice and scope the Ollama request tools to that selected name so local models cannot pick a different tool.\n\nFixes #1236
ACP clients own MCP server configuration via session/new.mcpServers and AcpAgent#configureMcpServers. The ACP session factory previously left enableMCP at its default (true), so createAgentSession ran discoverAndLoadMCPTools on every session/new and the resulting host MCP tools landed in the session tool registry alongside the client-supplied ones. search_tool_bm25 then surfaced only the host tools.
Force enableMCP: false on every session created through createAcpSessionFactory so on-disk discovery is bypassed in ACP mode. Non-ACP modes (omp interactive, print, RPC) keep auto-discovery.
Fixes#1234
Prevent disabled providers from being registered for implicit local discovery and from creating built-in model discovery managers. Added regression coverage for disabled local providers during model registry refresh.
Fixes#1232
- Updated the oracle agent description to present it as a senior engineer who can either consult or execute when delegated.
- Removed the read-only mandate and added requirements to implement changes, run checks, and report completed work in delegation mode.
- Adjusted the procedure and decision guidance to distinguish consult versus implement paths while keeping the scope limited to requested tasks.
- Updated Perplexity JWT parsing in the ai OAuth utilities to return a far-future sentinel when `exp` is missing and use that when computing token expiry.
- Updated `getOAuthApiKey` to prefer the JWT expiry and normalized legacy one-hour `expires` values to a non-expiring value.
- Updated coding-agent web search token lookup to verify Perplexity credentials against the JWT `exp` claim and treat missing claims as non-expiring.
- Added isStreaming-aware diff options and trimmed-input handling for streaming/non-streaming previews.
- Added helpers to trim trailing partial lines and strip unmatched trailing `-`/`@@` blocks during streaming.
- Reworked apply_patch and hashline preview builders to keep files in input order with per-line added mapping.
- Added streaming preview regression tests and Unreleased Fixed changelog notes for partial-line and ordering fixes.
- Added reusable loop auto-submit deferral and readiness helpers for loop mode.
- Updated loop iteration flow to defer next prompts while the session is streaming or compacting.
- Added tests verifying loop submissions wait until compaction/streaming completes before resolving.
The size+mtime check from 094273df5 is unreliable on filesystems with
coarse mtime resolution: a same-length rewrite within the same tick
(e.g. "a" → "b") leaves both fields unchanged and falsely trips the
"file content did not change on disk" guard. CI on ubuntu Bun 1.3.14
hit this in the create-then-update aggregation test.
Re-read the file post-write and compare bytes to the previous content
instead — deterministic regardless of FS timestamp granularity.
The narrowed object-with-title schema added in e26a17f3f is no longer
necessary: the upstream constrained-sampling fix lets models emit
arbitrary props on a record now, so the opaque shape no longer hides
title from discovery. Plan-approval callers still pass extra.title and
the renderer/handler logic accept it unchanged.
The companion changes from e26a17f3f (resolve.md context enumeration,
runResolveInvocation apply-throw requeue) stay intact.
The timeout/async section is only meaningful when the async option is
exposed to the model. Wrap the new section in {{#if asyncEnabled}} to
match the existing async-bullet guard above.
Three coordinated tweaks in runEphemeralTurn and the supporting
#buildEphemeralSnapshot so IRC reply text stops leaking tool-call
markup, duplicating verbatim, and breaking DeepSeek-class encoders:
- Drop the recipient's tools array entirely instead of relying on
toolChoice:"none" (not every backend enforces it). The model now has
no tool surface to emit so leaked function_call / DSML markup stops.
- Preserve thinking content blocks when snapshotting the in-flight
streaming assistant message so the openai-completions encoder can
re-emit reasoning_content for DeepSeek-routed recipients (10 reports
of HTTP 400 "'reasoning_content' in thinking mode must be passed
back").
- Collapse consecutive duplicate sentences in replyText and cap reply
length so a looping recipient does not spam the IRC channel with the
same line repeated N times.
The 60s autoclear was mutating canonical #todoPhases via setTimeout, so
earlier completions vanished from the model's view of phase progress.
Default delay bumped well above any plausible turn duration and a
dedup helper added so the canonical list remains intact until the next
explicit prompt boundary.
The unconditional clear of #checkpointState on stopReason==="aborted"
fired on user interrupts, TTSR rule injection, streaming-edit guards,
plan-compact, and auto-compaction, silently dropping the user's
checkpoint with no signal to the model. Downstream #applyRewind already
tolerates message-count drift via its safeCount clamp, so the clear is
safe to remove. Accounts for 100% of rewind tool grievances.
Documents that async:true defers reporting but does not extend or
disable the timeout, so long-running daemons should pass a generous
timeout. Also notes the output minimizer may rewrite results and that
the full bytes are always available at the artifact:// footer.
- The plan-approval gate required extra.title but the wire schema only
declared an opaque additionalProperties record so codex/gpt-5.x could
not discover the field. Schema now declares title with a description
while still allowing passthrough for future per-context keys.
- resolve.md replaces the truncated "Schema depends on context:" line
with the actual enumeration.
- runResolveInvocation wraps apply() in try/catch; a thrown apply (e.g.
ast_edit overlap) requeues the resolve directive so the model can
discard or fix-and-retry instead of losing the preview.
- todo-write.md adds an explicit note that tasks are referenced by
verbatim content text; the tool never emits task-N IDs.
- resolveTaskOrError rejects ^task-\d+$ inputs with a clarifying error.
- execute sets isError:true when any op failed.
- appendItems short-circuits on the first "already exists" error so the
call no longer applies the prefix of a doomed batch.
Three independent lsp fixes that share the action dispatcher and import
graph in lsp/index.ts:
- symbol-required: project-aware references/rename/definition now reject
an omitted symbol with a structured error pointing at symbol=<name>
(and symbol#N for repeats). resolveSymbolColumn used to fall back to
the first non-whitespace column and the server happily answered for
whatever identifier sat there. findSymbolMatchIndexes also enforces
word boundaries for bare identifiers.
- timeout-distinguish: the outer catch around dispatched LSP actions
rethrew new ToolAbortError() for both an internal timeoutSignal abort
and a caller signal abort, so callers got an opaque "Operation
aborted". Timeouts now map to a ToolError that names the action and
the elapsed budget; caller cancels stay ToolAbortError.
- edit-coalesce: rename and rename_file collected edits from every
project-aware server and applied them sequentially, but each server
computed positions against pre-edit text. Once server A wrote, server
B's edits had stale offsets and produced malformed imports. Merge
per-uri edits across servers, sort by descending position, and throw
on overlap so the model retries instead of silently corrupting.
When the Codex backend returned the literal "(see attached image)" as
the answer for a text query and streamedAnswer was empty, the wrapper
accepted the placeholder as the response and the chain never advanced.
Throw SearchProviderError so the next provider is tried.
Worker emits BuildMessage errors via the async error event, after the
surrounding try/catch in spawnTabWorker has already resolved, so the
documented spawnInlineWorker fallback was unreachable for the very case
it was added to cover. initializeTabWorker now terminates the broken
worker and retries once via spawnInlineWorker, with the original error
attached as cause if the inline fallback also fails.
clampTimeout silently floored the caller-supplied timeout to 30s, so a
requested 120s for a slow waitForResponse came back as a 30s failure
indistinguishable from the default. Raise the cap and document the new
max in the schema field description.
networkidle2 requires <=2 in-flight requests for 500ms which never
resolves on dev servers (HMR, websockets, telemetry beacons), so
browser.open and tab.goto timed out before user code ran. "load" matches
Puppeteer's documented default and works on real-world pages.
Both empty-patch and whitespace-only-patch branches in the isolated-task
merge unconditionally appended "Applied patches: yes". The marker now
fires only when the combined patch is non-empty AND applyText succeeded.
No-op merges report "No changes to apply.". The renderer marker matcher
accepts the new string.
The "Started N background task jobs" announcement listed the per-task
label (e.g. 7-MyScout) without the underlying jobId, but JobTool.execute
looks up by jobId. Append (job: ${jobId}) to each task entry and accept
the task label as a fallback alias in JobTool.execute.
renderDescription filtered the listing only by disabledAgents. When a
subagent had parentSpawns empty the model still saw the full menu, tried
to spawn oracle, and got "Cannot spawn 'oracle'. Allowed: none". The
listing now intersects with the allowed-spawn set.
executePatchSingle returned success based on the writethrough callback
resolving, but the LSP-backed writethrough could resolve to an in-memory
editor buffer while disk stayed unchanged. Capture pre-write mtime+size,
re-stat post-write, and throw a ToolError when nothing changed.
Diff.applyPatch with fuzzFactor:3 absorbed orphan duplicate closers from
elsewhere in the file, editing the wrong site. Recovery now requires the
cached snapshot to hash-match the model-supplied anchors AND applies
with fuzzFactor:0; otherwise the original mismatch is rethrown so the
caller re-reads.
computeLineHash mixed in the line index for blank and punctuation-only
lines so any line shift invalidated anchors whose content was unchanged.
Two adjacent blank lines now collide on the same hash by design; the
range and op disambiguate by line number.
- search.md no longer claims "full regex syntax". Engine is rust-regex
(RE2) so lookaround and backreferences are unsupported; the doc now
says so and points at the post-filter alternative.
- search.ts rejects array entries containing a top-level comma with an
actionable ToolError, instead of silently demoting to a footer note
and returning zero matches.
- Add a paths-as-array example to search.md.
READY_TIMEOUT_MS was a hard 5s ceiling, ignoring the per-cell timeout the
caller supplied. Cold Bun starts on slow machines routinely exceed 5s
and the worker init failed before user code ran. The window now takes
the larger of the default and the caller timeout.
A top-level `return value;` in a JS eval cell was previously swallowed:
returnFinalExpression only handled ExpressionStatement, so ReturnStatement
flipped the IIFE wrapper which discarded the value. Rewrite the trailing
return into __omp_set_final_expr__((expr)) so the existing
final-expression channel surfaces the value just like a trailing
expression.
- Expose optional timeoutMs (clamped 0.5..60s) and pipe through to the
native walker. Default stays at 5s.
- On timeout, drain accumulated matches and return them with
truncated:true plus a notice line instead of throwing.
- Add gitignore boolean to the schema so callers can opt out of the
default exclude when looking for .env/.jsonl/build artifacts.
- Reject comma-in-paths array entries with an actionable ToolError.
- Update find.md with the knobs and the array-shape example.
ResolveBinary picked pip3 when only pip3 exists, but the spawn passed the
bare string "pip" which throws on macOS hosts without a pip alias. Use
the resolved path and wrap the helper in try/catch so a missing or
broken installer never escapes as an uncaught throw breaking URL reads.
The ignoreResultLimits flag now gates only the byte-budget tail-truncate,
never the explicit line window. Reads of internal URLs (skill, local,
memory) with a line range previously returned the tail of the file
instead of the requested window for files larger than the byte budget.
- Expanded `isAnthropicFastModeUnsupportedError` to treat 429 `rate_limit_error` responses mentioning fast mode as unsupported alongside 400 `invalid_request_error` speed-rejection cases.
- Added tests for unsupported-fast-mode detection covering 400, 429, and unrelated error payloads.
- Added `AgentSession.isFastModeActive()` with provider-scoped resolution and switched status-line rendering to use it for the fast-mode icon.
Two new `ServiceTier` values let users target priority/fast mode at one
provider family without paying premium costs on the other when switching
models mid-session:
- `"openai-only"` → resolves to `"priority"` on `openai` and
`openai-codex`; `undefined` everywhere else.
- `"claude-only"` → resolves to `"priority"` on direct `anthropic`;
`undefined` on Bedrock/Vertex Claude and elsewhere.
Implementation centers on a new `resolveServiceTier(serviceTier, provider)`
helper exported from `@oh-my-pi/pi-ai`. The three OpenAI providers and the
Anthropic provider all route through it, replacing the previous
`shouldSendServiceTier` type-guard pattern (which couldn't survive scoped
values — the input variable's literal type stops matching the wire type
once scopes are introduced). `shouldSendServiceTier` is kept as a plain
boolean for external callers but no longer narrows the input.
`getPriorityPremiumRequests` is reworked: it now counts Anthropic +
`"priority"` (fast mode) as one premium request — the original PR
introduced the realization but didn't update billing — and continues to
ignore providers that silently drop the field on the wire.
User-facing:
- `serviceTier` setting enum gains `"openai-only"` and `"claude-only"`
with clear UI descriptions.
- `/fast on` still sets the unscoped `"priority"`, but `/fast status`
and `isFastModeEnabled()` now report `on` for any priority-granting
tier (including scoped values). `/fast off` clears to `undefined`
regardless of scope.
- The Anthropic auto-fallback listener and re-arm clearing both cover
`"priority"` and `"claude-only"` (the two values that grant priority
on Anthropic). `"openai-only"` doesn't trigger the anthropic
fallback even if the user is on an Anthropic model — by design.
Tests cover all four resolver branches (unscoped passthrough, openai-only
match/miss, claude-only match/miss), Anthropic provider's wire `speed`
field under each scope, and updated premium accounting.
Replaces the parallel `speed` knob with the existing `serviceTier`
concept. The anthropic-messages provider now realizes
`serviceTier: "priority"` by setting `speed: "fast"` on the wire and
appending the `fast-mode-2026-02-01` beta header; other providers
continue to pass `service_tier` through natively or ignore it.
User-facing impact:
- `/fast` no longer dispatches on model.api. It just toggles
`serviceTier: "priority"`. Anthropic-specific translation lives
entirely in the provider.
- Anthropic auto-fallback marker is now the generic `"priority"`
identifier in `AssistantMessage.disabledFeatures` instead of
`"anthropic.fast_mode"`.
- New `clearAnthropicFastModeFallback(providerSessionState)` export is
invoked from `AgentSession.setServiceTier` when transitioning into
`"priority"`, so re-running `/fast on` after the provider
auto-disabled fast mode actually re-arms the next request instead of
silently no-oping.
Provider-side cleanups:
- Tightened cast site (`ParamsWithSpeed` alias) for the typed
`speed: "fast"` injection.
- Widened the rejection matcher (`\bspeed\b` + `not support`) so
phrasing drift ("is not supported" vs "does not support", quoted vs
backticked) doesn't break the fallback.
Dropped from the PR:
- `Agent.speed` / `AgentOptions.speed` / `SimpleStreamOptions.speed`
fields.
- `SpeedChangeEntry` and `appendSpeedChange` from the session entry
schema; service-tier change entries already cover this.
- `AgentSession.setSpeed` / `.speed` and the previousSpeed
capture/restore in `switchSession` — collapsed back into
`setServiceTier` + previousServiceTier, which now covers the rollback
too.