Three coordinated tweaks in runEphemeralTurn and the supporting
#buildEphemeralSnapshot so IRC reply text stops leaking tool-call
markup, duplicating verbatim, and breaking DeepSeek-class encoders:
- Drop the recipient's tools array entirely instead of relying on
toolChoice:"none" (not every backend enforces it). The model now has
no tool surface to emit so leaked function_call / DSML markup stops.
- Preserve thinking content blocks when snapshotting the in-flight
streaming assistant message so the openai-completions encoder can
re-emit reasoning_content for DeepSeek-routed recipients (10 reports
of HTTP 400 "'reasoning_content' in thinking mode must be passed
back").
- Collapse consecutive duplicate sentences in replyText and cap reply
length so a looping recipient does not spam the IRC channel with the
same line repeated N times.
The 60s autoclear was mutating canonical #todoPhases via setTimeout, so
earlier completions vanished from the model's view of phase progress.
Default delay bumped well above any plausible turn duration and a
dedup helper added so the canonical list remains intact until the next
explicit prompt boundary.
The unconditional clear of #checkpointState on stopReason==="aborted"
fired on user interrupts, TTSR rule injection, streaming-edit guards,
plan-compact, and auto-compaction, silently dropping the user's
checkpoint with no signal to the model. Downstream #applyRewind already
tolerates message-count drift via its safeCount clamp, so the clear is
safe to remove. Accounts for 100% of rewind tool grievances.
Documents that async:true defers reporting but does not extend or
disable the timeout, so long-running daemons should pass a generous
timeout. Also notes the output minimizer may rewrite results and that
the full bytes are always available at the artifact:// footer.
- The plan-approval gate required extra.title but the wire schema only
declared an opaque additionalProperties record so codex/gpt-5.x could
not discover the field. Schema now declares title with a description
while still allowing passthrough for future per-context keys.
- resolve.md replaces the truncated "Schema depends on context:" line
with the actual enumeration.
- runResolveInvocation wraps apply() in try/catch; a thrown apply (e.g.
ast_edit overlap) requeues the resolve directive so the model can
discard or fix-and-retry instead of losing the preview.
- todo-write.md adds an explicit note that tasks are referenced by
verbatim content text; the tool never emits task-N IDs.
- resolveTaskOrError rejects ^task-\d+$ inputs with a clarifying error.
- execute sets isError:true when any op failed.
- appendItems short-circuits on the first "already exists" error so the
call no longer applies the prefix of a doomed batch.
Three independent lsp fixes that share the action dispatcher and import
graph in lsp/index.ts:
- symbol-required: project-aware references/rename/definition now reject
an omitted symbol with a structured error pointing at symbol=<name>
(and symbol#N for repeats). resolveSymbolColumn used to fall back to
the first non-whitespace column and the server happily answered for
whatever identifier sat there. findSymbolMatchIndexes also enforces
word boundaries for bare identifiers.
- timeout-distinguish: the outer catch around dispatched LSP actions
rethrew new ToolAbortError() for both an internal timeoutSignal abort
and a caller signal abort, so callers got an opaque "Operation
aborted". Timeouts now map to a ToolError that names the action and
the elapsed budget; caller cancels stay ToolAbortError.
- edit-coalesce: rename and rename_file collected edits from every
project-aware server and applied them sequentially, but each server
computed positions against pre-edit text. Once server A wrote, server
B's edits had stale offsets and produced malformed imports. Merge
per-uri edits across servers, sort by descending position, and throw
on overlap so the model retries instead of silently corrupting.
When the Codex backend returned the literal "(see attached image)" as
the answer for a text query and streamedAnswer was empty, the wrapper
accepted the placeholder as the response and the chain never advanced.
Throw SearchProviderError so the next provider is tried.
Worker emits BuildMessage errors via the async error event, after the
surrounding try/catch in spawnTabWorker has already resolved, so the
documented spawnInlineWorker fallback was unreachable for the very case
it was added to cover. initializeTabWorker now terminates the broken
worker and retries once via spawnInlineWorker, with the original error
attached as cause if the inline fallback also fails.
clampTimeout silently floored the caller-supplied timeout to 30s, so a
requested 120s for a slow waitForResponse came back as a 30s failure
indistinguishable from the default. Raise the cap and document the new
max in the schema field description.
networkidle2 requires <=2 in-flight requests for 500ms which never
resolves on dev servers (HMR, websockets, telemetry beacons), so
browser.open and tab.goto timed out before user code ran. "load" matches
Puppeteer's documented default and works on real-world pages.
Both empty-patch and whitespace-only-patch branches in the isolated-task
merge unconditionally appended "Applied patches: yes". The marker now
fires only when the combined patch is non-empty AND applyText succeeded.
No-op merges report "No changes to apply.". The renderer marker matcher
accepts the new string.
The "Started N background task jobs" announcement listed the per-task
label (e.g. 7-MyScout) without the underlying jobId, but JobTool.execute
looks up by jobId. Append (job: ${jobId}) to each task entry and accept
the task label as a fallback alias in JobTool.execute.
renderDescription filtered the listing only by disabledAgents. When a
subagent had parentSpawns empty the model still saw the full menu, tried
to spawn oracle, and got "Cannot spawn 'oracle'. Allowed: none". The
listing now intersects with the allowed-spawn set.
executePatchSingle returned success based on the writethrough callback
resolving, but the LSP-backed writethrough could resolve to an in-memory
editor buffer while disk stayed unchanged. Capture pre-write mtime+size,
re-stat post-write, and throw a ToolError when nothing changed.
Diff.applyPatch with fuzzFactor:3 absorbed orphan duplicate closers from
elsewhere in the file, editing the wrong site. Recovery now requires the
cached snapshot to hash-match the model-supplied anchors AND applies
with fuzzFactor:0; otherwise the original mismatch is rethrown so the
caller re-reads.
computeLineHash mixed in the line index for blank and punctuation-only
lines so any line shift invalidated anchors whose content was unchanged.
Two adjacent blank lines now collide on the same hash by design; the
range and op disambiguate by line number.
- search.md no longer claims "full regex syntax". Engine is rust-regex
(RE2) so lookaround and backreferences are unsupported; the doc now
says so and points at the post-filter alternative.
- search.ts rejects array entries containing a top-level comma with an
actionable ToolError, instead of silently demoting to a footer note
and returning zero matches.
- Add a paths-as-array example to search.md.
READY_TIMEOUT_MS was a hard 5s ceiling, ignoring the per-cell timeout the
caller supplied. Cold Bun starts on slow machines routinely exceed 5s
and the worker init failed before user code ran. The window now takes
the larger of the default and the caller timeout.
A top-level `return value;` in a JS eval cell was previously swallowed:
returnFinalExpression only handled ExpressionStatement, so ReturnStatement
flipped the IIFE wrapper which discarded the value. Rewrite the trailing
return into __omp_set_final_expr__((expr)) so the existing
final-expression channel surfaces the value just like a trailing
expression.
- Expose optional timeoutMs (clamped 0.5..60s) and pipe through to the
native walker. Default stays at 5s.
- On timeout, drain accumulated matches and return them with
truncated:true plus a notice line instead of throwing.
- Add gitignore boolean to the schema so callers can opt out of the
default exclude when looking for .env/.jsonl/build artifacts.
- Reject comma-in-paths array entries with an actionable ToolError.
- Update find.md with the knobs and the array-shape example.
ResolveBinary picked pip3 when only pip3 exists, but the spawn passed the
bare string "pip" which throws on macOS hosts without a pip alias. Use
the resolved path and wrap the helper in try/catch so a missing or
broken installer never escapes as an uncaught throw breaking URL reads.
The ignoreResultLimits flag now gates only the byte-budget tail-truncate,
never the explicit line window. Reads of internal URLs (skill, local,
memory) with a line range previously returned the tail of the file
instead of the requested window for files larger than the byte budget.
- Expanded `isAnthropicFastModeUnsupportedError` to treat 429 `rate_limit_error` responses mentioning fast mode as unsupported alongside 400 `invalid_request_error` speed-rejection cases.
- Added tests for unsupported-fast-mode detection covering 400, 429, and unrelated error payloads.
- Added `AgentSession.isFastModeActive()` with provider-scoped resolution and switched status-line rendering to use it for the fast-mode icon.
Two new `ServiceTier` values let users target priority/fast mode at one
provider family without paying premium costs on the other when switching
models mid-session:
- `"openai-only"` → resolves to `"priority"` on `openai` and
`openai-codex`; `undefined` everywhere else.
- `"claude-only"` → resolves to `"priority"` on direct `anthropic`;
`undefined` on Bedrock/Vertex Claude and elsewhere.
Implementation centers on a new `resolveServiceTier(serviceTier, provider)`
helper exported from `@oh-my-pi/pi-ai`. The three OpenAI providers and the
Anthropic provider all route through it, replacing the previous
`shouldSendServiceTier` type-guard pattern (which couldn't survive scoped
values — the input variable's literal type stops matching the wire type
once scopes are introduced). `shouldSendServiceTier` is kept as a plain
boolean for external callers but no longer narrows the input.
`getPriorityPremiumRequests` is reworked: it now counts Anthropic +
`"priority"` (fast mode) as one premium request — the original PR
introduced the realization but didn't update billing — and continues to
ignore providers that silently drop the field on the wire.
User-facing:
- `serviceTier` setting enum gains `"openai-only"` and `"claude-only"`
with clear UI descriptions.
- `/fast on` still sets the unscoped `"priority"`, but `/fast status`
and `isFastModeEnabled()` now report `on` for any priority-granting
tier (including scoped values). `/fast off` clears to `undefined`
regardless of scope.
- The Anthropic auto-fallback listener and re-arm clearing both cover
`"priority"` and `"claude-only"` (the two values that grant priority
on Anthropic). `"openai-only"` doesn't trigger the anthropic
fallback even if the user is on an Anthropic model — by design.
Tests cover all four resolver branches (unscoped passthrough, openai-only
match/miss, claude-only match/miss), Anthropic provider's wire `speed`
field under each scope, and updated premium accounting.
Replaces the parallel `speed` knob with the existing `serviceTier`
concept. The anthropic-messages provider now realizes
`serviceTier: "priority"` by setting `speed: "fast"` on the wire and
appending the `fast-mode-2026-02-01` beta header; other providers
continue to pass `service_tier` through natively or ignore it.
User-facing impact:
- `/fast` no longer dispatches on model.api. It just toggles
`serviceTier: "priority"`. Anthropic-specific translation lives
entirely in the provider.
- Anthropic auto-fallback marker is now the generic `"priority"`
identifier in `AssistantMessage.disabledFeatures` instead of
`"anthropic.fast_mode"`.
- New `clearAnthropicFastModeFallback(providerSessionState)` export is
invoked from `AgentSession.setServiceTier` when transitioning into
`"priority"`, so re-running `/fast on` after the provider
auto-disabled fast mode actually re-arms the next request instead of
silently no-oping.
Provider-side cleanups:
- Tightened cast site (`ParamsWithSpeed` alias) for the typed
`speed: "fast"` injection.
- Widened the rejection matcher (`\bspeed\b` + `not support`) so
phrasing drift ("is not supported" vs "does not support", quoted vs
backticked) doesn't break the fallback.
Dropped from the PR:
- `Agent.speed` / `AgentOptions.speed` / `SimpleStreamOptions.speed`
fields.
- `SpeedChangeEntry` and `appendSpeedChange` from the session entry
schema; service-tier change entries already cover this.
- `AgentSession.setSpeed` / `.speed` and the previousSpeed
capture/restore in `switchSession` — collapsed back into
`setServiceTier` + previousServiceTier, which now covers the rollback
too.
- Preserved launch and attach request failures when configurationDone also fails.
- Handled initial stop-outcome watcher rejections for failed launch and attach attempts.
- Rejected directory-valued debug launch programs before adapter selection and documented the debugpy launch shape.
Fixes#1187
- Added splitInternalUrlSel to iteratively peel internal-URL selector chunks while preserving unsupported schemes like mcp://.
- Updated ReadTool to use the internal splitter before routing so selector parsing is handled via parseSel.
- Added unit tests covering malformed selectors, namespaced skill hosts, and unchanged behavior for non-URLs or unsupported schemes.
The schema-level fix (normalizing {} to true in open schema positions)
makes the prompts correct-by-construction: models that follow the schema
grammar now emit string values for extra.title without needing softer
prompt language. Revert the SHOULD/optional hedging to the original MUST
wording per maintainer review.
Fixes#1179
Grammar-constrained models (e.g. Qwen3.6-35B-MTP via llama.cpp) emit
`extra: { title: {} }` instead of `extra: { title: "<string>" }` because
the resolve schema declares `extra` as Record<string, unknown> with an
open value schema, leaving the model free to drop in an empty object.
The apply guard then threw 'Plan approval requires extra: { title: ... }'
on every retry, looping the model indefinitely (issue #1179).
Plan approval now uses a layered title resolution:
1. `extra.title` if it is a non-empty string (and sanitizes to non-empty)
2. First `# Heading` in the plan content
3. Filename stem of `planFilePath` (`'/data/workspaces/can1357__oh-my-pi__1179/.omp-session/2026-05-19T03-59-21-254Z_019e3e63-62a6-7000-be63-371f2cd6d67d/local/PLAN.md'` → `PLAN`)
4. Literal `plan` as a final safety net
Each candidate is run through `normalizePlanTitle`; rejected ones fall
through. Extracted as `resolvePlanTitle` in plan-mode/approved-plan.ts
so it's unit-testable.
Prompt language relaxed from MUST to SHOULD for `extra.title` in
plan-mode-active.md and plan-mode-tool-decision-reminder.md, noting the
fallback so models don't waste turns on a now-optional field.
Fixes#1179
- Added a new `oracle.md` prompt defining a read-only diagnostic, architecture, and debugging advisor.
- Imported `oracle.md` into the agent prompt loader in `packages/coding-agent/src/task/agents.ts`.
- Included `oracle.md` in `EMBEDDED_AGENT_DEFS` so the new agent is available at runtime.
- Added capParseErrors in shared render utilities and updated ast_grep and ast_edit to return capped parseErrors plus parseErrorsTotal.
- Threaded the preserved totals into parse-error formatting and renderer output so labels and overflow counts report the full number of issues.
- Added an ast_grep test asserting parse errors are capped at PARSE_ERRORS_LIMIT while parseErrorsTotal retains the original count.
- Implemented strict base64 validation and normalization for image `displayValue` payloads in the JS runtime, supporting strict strings, `Uint8Array`, `Buffer`, `ArrayBuffer`, typed-array views, and JSON `Buffer` objects.
- Dropped unrecognized image payloads while emitting a warning and fallback text instead of forwarding malformed data.
- Added tests covering successful coercions and invalid image data rejection paths.
- Guard renderInlineMarkdown against non-string input: partial JSON during
streaming can leave option label fields as undefined, causing marked.lexer
to throw 'undefined is not an object (evaluating e.replace)'. The ask tool
renderer now silently falls back to an empty string or baseColor output.
- Sanitize normalizePlanTitle instead of hard-rejecting: models that produce
natural-language plan titles like 'My Improvement Plan' were getting a
ToolError on every resolve call, causing an infinite retry loop. Spaces are
now converted to hyphens, remaining invalid chars are dropped, and only
truly unresolvable titles (empty after sanitization, path separators) throw.
- Fix ask.md prompt example: the example showed the legacy single-question
format (question/options/recommended at the top level) while the schema
requires questions: [{id, question, options}]. Models that follow examples
closely (Qwen3) generated calls that always failed schema validation.
Fixes#1176
Rewrites JSON-roundtripped Zod 4 schema objects that leak Zod internals as
JSON Schema keywords (e.g., `type:"enum"`, `enum:{...}`) into valid
JSON Schema 2020-12. This prevents validation failures when such schemas
are used as tool input schemas (e.g., from MCP servers).
Updates `isZodSchema` to reject deserialized Zod impostors that retain
`_zod` but lose their prototype.
Wires `toJSON` methods onto TypeBox shim schemas to ensure `JSON.stringify`
produces clean JSON Schema, preventing future leaks.
Fixes#1101