- Added the active session model to compaction candidate selection before role-based candidates.
- Updated compaction routing so role-based models are only considered after the current chat model.
- Added a regression test proving an Anthropic session prefers its active model over `modelRoles.default` on OpenAI.
#scheduleTodoAutoClear / #runTodoAutoClear used to splice completed and
abandoned tasks out of #todoPhases on a 60s (later 30min) timer. The
mutation made earlier completions vanish from phase counts ("5 tasks"
dropped to 4) and contradicted the model's own claim of progress.
The autoclear path is removed entirely. Canonical #todoPhases is only
mutated by explicit todo_write calls. formatSummary's denominator
(`current.tasks.length`) now stays stable across tool calls, so phase
counts include completed tasks until the model explicitly removes them.
Leaves the `tasks.todoClearDelay` setting in place (inert) to avoid
changing the schema in this patch.
Three coordinated tweaks in runEphemeralTurn and the supporting
#buildEphemeralSnapshot so IRC reply text stops leaking tool-call
markup, duplicating verbatim, and breaking DeepSeek-class encoders:
- Drop the recipient's tools array entirely instead of relying on
toolChoice:"none" (not every backend enforces it). The model now has
no tool surface to emit so leaked function_call / DSML markup stops.
- Preserve thinking content blocks when snapshotting the in-flight
streaming assistant message so the openai-completions encoder can
re-emit reasoning_content for DeepSeek-routed recipients (10 reports
of HTTP 400 "'reasoning_content' in thinking mode must be passed
back").
- Collapse consecutive duplicate sentences in replyText and cap reply
length so a looping recipient does not spam the IRC channel with the
same line repeated N times.
The 60s autoclear was mutating canonical #todoPhases via setTimeout, so
earlier completions vanished from the model's view of phase progress.
Default delay bumped well above any plausible turn duration and a
dedup helper added so the canonical list remains intact until the next
explicit prompt boundary.
The unconditional clear of #checkpointState on stopReason==="aborted"
fired on user interrupts, TTSR rule injection, streaming-edit guards,
plan-compact, and auto-compaction, silently dropping the user's
checkpoint with no signal to the model. Downstream #applyRewind already
tolerates message-count drift via its safeCount clamp, so the clear is
safe to remove. Accounts for 100% of rewind tool grievances.
- Expanded `isAnthropicFastModeUnsupportedError` to treat 429 `rate_limit_error` responses mentioning fast mode as unsupported alongside 400 `invalid_request_error` speed-rejection cases.
- Added tests for unsupported-fast-mode detection covering 400, 429, and unrelated error payloads.
- Added `AgentSession.isFastModeActive()` with provider-scoped resolution and switched status-line rendering to use it for the fast-mode icon.
Two new `ServiceTier` values let users target priority/fast mode at one
provider family without paying premium costs on the other when switching
models mid-session:
- `"openai-only"` → resolves to `"priority"` on `openai` and
`openai-codex`; `undefined` everywhere else.
- `"claude-only"` → resolves to `"priority"` on direct `anthropic`;
`undefined` on Bedrock/Vertex Claude and elsewhere.
Implementation centers on a new `resolveServiceTier(serviceTier, provider)`
helper exported from `@oh-my-pi/pi-ai`. The three OpenAI providers and the
Anthropic provider all route through it, replacing the previous
`shouldSendServiceTier` type-guard pattern (which couldn't survive scoped
values — the input variable's literal type stops matching the wire type
once scopes are introduced). `shouldSendServiceTier` is kept as a plain
boolean for external callers but no longer narrows the input.
`getPriorityPremiumRequests` is reworked: it now counts Anthropic +
`"priority"` (fast mode) as one premium request — the original PR
introduced the realization but didn't update billing — and continues to
ignore providers that silently drop the field on the wire.
User-facing:
- `serviceTier` setting enum gains `"openai-only"` and `"claude-only"`
with clear UI descriptions.
- `/fast on` still sets the unscoped `"priority"`, but `/fast status`
and `isFastModeEnabled()` now report `on` for any priority-granting
tier (including scoped values). `/fast off` clears to `undefined`
regardless of scope.
- The Anthropic auto-fallback listener and re-arm clearing both cover
`"priority"` and `"claude-only"` (the two values that grant priority
on Anthropic). `"openai-only"` doesn't trigger the anthropic
fallback even if the user is on an Anthropic model — by design.
Tests cover all four resolver branches (unscoped passthrough, openai-only
match/miss, claude-only match/miss), Anthropic provider's wire `speed`
field under each scope, and updated premium accounting.
Replaces the parallel `speed` knob with the existing `serviceTier`
concept. The anthropic-messages provider now realizes
`serviceTier: "priority"` by setting `speed: "fast"` on the wire and
appending the `fast-mode-2026-02-01` beta header; other providers
continue to pass `service_tier` through natively or ignore it.
User-facing impact:
- `/fast` no longer dispatches on model.api. It just toggles
`serviceTier: "priority"`. Anthropic-specific translation lives
entirely in the provider.
- Anthropic auto-fallback marker is now the generic `"priority"`
identifier in `AssistantMessage.disabledFeatures` instead of
`"anthropic.fast_mode"`.
- New `clearAnthropicFastModeFallback(providerSessionState)` export is
invoked from `AgentSession.setServiceTier` when transitioning into
`"priority"`, so re-running `/fast on` after the provider
auto-disabled fast mode actually re-arms the next request instead of
silently no-oping.
Provider-side cleanups:
- Tightened cast site (`ParamsWithSpeed` alias) for the typed
`speed: "fast"` injection.
- Widened the rejection matcher (`\bspeed\b` + `not support`) so
phrasing drift ("is not supported" vs "does not support", quoted vs
backticked) doesn't break the fallback.
Dropped from the PR:
- `Agent.speed` / `AgentOptions.speed` / `SimpleStreamOptions.speed`
fields.
- `SpeedChangeEntry` and `appendSpeedChange` from the session entry
schema; service-tier change entries already cover this.
- `AgentSession.setSpeed` / `.speed` and the previousSpeed
capture/restore in `switchSession` — collapsed back into
`setServiceTier` + previousServiceTier, which now covers the rollback
too.
Wires `speed: "fast"` and the `fast-mode-2026-02-01` beta into the
Anthropic provider, plumbs a matching `speed` option through
`SimpleStreamOptions` and the Agent, and teaches `/fast` to dispatch
on the active model's API (Anthropic -> speed=fast, OpenAI -> existing
serviceTier=priority path). Server is the authority on which models
support fast mode.
When the server rejects an unsupported model, the provider mirrors the
strict-tools fallback: drops the field, retries the same turn
transparently, persists the disable via `providerSessionState`, and
surfaces the action through the new `AssistantMessage.disabledFeatures`
marker so the session can sync the toggle off and warn the user.
- Tracked ACP tool-call inputs per session and replayed them via `toolArgsById`/`getToolArgs` plumbing.
- Merged ACP tool execution end content from start and result events so command output replay preserves original args.
- Scoped ACP async-job draining by session `ownerId` and `agentId` with in-flight tracking and permission-gated deferred turns.
- Refactored compaction telemetry and async tests with per-test telemetry setup and asynchronous teardown resets.
- Held wire-level agent_end until #promptInFlightCount drops to 0, preventing AgentBusyError when subscribers fire the next prompt synchronously from agent_end.
- Added #pendingAgentEndEmit field and #flushPendingAgentEnd(), called from #endInFlight and #resetInFlight.
- Added regression test covering re-entrant prompt() from agent_end listener.
- Migrated per-object caches (chat/tool starts, model fingerprints, validation contexts, provider indexes, render IDs) from WeakMap to Symbol-keyed properties on the objects themselves.
- Rewrote SSE debug tee as a single-pass inline parser, eliminating the body.tee() + readSseEvents re-parse pipeline.
- Refactored MockModel from a factory function + external WeakMap state into a self-contained class.
- Added FIFO memoization caches for heuristic candidate expansion and namespace suffix lookups.
- Added `AuthBrokerClient`, `RemoteAuthCredentialStore`, `AuthBrokerRefresher`, and `startAuthBroker` server in `packages/ai/src/auth-broker`.
- Renamed `AuthCredentialStore` class to `SqliteAuthCredentialStore`; extracted `AuthCredentialStore` as a persistence interface.
- Added `exportSnapshot`, `forceRefreshCredentialById`, `disableCredentialById`, and `upsertCredential` to `AuthStorage` for broker wire protocol.
- Added `omp auth-broker` CLI subcommand (serve, token, login, logout, import, status) and `discoverAuthStorage` broker-mode path keyed on `OMP_AUTH_BROKER_URL`.
- Non-interrupting tool-source TTSR matches now prepend a system-reminder to the matched tool's `toolResult` content instead of queuing a loop-wide deferred follow-up turn.
- Text/thinking source matches retain the previous deferred-injection behavior.
- Added deduplication so one rule attaches to exactly one sibling tool call per batch.
- Stale per-tool injections are cleared on abort/error before tools produce results.
- Replaced all StringEnum(...) usages with z.enum([...]) across tools, examples, and tests.
- Removed StringEnum re-export from @oh-my-pi/pi-coding-agent public API.
- Condensed verbose tool parameter descriptions to minimal lowercase phrases.
- Renamed AuthCredentialStore to SqliteAuthCredentialStore at usage sites.
- Added optional AgentTelemetry to summary, handoff, branch-summary, and compact option types.
- Replaced one-shot `completeSimple` usage with `instrumentedCompleteSimple` across compaction, summary, and branch-summary calls and passed `oneshotKind`.
- Added `PiGenAIAttr.OneshotKind`, `InstrumentedChatSpanOptions`, and response-header forwarding in telemetry span lifecycle.
- Added `resolveTelemetry` propagation in coding-agent session and inspect-image paths to pass request-scoped telemetry.
- Added compaction telemetry test harness and span assertions for success, no-telemetry, and error cases.
- Relocated compaction, branch-summarization, pruning, and utils from coding-agent to packages/agent/src/compaction.
- Moved OpenAI remote compaction helpers from packages/ai to the new compaction module.
- Added handoff.ts with extractHandoffDocument, createHandoffContext, and renderHandoffPrompt helpers.
- Exposed new entries.ts with standalone SessionEntry types so coding-agent no longer owns them.
- buildOpenAiNativeHistory now emits custom_tool_call / custom_tool_call_output for blocks with customWireName (apply_patch and other freeform tools), matching the normal Responses replay path; previously demoted to function_call which broke remote-compaction replay or mismatched the original call.
- requestOpenAiRemoteCompaction and requestRemoteCompaction accept an optional AbortSignal; the coding-agent compaction caller forwards the existing signal so cancellation now terminates the in-flight fetch instead of stranding the session in the compacting state until the server replies.
`SessionManager.close()` queues `#closePersistWriterInternal()` on the
persist chain. The task awaits `#persistWriter.close()`, which flips
`#closing = true` synchronously before yielding on its inner writer
`close()`. A concurrent `appendMessage()` landing in that yield window
hit the hot path, got the still-cached (but closing) writer back from
`#ensurePersistWriter()`, and threw `Error("Writer closed")` from
`writeSync`. The throw was stashed into `#persistError`; the next async
caller (`flush()` or a later `appendMessage()`) re-threw it as an
unhandled rejection with the original line-1282 stack.
Expose `NdjsonFileWriter.isOpen()` and treat a mid-close cached writer
as a miss in `#ensurePersistWriter()`. `_persist` now falls back to the
async `#rewriteFile()` cold path so the entry — already in
`#fileEntries` — still lands on disk once the close drains.
- Added OpenAI remote-compaction API support with provider-specific endpoint gating.
- Added buildOpenAiNativeHistory, token-budget estimation, and message trimming for remote-compaction.
- Added helpers to preserve and validate remote-compaction metadata in request/response handling.
- Refactored coding-agent compaction to consume pi-ai remote-compaction helpers with converted message history.
- Exported remote-compaction from ai index and documented the new APIs in CHANGELOG.
- Adjusted OutputSink to disable head retention after replace(), resetting counters so later pushes append to the tail and do not trigger stale middle-elision in dump().
- Refined artifact link emission to insert a newline separator only when the minimized output lacked one.
- Added a regression test for replace-plus-push ordering that verifies no elision marker and aligned byte counts.
- Added GoalRuntime with wall-clock and token accounting, budget steering, and lifecycle operations (create, pause, resume, drop, complete).
- Exposed goal tool as a hidden agent tool, activated only when goal mode is enabled.
- Integrated goal continuation loop in InteractiveMode with auto-submit between turns.
- Added status line segment and theme icons for goal mode state.
- Added sync truncation helpers to recursively prepare session entries and externalize image data.
- Reworked session persistence to use synchronous preparation plus `writeSync` with close-state checks.
- Added synchronous session-storage APIs and rerouted write paths to `writeLineSync`/`readTextSync`.
- Added `BlobStore.putSync`, migrated hashing to `Bun.SHA256`, and updated hash tests accordingly.
- Removed ExitPlanModeTool and deleted exit-plan-mode docs/tests, dropping the old approval contract outputs.
- Replaced plan-mode approval flow from exit_plan_mode to resolve across session, SDK, controllers, and discovery.
- Added standing resolve handler accessors and updated resolve routing for queued or standing approval handlers.
- Added PlanApprovalDetails and enforced normalized, validated approval titles with readable plan-file requirements.
- Extended resolve schema and invocation signatures with optional extra metadata and reason trimming behavior updates.
- Updated plan and resolve prompts and changelog guidance to require resolve action, reason, and extra.title for apply/discard.
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
Codex review flagged that the silent-abort sentinel
("__omp.silent_abort__") persists into AssistantMessage.errorMessage
but three downstream consumers render errorMessage verbatim:
- session-observer-overlay.ts: renders "✗ Error: __omp.silent_abort__"
when content is empty (confirmed user-visible today)
- print-mode.ts: writes marker to stderr and exits non-zero (latent;
plan-mode→compact not reachable from print mode today, but unguarded)
- acp-agent.ts: emits marker as agent_message_chunk text to ACP
clients when message has no other notifications (latent)
Add isSilentAbort() guard at each site. Extend the SILENT_ABORT_MARKER
consumer list in messages.ts doc comment to include all six consumers.
Add regression tests: overlay (2 tests), print-mode (2 tests), ACP
replay (1 test).
Op: correct
Restores: spec:silent-abort-marker-never-surfaces
ACP clients (Zed, etc.) only received `config_option_update` notifications
when they themselves drove the change via `session/set_session_config_option`.
Internal thinking-level updates (slash commands, automatic model-driven
adjustments, extension UI) bypassed the notification path, so client config
panels went stale until the next user-initiated change.
AgentSession now emits a `thinking_level_changed` event from
`setThinkingLevel`, and AcpAgent installs a session-lifetime subscription on
each managed session that pushes a fresh `config_option_update` whenever the
event fires — independent of prompt-turn lifecycle. The
`session/set_session_config_option` handler no longer pushes its own
notification for the `thinking` config (lifetime subscription covers it);
the response still returns fresh `configOptions` so callers see the new
state synchronously. Subscriptions are released in `#disposeSessionRecord`.
Also consolidated four duplicate `config_option_update` send sites into a
new `#pushConfigOptionUpdate(record)` helper.
Tests: added two cases to `test/acp-agent.test.ts` — one verifying internal
`setThinkingLevel` calls produce a `config_option_update` and a no-op
re-set produces none, and one verifying client-driven
`setSessionConfigOption(thinking, …)` produces exactly one notification.
Co-Authored-By: omp <noreply@oh-my-pi.dev>
- Removed local `abortableSleep` in favour of Node's built-in `scheduler.wait` from `node:timers/promises`.
- Consolidated per-provider retry/fetch loops into a shared `fetchWithRetry` utility in `packages/utils`.
- Moved `extractHttpStatusFromError`, `isRetryableError`, and related helpers out of `packages/ai` into `packages/utils`.
- Deleted `extractRetryDelay` in favour of `extractRetryHint` with unified header and body parsing.
- Added `tools.artifactHeadBytes` and `tools.outputMaxColumns` settings with defaults in `SETTINGS_SCHEMA`.
- Expanded `OutputSink` with `headBytes`/`maxColumns` and middle truncate logic with elision markers and tracking.
- Updated output-meta to resolve sink settings, emit truncation metrics, and use `truncateMiddle` for spills.
- Integrated head and column limits into JS/Python/Bash/SSH/read output flows, with `:raw` skipping read truncation.
- Documented new output middle-elision and column-cap behavior in `CHANGELOG.md`.
- Added truncation tests for `OutputSink`, `truncateMiddle`, and read-tool line handling.