- Added `openrouterVariant` option to `SimpleStreamOptions` and `OpenAICompletionsOptions` to append routing suffixes (`:nitro`, `:floor`, `:online`, `:exacto`) to OpenRouter model IDs at request time.
- Skips appending when the model ID already carries an explicit colon-suffix.
- Exposed `providers.openrouterVariant` setting in the coding-agent UI under Settings → Providers.
- Plumbed through `pi-native-server` forwarder and `AgentSession` options preparation.
Patch axis: extend
Displacement: net-zero; reuses existing plan reference state instead of adding persistence or overwriting approved artifacts
Rule violations averted: no approved-plan overwrite, no transcript format migration, no public CLI/API expansion
PASS/FAIL: PASS after plan-mode focused tests and package check. Note: system-prompt-templates has an unrelated HOME=/tmp path-shortening expectation failure.
Triple-stacked failure on the same axis (thinking effort) produced the
user-visible
Error: Compaction failed: Thinking effort high is not supported by
xai-oauth/grok-build.
Supported efforts:
(empty list after the colon) whenever the active model was a curated
xAI catalog entry with compat.supportsReasoningEffort: false.
Three defects lined up. (1) Behavior: compaction at four call sites
in packages/agent/src/compaction/compaction.ts hardcoded
reasoning: Effort.High and never threaded session.thinkingLevel —
the user's /model :off selection (and any explicit low/medium) was
silently overridden. On every other model this was invisible.
(2) Validation: requireSupportedEffort threw at the openai-flavored
mapper layer before the wire-side omitReasoningEffort gate in
providers/xai-responses.ts ever ran; two contradictory guards on the
same wire param. (3) Message: when getSupportedEfforts returned [],
the rendered error tail was 'Supported efforts: ' with nothing after
the colon — disappears as a side-effect of fix#2.
Fix#1 — thread ThinkingLevel | undefined end-to-end. Add
SummaryOptions.thinkingLevel and HandoffOptions.thinkingLevel.
Convert via a single exhaustive switch (effortFromThinkingLevel) in
the new resolveCompactionEffort helper:
- Off → undefined (omit reasoning entirely)
- undefined/Inherit → Effort.High → clamp per model (preserves the
historical default for users
who never touched the dial)
- explicit Effort → respect user → clamp per model
resolveCompactionEffort lives in compaction.ts; all four call sites
(generateSummary, generateHandoff, generateShortSummary,
generateTurnPrefixSummary) route through it. agent-session.ts threads
this.thinkingLevel into all three production compaction entry points
(manual /compact at L6201, auto-compaction at L6458 — the most-fired
path, originally missed in plan review — and direct generateHandoff
at L5465). The audit-gate test
(test/agent-session-compaction-thinking-threading.test.ts) scans the
file with a brace-balanced extractor and refuses any unthreaded site.
Fix#2 — silent-clamp at the openai-flavored mapper layer. Extract
exported modelOmitsReasoningEffort(model) in model-thinking.ts as the
single source of truth for compat.supportsReasoningEffort: false on
openai-responses* APIs. getSupportedEfforts now calls it instead of
inlining the check (pure refactor — observable behavior preserved).
resolveOpenAiReasoningEffort in stream.ts early-returns undefined
when the predicate is true, so the wire-side omitReasoningEffort
gate (providers/xai-responses.ts:78) becomes the single source of
truth for the actual strip — no redundant throw.
Three regression tests pin the contract:
- packages/ai/test/xai-oauth-effort-strip.test.ts (5 tests):
modelOmitsReasoningEffort returns true for grok-build and
grok-4.20-0309-reasoning, false for grok-4.3 / Anthropic /
openai-completions.
- packages/agent/test/compaction-thinking-level.test.ts (5 tests):
every ThinkingLevel outcome through generateHandoff — Off stays
undefined (not coerced to High), Low stays Low, Inherit / undefined
default to High, grok-build clamps to undefined regardless of
requested level. Covers the Codex-caught Off-vs-not-provided
distinction.
- packages/coding-agent/test/agent-session-compaction-thinking-threading.test.ts
(2 tests): brace-balanced source scan asserts every direct
compact() / generateHandoff() in agent-session.ts threads
'thinkingLevel: this.thinkingLevel'; floor of 3 threaded sites.
TDD red-green verified for fix#1: temporarily reverted the handoff
call-site back to hardcoded Effort.High → compaction-thinking-level
went 2 pass / 3 fail (Off coerced, Low overridden, grok-build throws);
restored → 5 pass / 0 fail.
Verified:
- packages/agent: 127 pass / 0 fail
- packages/ai: 1061 pass / 337 skip / 0 fail
- packages/coding-agent (focused): 179 pass / 5 skip / 0 fail
- biome + tsgo --noEmit clean across all three packages
Out of scope (follow-ups):
- branch-summarization.ts:307 already passes no reasoning — no edit.
- The empty-list error message at model-thinking.ts:296 is now
structurally unreachable from the openai-responses path.
- modelOmitsReasoningEffort and grokSupportsReasoningEffort
(xai-responses.ts:22) overlap; collapse into a single predicate
in a future commit.
Op: correct
Restores: spec:compaction-honors-session-thinking-level
Restores: spec:xai-oauth-grok-build-compaction-no-throw
(cherry picked from commit e07b47ee46769053c658819437e2478389a4cee0)
- Imported isPromise from node:util/types in four modules.
- Replaced four instanceof Promise checks with isPromise calls for more reliable promise detection.
Two TS errors in CI:
1. `agent-session.ts:1317` — `error TS1345: An expression of type 'void'
cannot be tested for truthiness`. The listener type is
`(event: AgentSessionEvent) => void`, so the returned value can't be
directly tested. Same shape in `agent.ts:1079`.
2. `test/session/emit-listener-isolation.test.ts:19` — the test fixture
for `AgentEvent.tool_execution_start` was missing the required `args`
field.
Cast the return to `unknown` and check `instanceof Promise` instead of
duck-typing `.then` — type-safe and matches what async functions actually
return. Add `args: {}` to the test fixture.
Both `AgentSession.#emit` (session/agent-session.ts) and `Agent.#emit`
(packages/agent/src/agent.ts) iterated listeners with no error isolation.
A synchronous throw in any subscriber aborted the for-loop, so later
subscribers (TUI rendering, ACP bridge, task executor progress,
hindsight) silently missed events. Many listeners — see
`modes/controllers/event-controller.ts:141` and
`modes/controllers/input-controller.ts:576` — are registered as
`async (event) => { await this.handleEvent(event); }`; the returned
Promise was dropped, so any rejection became an unhandled rejection.
Wrap each listener invocation in try/catch and attach a `.catch` to any
returned thenable. Errors are logged via `logger.warn` (already imported
in agent-session.ts) and `console.error` (agent.ts has no logger
dependency, keep it that way).
Test: new `test/session/emit-listener-isolation.test.ts` registers two
listeners on both classes; first listener throws (or returns a rejecting
Promise); asserts the second listener still receives the event AND no
`unhandledRejection` fires. 4 cases (sync+async × Agent+AgentSession).
All fail on current main; all pass with the fix.
- Short-circuited `agent_end` when `#checkCompaction` deferred handoff, skipping rewind/todo passes and `agent.continue()` race.
- Aborted retry/compaction paths in `AgentSession.dispose()` before draining post-prompt tasks so `/exit` and Ctrl+C no longer hang.
- Added handoff-deadlock regression tests in `agent-session-handoff.test.ts` to prevent reordering races.
- Replaced external watchdog timers with per-request SDK timeouts for first-event budget across OpenAI, Anthropic, and Azure providers.
- Keyed Python shared kernels by (sessionId, cwd) to prevent cross-directory state bleed.
- Deduplicated concurrent cold-start session acquisition for JS and Python executors.
- Moved `isOpenAIResponsesProgressEvent` to shared module and scoped display output routing per run for interleaved async cells.
- Removed deprecated MCP-specific type aliases and functions from tool-discovery module, consolidating to unified generic tool discovery API.
- Migrated session and SDK code to use generic filterBySource() and collectDiscoverableTools() instead of MCP-specific variants.
- Removed deprecated interface members including hasQueuedMessages(), FocusPane, AcpBuiltinCommandRuntime, and legacy settings methods.
- Updated test suites to use renamed generic discovery methods and removed back-compat test coverage for legacy MCP shapes.
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
- Expanded transient error matching for Bun HTTP2StreamReset, RefusedStream, and EnhanceYourCalm.
- Dropped thinking-only/error/aborted turns without text/toolCall, reset aborted tool-call map, and stored timestamps.
- Updated TUI render planning to track scrollback high-water and suppress suffix-scroll artifacts in non-multiplexer sessions.
- Added regression tests and changelog notes for Bun HTTP/2 retry handling, thinking-only filtering, and scrollback regressions.
- Added OpenAI Codex and Gemini web search provider options with updated setup/auth descriptions.
- Updated Codex OAuth flow to refresh near-expiry tokens during web_search and persist the refreshed credentials.
- Plumbed AgentStorage through search orchestrator, scrapers, and fetch paths so providers share session credentials.
- Refactored web provider and credential helpers to accept caller-provided AgentStorage and resolve keys synchronously.
Queued extension-delivered user messages when deliverAs is set and waited for session_start extension message sends before prompting subagents.
Fixes#1343
- Exported `normalizeTools` so `AppendOnlyContext` uses the same tool normalization as the agent loop.
- Added `BuildOptions.intentTracing` to `build()`/`reset()`/`takeSnapshot()` so intent injection is consistent and included in the prefix fingerprint.
- Improved `#computeDigest` to cover tool_calls, tool_call_id, name, and id fields to catch in-place mutations.
- Fixed `#unsubscribeAppendOnly` leak and added no-op guard in `#syncAppendOnlyContext`.
- Added `recoverOrphanedBackups` to promote `.jsonl..bak` files back to their primary path when the primary is missing, preventing data loss after a mid-rename crash.
- Changed backup filename from dot-prefixed to plain `..bak` so the shared `*.bak` glob can find it on both real and in-memory storage backends.
- Surfaced the original EPERM as the error `cause` and included both original and retry messages when rollback also fails.
- Added `retry.maxDelayMs` to the settings schema and interfaces, with a default cap for provider backoff delays.
- Updated session auto-retry logic to fail fast when a requested wait exceeds the cap without fallback, emitting terminal auto-retry failure state.
- Propagated retry state and failure data into task progress and rendering so children show retry/wait details and reminder prompts stop after terminal errors.
Replaced overwrite-style session rewrites with an EPERM fallback that moves the old session file aside before retrying and restores it if the retry fails.
Added regression coverage for active-session rewrite recovery so the session remains writable after the fallback.
Fixes#1337
When session.prompt() returns, idle-flush tasks for async-job result
deliveries are scheduled via #schedulePostPromptTask (1ms delay) and
added to #postPromptTasks immediately. The 800ms loop timer could fire
in that window before isStreaming became true, causing the loop to
submit the next prompt while the delivery turn was still pending. The
delivery then hit AgentBusyError and the job result was silently dropped.
Add AgentSession.hasPostPromptWork (= #postPromptTasks.size > 0) and
include it in #isLoopAutoSubmitBlocked() alongside isStreaming and
isCompacting. Add a regression test that verifies the loop defers when
hasPostPromptWork is true and fires once it becomes false.
Fixes#1294
- Added optional `onBeforeYield` configuration and `setOnBeforeYield` in Agent, executed before follow-up checks.
- Added `YieldQueue` to `AgentSession`, with setup/teardown and streaming/idle flush via `setOnBeforeYield`.
- Replaced immediate async-result follow-up dispatch with queued batch entries, including stale-state suppression.
- Added MCP follow-up queueing in SDK, deduplicating updates by `serverName` and `uri`.
- Added changelog entries for `onBeforeYield`, async-result batching, MCP dedupe, and `display.shimmer` modes.
- Added yield queue unit tests for streaming emission, debounced idle batches, stale filtering, and error isolation.
- Added the active session model to compaction candidate selection before role-based candidates.
- Updated compaction routing so role-based models are only considered after the current chat model.
- Added a regression test proving an Anthropic session prefers its active model over `modelRoles.default` on OpenAI.
#scheduleTodoAutoClear / #runTodoAutoClear used to splice completed and
abandoned tasks out of #todoPhases on a 60s (later 30min) timer. The
mutation made earlier completions vanish from phase counts ("5 tasks"
dropped to 4) and contradicted the model's own claim of progress.
The autoclear path is removed entirely. Canonical #todoPhases is only
mutated by explicit todo_write calls. formatSummary's denominator
(`current.tasks.length`) now stays stable across tool calls, so phase
counts include completed tasks until the model explicitly removes them.
Leaves the `tasks.todoClearDelay` setting in place (inert) to avoid
changing the schema in this patch.
Three coordinated tweaks in runEphemeralTurn and the supporting
#buildEphemeralSnapshot so IRC reply text stops leaking tool-call
markup, duplicating verbatim, and breaking DeepSeek-class encoders:
- Drop the recipient's tools array entirely instead of relying on
toolChoice:"none" (not every backend enforces it). The model now has
no tool surface to emit so leaked function_call / DSML markup stops.
- Preserve thinking content blocks when snapshotting the in-flight
streaming assistant message so the openai-completions encoder can
re-emit reasoning_content for DeepSeek-routed recipients (10 reports
of HTTP 400 "'reasoning_content' in thinking mode must be passed
back").
- Collapse consecutive duplicate sentences in replyText and cap reply
length so a looping recipient does not spam the IRC channel with the
same line repeated N times.
The 60s autoclear was mutating canonical #todoPhases via setTimeout, so
earlier completions vanished from the model's view of phase progress.
Default delay bumped well above any plausible turn duration and a
dedup helper added so the canonical list remains intact until the next
explicit prompt boundary.
The unconditional clear of #checkpointState on stopReason==="aborted"
fired on user interrupts, TTSR rule injection, streaming-edit guards,
plan-compact, and auto-compaction, silently dropping the user's
checkpoint with no signal to the model. Downstream #applyRewind already
tolerates message-count drift via its safeCount clamp, so the clear is
safe to remove. Accounts for 100% of rewind tool grievances.
- Expanded `isAnthropicFastModeUnsupportedError` to treat 429 `rate_limit_error` responses mentioning fast mode as unsupported alongside 400 `invalid_request_error` speed-rejection cases.
- Added tests for unsupported-fast-mode detection covering 400, 429, and unrelated error payloads.
- Added `AgentSession.isFastModeActive()` with provider-scoped resolution and switched status-line rendering to use it for the fast-mode icon.
Two new `ServiceTier` values let users target priority/fast mode at one
provider family without paying premium costs on the other when switching
models mid-session:
- `"openai-only"` → resolves to `"priority"` on `openai` and
`openai-codex`; `undefined` everywhere else.
- `"claude-only"` → resolves to `"priority"` on direct `anthropic`;
`undefined` on Bedrock/Vertex Claude and elsewhere.
Implementation centers on a new `resolveServiceTier(serviceTier, provider)`
helper exported from `@oh-my-pi/pi-ai`. The three OpenAI providers and the
Anthropic provider all route through it, replacing the previous
`shouldSendServiceTier` type-guard pattern (which couldn't survive scoped
values — the input variable's literal type stops matching the wire type
once scopes are introduced). `shouldSendServiceTier` is kept as a plain
boolean for external callers but no longer narrows the input.
`getPriorityPremiumRequests` is reworked: it now counts Anthropic +
`"priority"` (fast mode) as one premium request — the original PR
introduced the realization but didn't update billing — and continues to
ignore providers that silently drop the field on the wire.
User-facing:
- `serviceTier` setting enum gains `"openai-only"` and `"claude-only"`
with clear UI descriptions.
- `/fast on` still sets the unscoped `"priority"`, but `/fast status`
and `isFastModeEnabled()` now report `on` for any priority-granting
tier (including scoped values). `/fast off` clears to `undefined`
regardless of scope.
- The Anthropic auto-fallback listener and re-arm clearing both cover
`"priority"` and `"claude-only"` (the two values that grant priority
on Anthropic). `"openai-only"` doesn't trigger the anthropic
fallback even if the user is on an Anthropic model — by design.
Tests cover all four resolver branches (unscoped passthrough, openai-only
match/miss, claude-only match/miss), Anthropic provider's wire `speed`
field under each scope, and updated premium accounting.
Replaces the parallel `speed` knob with the existing `serviceTier`
concept. The anthropic-messages provider now realizes
`serviceTier: "priority"` by setting `speed: "fast"` on the wire and
appending the `fast-mode-2026-02-01` beta header; other providers
continue to pass `service_tier` through natively or ignore it.
User-facing impact:
- `/fast` no longer dispatches on model.api. It just toggles
`serviceTier: "priority"`. Anthropic-specific translation lives
entirely in the provider.
- Anthropic auto-fallback marker is now the generic `"priority"`
identifier in `AssistantMessage.disabledFeatures` instead of
`"anthropic.fast_mode"`.
- New `clearAnthropicFastModeFallback(providerSessionState)` export is
invoked from `AgentSession.setServiceTier` when transitioning into
`"priority"`, so re-running `/fast on` after the provider
auto-disabled fast mode actually re-arms the next request instead of
silently no-oping.
Provider-side cleanups:
- Tightened cast site (`ParamsWithSpeed` alias) for the typed
`speed: "fast"` injection.
- Widened the rejection matcher (`\bspeed\b` + `not support`) so
phrasing drift ("is not supported" vs "does not support", quoted vs
backticked) doesn't break the fallback.
Dropped from the PR:
- `Agent.speed` / `AgentOptions.speed` / `SimpleStreamOptions.speed`
fields.
- `SpeedChangeEntry` and `appendSpeedChange` from the session entry
schema; service-tier change entries already cover this.
- `AgentSession.setSpeed` / `.speed` and the previousSpeed
capture/restore in `switchSession` — collapsed back into
`setServiceTier` + previousServiceTier, which now covers the rollback
too.
Wires `speed: "fast"` and the `fast-mode-2026-02-01` beta into the
Anthropic provider, plumbs a matching `speed` option through
`SimpleStreamOptions` and the Agent, and teaches `/fast` to dispatch
on the active model's API (Anthropic -> speed=fast, OpenAI -> existing
serviceTier=priority path). Server is the authority on which models
support fast mode.
When the server rejects an unsupported model, the provider mirrors the
strict-tools fallback: drops the field, retries the same turn
transparently, persists the disable via `providerSessionState`, and
surfaces the action through the new `AssistantMessage.disabledFeatures`
marker so the session can sync the toggle off and warn the user.
- Tracked ACP tool-call inputs per session and replayed them via `toolArgsById`/`getToolArgs` plumbing.
- Merged ACP tool execution end content from start and result events so command output replay preserves original args.
- Scoped ACP async-job draining by session `ownerId` and `agentId` with in-flight tracking and permission-gated deferred turns.
- Refactored compaction telemetry and async tests with per-test telemetry setup and asynchronous teardown resets.
- Held wire-level agent_end until #promptInFlightCount drops to 0, preventing AgentBusyError when subscribers fire the next prompt synchronously from agent_end.
- Added #pendingAgentEndEmit field and #flushPendingAgentEnd(), called from #endInFlight and #resetInFlight.
- Added regression test covering re-entrant prompt() from agent_end listener.
- Migrated per-object caches (chat/tool starts, model fingerprints, validation contexts, provider indexes, render IDs) from WeakMap to Symbol-keyed properties on the objects themselves.
- Rewrote SSE debug tee as a single-pass inline parser, eliminating the body.tee() + readSseEvents re-parse pipeline.
- Refactored MockModel from a factory function + external WeakMap state into a self-contained class.
- Added FIFO memoization caches for heuristic candidate expansion and namespace suffix lookups.