- Expired the per-credential usage report cache after recording observed OpenCode Go spend.\n- Threaded provider base URL into OpenCode Go cost recording so the invalidation targets the same cache key /usage uses.\n- Added regression coverage for refreshing cached OpenCode Go limits immediately after a completed turn.\n\nFixes #2942
- Added an OpenCode Go usage provider that synthesizes 5h, weekly, and monthly cap windows from OMP-observed request costs.\n- Recorded OpenCode Go assistant request costs against the active credential so /usage can report local cap utilization.\n- Added regression coverage for fresh keys and observed spend aggregation.\n\nFixes #2942
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
Stopped auto context-full maintenance from retrying repeated summarization timeouts on the same model before fallback. Added a regression test for timeout fast-fallback behavior.\n\nFixes #2913
- Extracted native token counting into a new localized `tokenizer.ts` wrapping `@oh-my-pi/pi-natives`.
- Introduced a lightning-fast byte-length estimation logic for token counting when accurate counting is disabled.
- Diverted token calculations to the faster estimator during test environments and when `PI_TOKENIZER_ACCURATE` is falsy.
- Updated agent base and coding-agent sessions to consume the new localized `countTokens` utility.
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Bound native Reflect methods to local variables in the browser launch script to prevent detection.
- Updated stealth injection scripts to use the bound Reflect methods instead of global Reflect calls.
- Fixed a type assertion issue in the AgentSession tool proxy.
- Changed provider delta input comparisons in `buildResponsesDeltaInput` to use `Bun.deepEquals` instead of stringified JSON checks.
- Updated session message diffing to return early on length changes and compare normalized entries with deep equality.
- Replaced JSON-string assertions in the issue-966 repro test with `Bun.deepEquals` for stable equality checks.
- Tracked completed primary turns in `AgentSession` and started an immune-turn window after each interrupting advisor steer.
- Routed follow-on `concern`/`blocker` notes to the aside channel while the immune window is active, while preserving prior auto-resume-suppressed handling.
- Added the `advisor.immuneTurns` setting and tests for immune-turn detection and delivery-channel decisions.
- Guarded `session_stop` continuations against aborted or superseded hook results before queueing hidden follow-up turns.
- Reported compaction recovery continuations from `#checkCompaction` and skipped stop hooks while internal recovery owns the next turn.
- Let terminal empty-stop retry caps fall through to `session_stop` and added regression coverage for aborts, empty-stop caps, and promotion recovery.
- Updated getApiKey signatures to accept a Model and return ApiKey or ApiKeyResolver.
- Updated stream key handling to resolve credentials per model and use seedApiKeyResolver for retries.
- Added antigravityEndpointMode setting with auto/production/sandbox endpoint selection.
- Added 429/5xx endpoint failover for Gemini stream, usage, search, and image calls.
- Fixed context breakdown to anchor estimates on the latest completed assistant usage message after compaction.
- Adjusted pending-context usage selection to prefer an in-turn provider anchor when available at/after cutoff.
- Added a contextUsageRevision cache token so status-line context memo invalidates after snapshot clear.
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
- Added `images.describeForTextModels` configuration defaulting to true for text models.
- Added `describeAttachedImagesForTextModel` to persist images and generate local:// descriptions.
- Added image-description notices to session flow with hidden typing and pre-user insertion.
- Added fallback behavior that returns notes when vision is unavailable or output is empty.
Recorded todo reminder developer messages in the session log so JSONL transcripts match model-visible context when reminders are enabled.\n\nFixes #2824
- Sanitized artifact filenames by normalizing tool names before composing spill paths.
- Applied `wrapToolWithMetaNotice` to custom tool adapters and RPC-host tools in agent-session setup.
- Wrapped SDK-registered extension/custom tools with the same meta-notice adapter during session creation.
- Queued advisor concern cards now get reclaimed as visible advice during settle when auto-resume suppression is active and the session is idle.
- Preserve logic was narrowed to keep advisor cards hidden only during abort teardown, allowing steers during resumed streaming turns.
- A regression test was added to verify stranded advisor steers are persisted as visible advice without triggering an advisor-only resume turn.
- Added resolveAdvisorDeliveryChannel in advisor tooling to map each note to aside, steer, or preserve using severity, auto-resume suppression, core-streaming, and abort state.
- Updated AgentSession advice enqueuing to route concern/blocker notes through that resolver, preserving them only when the interrupted turn is idle or tearing down and steering them during active resumed turns.
- Added regression tests for resolveAdvisorDeliveryChannel covering nit versus interrupting severities across streaming, aborting, and suppression combinations.
- Rewrote `formatSessionDumpText` in `session-dump-format.ts` to emit the pre-16.x full dump: system-prompt prelude, model/thinking config, tool inventory with parameters, and the transcript as markdown role headings (`## User`, `## Assistant`, `### Tool Call`/`### Tool Result`), reusing `renderDelimitedThinking` for `<thinking>` blocks.
- Dropped the compact default and the `[raw]` flag from `/dump`: removed the `isRaw` parameter from `handleDumpCommand` in `command-controller.ts`, `interactive-mode.ts`, and `types.ts`, and removed the `inlineHint: "[raw]"`/`compact` plumbing in `builtin-registry.ts`.
- Updated the `formatSessionAsText` doc comment in `agent-session.ts` to describe the verbose dump shape.
- Removed the obsolete `formatSessionDumpText raw thinking` suite from `advisor.test.ts` and refreshed `session-dump-format.test.ts` to assert the verbose dump output.
- Recorded the revert in the coding-agent changelog and trimmed `/dump` from the compact transcript tool-intent-prefix entry.
- Generalized `isUserQueuedMessage` to a user-attribution predicate (`role === "user"` or custom `attribution === "user"` and not display-suppressed) so visible agent-authored steers (advisor cards, IRC/extension asides) and hidden goal/plan/budget steers are all excluded from editor restore, not just advisor cards.
- Gave `AgentSession.clearQueue` a `{ forInterrupt }` option: plain Alt+Up dequeue restores user messages and preserves every other queued message for the continuing stream, while Esc+abort keeps only advisor cards (for `abort()`'s `#extractQueuedAdvisorCards` preservation) and drops other internal steers so the post-abort `#drainStrandedQueuedMessages` can't auto-resume the interrupted run.
- Threaded `forInterrupt: options?.abort` from `InputController.restoreQueuedMessagesToEditor` and kept `queuedMessageCount` on actual displayable-queue semantics so `hasPendingMessages()`/RPC and the empty-submit abort gate stay accurate.
- Updated skill-queue tests to cover both policies (hidden and visible agent-authored steers preserved on dequeue, dropped on interrupt) and refreshed the changelog entry.
- Added `isUserQueuedMessage()` to `agent-session.ts`, treating only plain user turns and visible `attribution: "user"` custom messages (e.g. `/skill`) as restorable, so advisor concern/blocker notes, hidden goal/plan/budget steers, and IRC/extension asides no longer leak into the editor on Esc/Alt+Up.
- Reworked `clearQueue()` to return only user-authored messages while re-queuing advisor cards via `replaceQueues()` (so the user-interrupt abort path still re-records them as advice) and dropping other agent-authored steers to prevent a silent auto-resume on leftover internal context.
- Filtered `getQueuedMessages()` chips and rewrote `popLastQueuedMessage()` to skip agent-authored cards and pull the last user-authored entry, while `queuedMessageCount` still counts all displayable queued work.
- Extended `input-controller-skill-queue.test.ts` with a `queueAdvisorSteer` helper and cases asserting advisor/IRC cards count as pending work but stay out of chips, restore, and `popLastQueuedMessage`, and survive `clearQueue()`.
- Updated #isRetryableReasonlessAbort to reject reasonless aborts when the streaming-edit guard flag is set.
- This prevented routing those aborts through retry logic and avoided prompt hangs or unintended guard bypasses during edit-stream recovery.
Install a subagent's ordered model candidates as child-session retry fallback chains so a retryable provider failure advances to the next candidate instead of killing the worker (issue #2750).
Auto-retry empty/reasonless provider aborts without model fallback (issue #2685). Includes review fix b7a3d01439: skip the retry while the session is disposing to avoid a shutdown hang.
A dispose-driven bare abort() yields the same empty/reason-less aborted turn as a transient provider abort, but with #isDisposed set and #abortInProgress unset. #isRetryableReasonlessAbort matched it and routed it through #handleRetryableError, which created #retryPromise and scheduled a continuation the disposed guard then skipped without resolving the promise — hanging the in-flight prompt() in #waitForPostPromptRecovery during shutdown.
Guard the predicate on !#isDisposed so lifecycle aborts settle the turn, and add a regression test. Addresses review feedback on #2689.
Empty provider-side aborted turns now enter the existing auto-retry path without switching retry model fallback, while aborted turns with partial content still settle normally.\n\nFixes #2685
The first-message eager-task / eager-todo preludes are the oldest messages in a
session, so auto-compaction summarizes them away and the agent silently loses the
delegate-via-tasks / phased-todo guidance mid-work. Re-assert those reminders on
the auto-continuation turn that follows a compaction.
- Widen #createEagerTaskPrelude / #createEagerTodoPrelude to accept
`string | undefined`; `undefined` (post-compaction) skips only the
first-message and prompt-suffix gates, keeping the mode / agent-kind /
plan-mode / surviving-todo / active-tool gates intact.
- Reminder-only post-compaction: the todo nudge never attaches a forced `todo`
tool_choice on the resumed turn (forcing a tool after a mid-turn compaction
would override the agent's in-flight action).
- Add #buildPostCompactionEagerNudges() and prepend its output on the single
#scheduleAutoContinuePrompt continuation hook. All three call sites are
willRetry-safe, so overflow/incomplete retry recoveries never carry the nudge.
Op: extend
Matched retry fallback roles against the plain model selector as well as the routed in-flight selector, preserving configured chains for compat-routed OpenRouter and Vercel models.
Added regression coverage for a compat-routed OpenRouter primary using a plain role selector.