- Introduced `SoftToolRequirement` to support non-invasive tool enforcement with lifecycle management and escalation.
- Added `ToolChoiceDirective` to coordinate hard and soft tool requirements within the agent loop.
- Optimized preview workflows in `coding-agent` by replacing forced tool choices with non-forcing pending invokers.
- Enhanced `CompactionSummaryMessage` to prioritize structured rendering for tool requirement reminders.
- Added `keepBoundaryId` and `cacheWarmSuffixTokens` guards to `pruneToolOutputs` and `pruneSupersededToolResults` so superseded/useless results sitting in the already-sent cached prefix are no longer rewritten mid-session, with new `computeMessageSuffixTokens`/`resolveBoundaryIndex` helpers.
- Added a `keepBoundaryId` floor to `collectShakeRegions` so shake skips entries summarized away by the latest compaction.
- Threaded `firstKeptEntryId`, `PRUNE_CACHE_WARM_SUFFIX_TOKENS` (8k) and `PRUNE_IDLE_FLUSH_MS` (90m, above the 1h cache TTL) from `#pruneToolOutputs`, `#pruneStaleToolResults` and `shake` in `agent-session.ts`.
- Added six boundary tests in `supersede-prune.test.ts` covering warm-prefix protection, tail-case pruning, and the pre-boundary floor.
- Added utility functions to strip descriptions from JSON schemas and tool definitions for optimized token output.
- Integrated `pruneToolDescriptions` configuration across agent loops and sessions to enable optional schema pruning.
- Updated agent context and snapshot logic to propagate pruning settings and maintain fingerprinting integrity.
- Verified schema structural integrity and removal of annotation descriptions through new unit tests.
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
- Standardized elision markers across all tool outputs and filters to use cohesive `[...N [type] elided...]`, `[...Nln elided...]`, and `[...xB elided...]` syntax.
- Updated documentation, prompts, and test expectations to reflect the unified elision format.
- Improved transcript viewer robustness by preventing content aliasing through path-inclusive signature hashing.
- Added logic to clear stale transcript content when associated session files are deleted, accompanied by verifying test cases.
- Added an exclusion check to `extractFileOpsFromMessage` to ignore paths containing URL schemes (e.g., `artifact://`, `conflict://`, `https://`).
- Implemented `isUrlSchemePath` helper to identify non-filesystem resources that cannot be re-grounded by the agent.
- Added comprehensive unit tests to verify exclusion logic for various scheme-prefixed paths.
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
- Extracted native token counting into a new localized `tokenizer.ts` wrapping `@oh-my-pi/pi-natives`.
- Introduced a lightning-fast byte-length estimation logic for token counting when accurate counting is disabled.
- Diverted token calculations to the faster estimator during test environments and when `PI_TOKENIZER_ACCURATE` is falsy.
- Updated agent base and coding-agent sessions to consume the new localized `countTokens` utility.
- Introduce a `transformAssistantMessage` lifecycle hook to `AgentOptions`, `Agent`, and `AgentLoopConfig`.
- Invoke the hook on finalized assistant messages within the agent streaming loop prior to context updates, UI emission, and tool dispatch.
- Enable mutating the text content and tool-call arguments of assistant messages in place during execution.
- Added native support for parsing, normalized validation, and serialization of ArkType schemas throughout the agent pipeline.
- Implemented pruning of unconstrained union branches and normalization helpers to handle unrepresentable strict-mode branches.
- Patched ArkType's schema package to preserve declared object key order during serialization.
- Integrated ArkType schemas into `agent-loop` and migrated coding agent tool parameters to ArkType format.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Added Gemini thinking-loop detection helpers for near-duplicate and verbatim output checks.
- Wrapped `stream`, `streamPiNative`, and `streamSimple` dispatches with the loop guard.
- Emitted retryable empty-content loop errors and stopped completion events on loop hits.
- Added `enableGeminiThinkingLoopGuard` options for OpenAI compatibility with Gemini defaults and overrides.
The deadline was only checked at loop entry and the top of the inner loop,
so a provider response returning tool calls after the deadline still
executed them, and a deadline crossed during onBeforeYield could dequeue
then drop queued steering/follow-up messages.
- Re-check the deadline after streamAssistantResponse before running tools;
pair leftover tool calls with aborted "Deadline exceeded" placeholders to
keep the tool_use/tool_result contract valid.
- Re-check after emitTurnEnd and after onBeforeYield before draining the
steering/aside/follow-up queues so queued messages are never dequeued and
dropped.
- Merge a deadline AbortController into the loop signal so in-flight provider
requests and tools are cancelled at the boundary.
- Updated getApiKey signatures to accept a Model and return ApiKey or ApiKeyResolver.
- Updated stream key handling to resolve credentials per model and use seedApiKeyResolver for retries.
- Added antigravityEndpointMode setting with auto/production/sandbox endpoint selection.
- Added 429/5xx endpoint failover for Gemini stream, usage, search, and image calls.
- Added awaited `onTurnEnd` and `setOnTurnEnd` wiring for turn-end callbacks.
- Added `advisor.syncBacklog` settings (off/1/3/5) and documented 30-second catch-up caps.
- Fixed advisor runtime backlog handling with failure counters, waiters, and retry requeue.
- Updated agent sessions to enqueue advisor updates on turn end and removed direct `turn_end` branch logic.
Prevented auto-retry from regenerating write calls after a provider stream timeout has already exposed assistant content or tool-call arguments.
Fixes#2683
- Replaced `PI_OWNED_TOOLS` lookups and owned-dialect documentation in `agent-loop.ts` and `types.ts` with `PI_DIALECT`.
- Added `prompt-tools-loop.test.ts` coverage that an unset `config.dialect` still selects Hermes from `Bun.env.PI_DIALECT`.
- Documented the breaking `PI_DIALECT` rename in `packages/agent/CHANGELOG.md`.
- Renamed ToolCallSyntax type to Dialect and Grammar interface to DialectDefinition across all packages.
- Moved grammar directory to dialect and updated all import paths in agent, ai, catalog, and coding-agent packages.
- Added renderTranscript and renderThinking methods to DialectDefinition, enabling native dialect-aware conversation serialization.
- Consolidated rendering utilities into new dialect/rendering.ts with shared helpers for ChatML, legacy text, and dialect-specific formatting.
- Updated conversation serialization in agent and coding-agent to use dialect.renderTranscript() for native turn envelope rendering.
- Canonicalized assistant and thinking messages by trimming and collapsing dot text.
- Skipped rendering assistant and thinking blocks when canonicalized content was empty.
- Filtered ACP thinking notifications and session outputs to ignore placeholder content.
- Added canonicalizeMessage tests for undefined, blank, whitespace, and dot-only inputs.
- Updated compaction, branch summarization, and session dump formatting to pass preferred model tool syntax into conversation serialization.
- Enhanced shared serializers to render assistant tool calls and tool results through grammar envelopes when syntax is available, with the prior compact format as fallback.
- Aligned prompt, preview, and test fixtures to the new transcript tags: `[Think]`, `[Tool Call]`, and `[Tool Result]`.
- Added an `interruptible` field to AgentTool and documented when it is honored.
- Updated immediate-mode tool execution to poll steering during in-flight interruptible calls and abort them when steering is queued.
- Marked the coding `job` tool as interruptible and added tests covering mid-wait aborts versus boundary-only steering drain.
- Added Gemini and Gemma syntax routing by model family and owned syntax env values.
- Added Gemini and Gemma in-band parsers for tool_code and token-based tool_call streams.
- Added rendering support for Gemini fenced tool_code/tool_outputs and Gemma tool tokens.
- Fixed parsing edge cases for comments, string escapes, nested args, and truncated blocks.
- Added an abortOnFabricatedToolResult option to Agent and AgentLoopConfig to choose whether in-band fabricated tool results are aborted or drained.
- Propagated the option through agent loop wiring into wrapInbandToolStream so fabrication is aborted only when enabled.
- Exposed the setting in coding-agent as tools.abortOnFabricatedResult and wired it through session creation with a true default.
- Added model-to-syntax mapping in catalog with preferred tool-call syntax API.
- Added `ToolExample` typing and `ToolCallSyntax` exports across tool/grammar interfaces.
- Added syntax-aware tool example rendering through provider-specific grammar invocations.
- Added `exampleSyntax` context flow and example metadata so rendered prompts include examples.
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
- Normalized agent `setSystemPrompt` to wrap string inputs into one-item arrays.
- Updated session creation to accept string `systemPrompt` values and normalize callback or direct results to string arrays.
- Adjusted extension result handling and test fixtures to accept string `systemPrompt` and missing `assistant_message` fields without crashing.
- agent-loop: raise repetition-detection floor to 180 chars and clear thinking
replay anchors when collapsing a detected loop.
- providers/google: ignore empty text parts, retain terminal thoughtSignatures,
and stop function-call signatures clobbering the prior block.
- autolearn: capture goal-mode at the turn boundary; harden managed-skill writes
against hard-links/symlinks (O_NOFOLLOW + nlink); refuse minting managed skills
whose name an authored skill already claims.
- eager tasks: thread agentKind through the session so a custom top-level agentId
still gets always-mode delegation; split Eager Tasks prompt into hard vs soft.
- title-generator: race the online title model against a local tiny-model fallback.
- eager-todo: keep the soft reminder aligned with the todo init schema.
- mcp/stdio: keep close() detaching the read loop instead of awaiting it.
- stream loop: fix collapsing and tool-call thought-signature handling.
- Validated queued toolChoice against active tools in agent and coding-agent sessions.
- Rejected queued forced choices with reason "unavailable" when selected tools were inactive.
- Dropped provider toolChoice payloads when requested function tools were not offered.
- Probed Tokio worker-thread support and fell back to current-thread runtime creation.
- Replaced queued-message interrupt flow with session abort calls on empty submit and escape.
- Removed interrupting state and notifyInterrupting teardown paths from abort handling.
- Updated AgentSession queue operations to use shared steering and follow-up queue views.
- Propagated isAborting through session state and collab payloads to suppress late updates.
- Coalesced concurrent interruptAndFlushQueuedMessages() calls through one in-flight promise.
- Replaced continue()-retry logic with agent.prompt() to flush queued messages from empty contexts.
- Skipped queued-message flush replay while compacting or streaming to avoid turn overlap.
- Updated flush path to consume queued steering first, then follow-ups, via dequeuing helper.
- Added a regression test for empty-state interrupt-and-flush delivering queued steers safely.
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
- Added `REMOTE_COMPACTION_TIMEOUT_MS` and wrapped remote compaction fetch signals with a timeout-backed `AbortSignal` to prevent indefinite hangs.
- Extended `requestOpenAiRemoteCompaction` and `requestRemoteCompaction` options with `timeoutMs` so callers can customize or rely on the new default watchdog.
- Optimized `trimOpenAiCompactInput` by caching serialized item sizes and decrementing a running budget as entries are removed, avoiding repeated JSON re-stringification.
- InputController now dispatched Esc to active viewSession operations, aborting compaction, handoff, and retry directly.
- Removed competing onEscape handler swaps across command and event controllers so overlapping auto/manual flow events no longer overwrote cancellation callbacks.
- Compaction now propagated fetch options and rethrew aborted signals so cancellations were not treated as remote failures.
- Re-polled steering at the loop yield boundary and included it in the pre-stop pending batch so late messages are processed immediately.
- Added session-side draining for stranded queued messages, scheduling an auto-continue when a prompt settles and follow-ups or steers remain.
- Added a regression test for late steering injection at yield and updated mid-turn collab prompt handling to keep steering messages in the pending display queue until consumed.
- Added optional `useless` flags to tool result types and payload builders.
- Added `pruneUseless` and `dropUeless` options to control uneventful result pruning.
- Changed compaction and shake passes to prune or ignore non-error useless tool results.
- Changed conversation serialization to omit useless toolCall/toolResult pairs from output.
- Added coverage for useless tagging, pruning, and serialization behavior.
- Added `hasSteeringMessages` to `executeToolCalls` polling so queued steering is detected without dequeuing.
- Retained fallback to `getSteeringMessages` by consuming messages when no peek callback exists.
- Added optional `hasSteeringMessages` config hook and limited steering checks to boundaries.
- Fixed interrupted tool-batch steering by keeping queued messages until boundary handling.
- Added idle text and image submissions to steer queueing when no input waiter exists.
- Auto-continued resumable sessions after queued steering and preserved submit metadata.
- Renamed all functions, types, and constants in @oh-my-pi/snapcompact to namespace-relative names (`snapcompactCompact` → `compact`, `renderSnapcompactFrames` → `renderMany`, `snapcompactFrameCount` → `frames`, `SnapcompactShape` → `Shape`, `SNAPCOMPACT_SHAPES` → `SHAPES`, …).
- Converted every consumer to `import * as snapcompact` member access: `agent/compaction.ts`, `coding-agent` `agent-session.ts`/`session-manager.ts`/`snapcompact-inline.ts`, and all affected tests.
- Renamed internal `geometry` locals to `geo` in `snapcompact.ts` to avoid TDZ collisions with the new `geometry` export.
- Updated `docs/compaction.md` prose and added a Breaking Changes entry to the snapcompact changelog documenting the full rename map.
- Added `renderSnapcompactFrames()` and `snapcompactFrameCount()` to @oh-my-pi/snapcompact for paging arbitrary text into PNG image blocks without dim-marker bookkeeping.
- Widened the agent loop's `transformProviderContext` hook to `(context, model) => Context` so per-request transforms can gate on the dispatch model's capabilities.
- Added `SnapcompactInlineTransformer` rendering the system prompt and large historical tool results as snapcompact frames on vision models: vision gate, per-provider image budgets, 3k-token floor, savings-margin gate, skip-last rule, and hash-keyed render caches swept to live tool calls.
- Added default-off `snapcompact.systemPrompt` and `snapcompact.toolResults` settings under a new Context → Experimental group, composed after secret obfuscation in `sdk.ts` so frames are built per-request and never persisted to session.jsonl.
- Added prompt stubs (`snapcompact-system-stub.md`, `snapcompact-system-frames-note.md`, `snapcompact-toolresult-note.md`) and unit tests covering frame paging, no-mutate guarantees, budget caps, gates, and render caching.