- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.
- Export `directoryExists` in `utils` to safely validate working directories before traversal.
- Update `SessionManager` and startup logic to fallback to the launch directory if a session's recorded working directory no longer exists.
- Add regression tests to ensure sessions now correctly adopt the launch directory instead of crashing on missing paths.
- Add `directoryExists` check to validate session working directories.
- Skip updating the session cwd if the recorded directory no longer exists on disk to prevent runtime errors.
- Replaced four individual power boolean settings with a single `power.sleepPrevention` enum for improved configuration management.
- Implemented automatic migration logic in `Settings.init` to map legacy macOS power booleans to the new enum levels.
- Updated `AgentSession` to utilize the new `sleepPrevention` enum for macOS power assertion lifecycle management.
- Changed default value of `display.cacheMissMarker` to `false` to suppress cache-miss markers.
- Added comprehensive unit tests in `settings-manager.test.ts` to verify backward-compatible migration behavior.
- Added tracking for expected cache invalidations during model changes, compactions, and plan-mode transitions.
- Included `cacheMissExplainedAt` metadata in session context to prevent displaying misleading cache miss warnings in the transcript.
- Updated controller logic to reset assistant usage markers when mode-switching or performing actions that invalidate the prompt cache.
- Renamed `repeatToolDescriptions` to `inlineToolDescriptors` throughout configuration, SDK, and internal session management.
- Set the default value to `true` and updated descriptions to clarify the descriptor inlining behavior.
- Fixed the `/dump` command to prevent duplicate tool inventory output when inlining is enabled.
- Introduced `SoftToolRequirement` to support non-invasive tool enforcement with lifecycle management and escalation.
- Added `ToolChoiceDirective` to coordinate hard and soft tool requirements within the agent loop.
- Optimized preview workflows in `coding-agent` by replacing forced tool choices with non-forcing pending invokers.
- Enhanced `CompactionSummaryMessage` to prioritize structured rendering for tool requirement reminders.
- Changed `wrapSteeringForModel` to wrap every `steering:true` user message regardless of position, so a steer's wire bytes stay identical once buried instead of reverting to raw and busting the prompt cache.
- Neutralized the present-tense framing in `user-interjection.md` so always-wrapping no longer leaves stale "current task" wording on buried interjections.
- Updated `session-messages.test.ts` to assert buried steering messages are wrapped too.
- Added `keepBoundaryId` and `cacheWarmSuffixTokens` guards to `pruneToolOutputs` and `pruneSupersededToolResults` so superseded/useless results sitting in the already-sent cached prefix are no longer rewritten mid-session, with new `computeMessageSuffixTokens`/`resolveBoundaryIndex` helpers.
- Added a `keepBoundaryId` floor to `collectShakeRegions` so shake skips entries summarized away by the latest compaction.
- Threaded `firstKeptEntryId`, `PRUNE_CACHE_WARM_SUFFIX_TOKENS` (8k) and `PRUNE_IDLE_FLUSH_MS` (90m, above the 1h cache TTL) from `#pruneToolOutputs`, `#pruneStaleToolResults` and `shake` in `agent-session.ts`.
- Added six boundary tests in `supersede-prune.test.ts` covering warm-prefix protection, tail-case pruning, and the pre-boundary floor.
- Added utility functions to strip descriptions from JSON schemas and tool definitions for optimized token output.
- Integrated `pruneToolDescriptions` configuration across agent loops and sessions to enable optional schema pruning.
- Updated agent context and snapshot logic to propagate pruning settings and maintain fingerprinting integrity.
- Verified schema structural integrity and removal of annotation descriptions through new unit tests.
- Update `MAX_FRAMES_DEFAULT` to 80 to better utilize high-capacity model context windows.
- Remove `providerFrameBudget` from `snapcompact` to decouple archival limits from provider-specific image caps.
- Update `compact` logic to treat `maxFrames` as a hard upper bound rather than a provider-clamped limit.
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
- Standardized elision markers across all tool outputs and filters to use cohesive `[...N [type] elided...]`, `[...Nln elided...]`, and `[...xB elided...]` syntax.
- Updated documentation, prompts, and test expectations to reflect the unified elision format.
- Improved transcript viewer robustness by preventing content aliasing through path-inclusive signature hashing.
- Added logic to clear stale transcript content when associated session files are deleted, accompanied by verifying test cases.
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
- Added a `mode` property to `CompactOptions` to allow fine-grained control over compaction strategies.
- Implemented `soft`, `remote`, and `snapcompact` submode overrides for the `/compact` command.
- Integrated `parseCompactArgs` to enable robust subcommand routing and validation, including focus instruction rejection for specific modes.
- Established a `CompactMode` registry to manage compaction strategies and verify remote availability.
The dispose() disconnect added by this PR awaited mcpManager.disconnectAll()
unbounded. An owned manager holding an HTTP/SSE server whose session-
termination DELETE hangs would block dispose for the full MCP request timeout
(30s default, unbounded when OMP_MCP_TIMEOUT_MS=0), stalling /exit and
print-mode shutdown on a broken remote endpoint.
Wrap the disconnect in withTimeout(..., 3_000) — mirroring the bounded
async-job teardown two lines above and the startup bound from issue #2100.
stdio close (the subprocess reap this PR targets) completes well within the
bound; a slow transport close is left to finish detached.
Adds a regression test driving the real MCPManager.disconnectAll() with a
stalled transport close.
After plan approval the executor delivers the plan-mode-reference exactly
once and sets `#planReferenceSent = true`. Both compaction paths — `compact()`
and `#runAutoCompaction()` — replace the conversation history that carried
that reference but never cleared the flag, so `#buildPlanReferenceMessage()`
short-circuited to null on every subsequent turn and the executor permanently
lost the plan it was working on (exactly the long-session failure reported).
Clear `#planReferenceSent` right after `replaceMessages()` in both paths so the
next turn re-reads the plan from disk and re-injects it. The reset is a no-op
for ordinary sessions: the default plan path (PLAN.md in session-local scratch)
has no file on disk, so `#buildPlanReferenceMessage()` still returns null there.
Adds a deterministic regression test (short-circuited compaction, mock stream)
that fails before this change and passes after, plus a guard proving normal
sessions get no spurious plan injection.
Fixes#1246
- Refactor deep imports by targeting specific sub-modules in `@oh-my-pi/pi-ai` to reduce barrel file overhead.
- Utilize jitless ArkType scopes in schema definitions to reduce startup JIT codegen costs by approximately 65%.
- Reorganize internal `auth-storage` exports to maintain clean boundaries between core and broker-specific functionality.
- Refactor replay safety logic to specifically prevent retries when a tool call is present in the assistant message.
- Enable retries for transient stream errors occurring during partial text or thinking sequences, ensuring consistency when no tool call has been completed.
- Add regression tests to verify that transient socket closures are successfully recovered while completed tool calls remain protected from redundant retries.
- Add an epoch counter to `AdvisorRuntime` to discard in-flight advisor batches when a reset or disposal occurs.
- Introduce `resetAdvisorSessionState` to clear advisor-specific queues, latches, and pending cards, ensuring pre-reset state does not interfere with new conversations.
- Extend `YieldQueue.clear` to support conditional clearing by entry kind.
- Expired the per-credential usage report cache after recording observed OpenCode Go spend.\n- Threaded provider base URL into OpenCode Go cost recording so the invalidation targets the same cache key /usage uses.\n- Added regression coverage for refreshing cached OpenCode Go limits immediately after a completed turn.\n\nFixes #2942
- Added an OpenCode Go usage provider that synthesizes 5h, weekly, and monthly cap windows from OMP-observed request costs.\n- Recorded OpenCode Go assistant request costs against the active credential so /usage can report local cap utilization.\n- Added regression coverage for fresh keys and observed spend aggregation.\n\nFixes #2942
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
Stopped auto context-full maintenance from retrying repeated summarization timeouts on the same model before fallback. Added a regression test for timeout fast-fallback behavior.\n\nFixes #2913
Two defects dropped the first steering/follow-up message typed as
auto-compaction began:
- The compaction AbortController (which backs isCompacting) was installed
AFTER auto_compaction_start was emitted. The emit awaits extension
delivery and yields to the event loop, so a message typed as the loader
appeared was read while isCompacting was false and mis-routed into the
core agent queue (which the handoff reset then wiped). Install the
controller before the emit, and move the emit to the first statement
inside the existing try so the catch/finally cleanup still runs, in both
#runAutoCompaction and #runAutoShake. The handoff branch now passes the
run's local abort signal instead of the mutable controller field, so a
superseded run bails at the handoff entry check rather than resetting the
session.
- handoff() calls agent.reset(), which clears the core steering/follow-up
queues. Capture both queues immediately before reset and restore them
immediately after (synchronous, no await gap), so queued steers and
follow-ups -- including in-flight RPC/SDK steer()/followUp() and a hidden
user companion such as an ultrathink notice -- survive the new-session
reset instead of being silently dropped.
Adds regression tests for both defects (controller-before-emit for the
context-full and shake paths, queue preservation across the reset for both
pre-enqueued and in-flight messages).
- Updated the `primaryArg` selection logic to check for the `advise` tool explicitly.
- Preformatted the summary format for advice as `{severity}: {note}` when both are present, falling back to either if only one exists.
- Updated and verified the associated test suite to confirm the output structure.
- Added "note" to the list of primary argument keys for formatting session history.
- Added an explicit type annotation to the inline savings unit test options.
- Extracted native token counting into a new localized `tokenizer.ts` wrapping `@oh-my-pi/pi-natives`.
- Introduced a lightning-fast byte-length estimation logic for token counting when accurate counting is disabled.
- Diverted token calculations to the faster estimator during test environments and when `PI_TOKENIZER_ACCURATE` is falsy.
- Updated agent base and coding-agent sessions to consume the new localized `countTokens` utility.
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Bound native Reflect methods to local variables in the browser launch script to prevent detection.
- Updated stealth injection scripts to use the bound Reflect methods instead of global Reflect calls.
- Fixed a type assertion issue in the AgentSession tool proxy.