parseModelString now extracts valid thinking level suffixes (e.g.,
"anthropic/claude-opus-4-6:high") instead of treating them as part of
the model ID. This enables per-role thinking levels in config:
modelRoles:
slow: anthropic/claude-opus-4-6:high
default: anthropic/claude-opus-4-6:low
smol: google/gemini-3-flash:medium
The thinking level is applied at startup, in SDK fallback resolution,
and during Ctrl+P role cycling. The original config string is preserved
on role cycle so the suffix round-trips correctly.
* feat(mcp): resource notifications, subscriptions, and read_resource builtin tool
- Add MCP resource subscription lifecycle (subscribe/unsubscribe on connect/disconnect)
- Wire mcp.notifications setting with live toggle support
- Add debounced followUp injection for resource change notifications
- Add global read_resource builtin tool with server resolution by URI/template scheme
- Add MCP prompt commands (buildMCPPromptCommands) with array content support
- Add server instructions injection into system prompt with attribution
- Add mcp.notificationDebounceMs configurable setting
Client (client.ts):
listResources, listResourceTemplates, readResource with pagination
subscribeToResources, unsubscribeFromResources
listPrompts, getPrompt, serverSupportsPrompts
serverSupportsResources, serverSupportsResourceSubscriptions
Manager (manager.ts):
Notification dispatch with subscribed-URI guard
Concurrent refresh deduplication via pending promise map
setNotificationsEnabled with subscribe/unsubscribe toggle
Tests:
client-resources.test.ts (31 tests)
client-prompts.test.ts (20 tests)
mcp-read-resource.test.ts (13 tests)
* fix(mcp): address PR review - eager prompt init and stale subscription cleanup
P1: Make setOnPromptsChanged eagerly fire for servers that already
have prompts loaded. The callback is registered after MCP discovery
has already loaded prompts and fired the hook, so without this the
handler is never called on the common startup path. The fix is in
the manager itself (not the caller), eliminating the race condition
regardless of when the callback is wired.
P2: Unsubscribe removed resource URIs on resource refresh.
refreshServerResources was subscribing to the new URI set and
overwriting #subscribedResources without unsubscribing URIs that
were previously subscribed but no longer present, leaving stale
subscriptions active on the server.
* fix(mcp): add resources and prompts to /mcp help text and subcommand completions
* feat(mcp): add /mcp notifications command
Shows per-server notification capabilities with subscription state:
- Lists supported notification types (tools/list_changed, resources/list_changed,
prompts/list_changed) with check marks
- Shows resources/subscribe status with active subscription count
- Lists subscribed URIs with green ticks when notifications are enabled
- Displays overall enabled/disabled state (mcp.notifications setting)
* fix(mcp): address PR review comments on race conditions and stale state
- Await subscribe/unsubscribe in refreshServerResources so the refresh
promise doesn't resolve before subscriptions are settled, preventing
a second refresh from racing and overwriting tracking state (P2 #3)
- Guard setNotificationsEnabled subscribe .then() against a disable
that happens while the subscribe request is in-flight (P2 #5)
- Re-check mcp.notifications setting inside debounce setTimeout
callback so toggling off mid-window actually suppresses the
follow-up message (P2 #4)
- Fire onToolsChanged and onPromptsChanged callbacks in
disconnectServer so stale slash commands and tool registrations
are cleaned up when a server is removed (P2 #2)
---------
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Add `rate-limit-utils.ts` to classify 429/503 errors (Quota, Rate Limit, Capacity)
- Fix `google-gemini-cli` to fail-fast on 429s instead of getting stuck in internal retries
- Update `isUsageLimitErrorMessage` regex in `AgentSession` to catch all Google-specific error variations
- Implement smart backoff timings (30m for quota, 30s for rate limit, 45s+jitter for capacity)
- Remove 0% hiding logic in `/usage` to always display account rotation pool
- Add comprehensive unit tests for rate limit parsing and provider behavior
- Enforced tool decision in plan mode--agent now requires calling either `ask` or `exit_plan_mode` when a turn ends without a required tool call.
- Fixed cancellation behavior of `ask` tool to abort the current turn instead of returning a normal cancelled selection, while timeout-driven auto-cancel still returns without aborting.
- Added plan-mode-tool-decision-reminder system prompt to guide agent when required tools are not called.
- Improved agent_end event handling to use fallback assistant message when #lastAssistantMessage is unavailable.
- Added checkpoint and rewind tools to create context checkpoints before exploratory work and rewind to replace exploration messages with concise reports.
- Added checkpoint.enabled setting to control availability of checkpoint and rewind tools in agent sessions.
- Added getCheckpointState() and setCheckpointState() methods to agent session API for checkpoint state management.
- Implemented checkpoint state tracking with message count, entry ID, and timestamp to enable context cost optimization during investigations.
- Removed normativeRewrite setting and buildNormativeUpdateInput() function from public API.
- Removed $normative property and TNormative generic parameter from ToolResultMessage and AgentToolResult interfaces.
- Deleted normative.ts patch normalization module and related helper functions for diff anchor processing.
- Removed rewriteAssistantToolCallArgs() and #rewriteToolCallArgs() methods that modified tool call arguments.
## New: schema compatibility validation API
Add `validateSchemaCompatibility(schema, provider)` in
`packages/ai/src/utils/schema/compatibility.ts` that performs a static
audit of a JSON Schema against three provider targets:
- `openai-strict`: checks forbidden keys, required/properties symmetry,
additionalProperties constraint, and that every node declares a type,
combinator, or $ref
- `google`: checks unsupported keyword set and array-valued type
- `cloud-code-assist-claude`: checks forbidden keywords, array type,
null type, nullable keyword, and combiner presence; also validates via
AJV 2020 draft
Add `validateStrictSchemaEnforcement(original, result)` to assert the
fail-open contract: when strict enforcement succeeds the output must pass
openai-strict validation; when it fails the output must be the original
schema object (same reference).
Export both functions and their types from `./utils/schema/index.ts`.
## New: shared constants in fields.ts
Extract `COMBINATOR_KEYS` (`anyOf`, `allOf`, `oneOf`) and add
`CCA_UNSUPPORTED_SCHEMA_FIELDS` as exported constants, eliminating the
local duplicate in `strict-mode.ts` and providing a canonical field set
for Cloud Code Assist (much narrower than the Google set — CCA supports
validation keywords like `additionalProperties`, `minLength`,
`pattern`, etc.).
## Fix: cycle detection in all recursive schema traversals
All recursive walkers now carry a `WeakSet<object>` guard. Previously any
schema with a reference cycle (or a schema object that appears at two
nodes in the tree) would cause an infinite loop or a stack overflow:
- `sanitizeSchemaForStrictMode` / `enforceStrictSchema`
- `normalizeSchemaForCloudCodeAssistClaude`
- `normalizeNullablePropertiesForCloudCodeAssist`
- `stripResidualCombiners`
- `sanitizeSchemaImpl` (Google sanitizer)
- `hasResidualCloudCodeAssistIncompatibilities`
`hasResidualCloudCodeAssistIncompatibilities` previously returned `true`
for already-visited nodes, producing false positives that forced the CCA
fallback schema on valid (but multiply-referenced) schemas. It now
correctly returns `false`.
## Fix: stripResidualCombiners iterates to fixpoint
The previous single-pass approach missed chained combiner reductions
where one collapsed variant exposed another reducible combiner. The
rewriter now loops until no further reduction occurs.
## Fix: mergeObjectCombinerVariants required-field computation
The merged object schema now takes the intersection of all variants'
`required` arrays, then unions in own-level required properties that
exist in the merged schema. Previously the `required` field was silently
dropped from the flattened schema, making all properties effectively
optional.
## Fix: sanitizeSchemaForGoogle improvements
- Type inference for const-collapsed enums: type is derived from all
variants (must unanimously agree), falling back to inference from enum
values; mixed null/non-null infers the non-null scalar type and sets
`nullable: true`
- Const→enum deduplication now uses deep structural equality instead of
`Object.is`
- Recursion spreads the full options object so new fields (`unsupportedFields`,
`seen`) are not silently dropped when descending into sub-schemas
- Array-valued `type` is filtered to strings before processing
- Removed incorrect stripping of `additionalProperties: false` (the
field is valid and should be preserved)
- Parameterized `unsupportedFields` in `SanitizeSchemaOptions` enables
code reuse between the Google and CCA sanitizers
## Fix: sanitizeSchemaForStrictMode / enforceStrictSchema
- `nullable: true` is now stripped during sanitization and expanded into
`anyOf: [schema, {type: "null"}]` in the enforcer output, matching
what OpenAI strict mode requires
- Type inference: `type: "array"` is inferred when `items` is present;
a scalar type is inferred from uniform `enum` values
- Const→enum merge uses deep equality to avoid duplicate entries when
both `const` and `enum` exist with the same value
- `additionalProperties` is now dropped unconditionally in sanitization
(previously only object-valued `additionalProperties` was recursed;
non-object values were passed through)
- `enforceStrictSchema` recurses into `$defs` and `definitions` blocks
- `enforceStrictSchema` handles tuple-style `items` arrays
- `enforceStrictSchema` skips double-wrapping: optional properties
already expressed as `anyOf: [..., {type: "null"}]` are not wrapped again
- `tryEnforceStrictSchema` now caches results in a `WeakMap` keyed on
the input schema object to avoid redundant work on repeated calls
## Fix: mergeCompatibleEnumSchemas deep equality
Uses `areJsonValuesEqual` instead of `Object.is` when deduplicating
enum members, so structurally equal objects are not duplicated.
## New: test coverage
- `packages/ai/test/schema-normalization.test.ts`: comprehensive unit
tests for strict mode, Google, and Cloud Code Assist normalization
- `packages/ai/test/schema-compatibility.test.ts`: unit tests for all
three provider targets in the new compatibility validator
- `packages/coding-agent/test/tools/provider-schema-compatibility.test.ts`:
integration test that instantiates every builtin and hidden tool, runs
their parameter schemas through all three provider pipelines, and
asserts zero compatibility violations
- Added public `waitForIdle()` and `getLastAssistantMessage()` APIs to AgentSession for deterministic session state access.
- Refactored deferred continuation scheduling from raw `setTimeout()` to centralized post-prompt task tracking system for concurrent recovery operations.
- Fixed race conditions between deferred TTSR/context-promotion continuations and `prompt()` completion via shared recovery orchestrator.
- Replaced `#waitForRetry()` with `#waitForPostPromptRecovery()` to unify retry and TTSR resume gate handling.
- Fixed TTSR violations during subagent execution aborting the entire subagent run; `#waitForPostPromptRecovery()` now awaits agent idle after TTSR/retry gates resolve, preventing `prompt()` from returning while fire-and-forget `agent.continue()` is still streaming.
- Added comprehensive test case verifying `prompt()` blocks until TTSR continuation with tool calls completes, preventing premature session disposal.
- Implemented TTSR resume gate to ensure `prompt()` blocks until TTSR interrupt continuations complete, preventing race conditions between TTSR injections and subsequent prompts.
- Replaced `#waitForRetry()` with `#waitForPostPromptRecovery()` to handle both retry and TTSR resume gates, ensuring prompt completion waits for all post-prompt recovery operations.
- Added comprehensive test coverage for TTSR resume gate behavior under interrupt and deferred injection modes.
- Consolidated @oh-my-pi/pi-utils subpath imports into single package root import across 100+ files.
- Moved tryParseJson utility from local web scrapers module to @oh-my-pi/pi-utils package for centralized JSON parsing.
- Renamed loadSkillsFromDir to scanSkillsFromDir and refactored skill discovery to use fs.promises.readdir instead of glob-based approach.
- Replaced custom parseJSON with tryParseJson across discovery modules for consistent error handling.
- Removed emitCustomToolSessionEvent method and cleanupSshResources function, consolidating shutdown logic into dispose method.
- Updated glob pattern construction to use GlobBuilder with literal_separator(true) for improved path handling.
- Added getTodoPhases() and setTodoPhases() methods to ToolSession API for in-memory todo phase management.
- Added getLatestTodoPhasesFromEntries() export to retrieve todo phases from session history entries.
- Changed todo state management from file-based (todos.json) to in-memory session cache with automatic persistence.
- Changed todo phases to sync from session branch history during branching and rewriting operations.
- Removed file-based todo loading logic and replaced with session-based todo phase retrieval throughout codebase.
- Renamed the `notes://` protocol to `local://` for better clarity.
- Updated all internal references, prompts, and tool documentation.
- Migrated plan storage paths to use the new `local://` scheme.
- Renamed XML tags from underscore to kebab-case format for consistency across prompts and system messages.
- Updated context tag from `swarm_context` to `context` in render logic and test assertions.
- Consolidated conditional logic in subagent user prompt by removing duplicate assignment blocks.
- Updated system prompt documentation to reflect kebab-case naming convention for XML tags.
- Replaced plan:// protocol with notes:// for session-scoped artifact storage and plan finalization.
- Added title parameter to exit_plan_mode tool to enable plan file renaming during approval workflow.
- Implemented NotesProtocolHandler for notes:// URL scheme with path traversal protection and session fallback.
- Added renameApprovedPlanFile function to handle plan artifact finalization with validation and error handling.
- Updated system prompt documentation to reference notes:// protocol and internal URL schemes for artifact access.
- Standardized XML tag naming from snake_case to kebab-case across 50+ prompt files for consistency.
- Replaced imperative language with RFC 2119 keywords (MUST/SHOULD/MAY/MUST NOT) throughout system and tool prompts for clarity.
- Removed artifactsDir parameter from Python executor and simplified environment variable handling to use PI_SESSION_FILE only.
- Renamed read_path.md to read-path.md and updated memory guidance with hierarchy rules and conflict resolution workflow.
- Added noEscape option to bash URL expansion and extracted cwd parameter from leading cd commands for improved path handling.
- Exported NO_PAGER_ENV constant from bash-interactive module for centralized environment variable management.
- Added stripInternalArgs() utility function to filter harness-internal keys from tool arguments.
- Hidden agent__intent parameter from UI and log displays across agent, session, MCP, and tool-execution components.
- Implemented HIDDEN_ARG_KEYS constant to centrally manage internal argument filtering.
- Updated formatArgsInline() to exclude internal keys when rendering tool arguments.
- Added async background job execution for bash and task tools with configurable concurrency limits and automatic result delivery.
- Added cancel_job tool and /jobs slash command to manage and inspect running background jobs with status display.
- Added jobs:// internal protocol handler for querying job status and retrieving job execution details.
- Added async.enabled and async.maxJobs settings to control background job execution behavior.
- Enhanced status line to display count of running background jobs with visual indicator.
- Implemented AsyncJobManager with exponential backoff retry delivery, job lifecycle tracking, and automatic eviction.
Fixes#56.
- Moved artifact management from ToolSession to SessionManager for centralized lifecycle control and caching.
- Replaced getArtifactManager() with allocateOutputArtifact() async method in ToolSession interface for simplified artifact allocation.
- Updated bash, fetch, python, and ssh tools to call session.allocateOutputArtifact() directly with optional chaining fallback.
- Fixed Lobsters scraper to handle user fields as strings instead of nested objects in API responses.
- Refactored byte truncation to use unified `truncateBytesWindowed` function supporting both head and tail modes, reducing code duplication.
- Optimized `truncateHead` and `truncateTail` to avoid full Buffer allocation by processing content incrementally with character-level scanning.
- Improved `TailBuffer.append()` to handle large incoming chunks more efficiently by detecting when a single chunk dominates the tail budget.
- Enhanced `OutputSink.push()` to avoid creating giant intermediate strings when spilling to files by windowing large chunks before concatenation.
- Refactored newline counting to use a constant `NL` for consistency across the module.
- Migrated AuthCredentialStore and AuthStorage to shared modules, standardizing credential management and soft deletion.
- Consolidated Anthropic authentication and various other formatting/utility helpers into shared modules.
- Improved unicode normalization in patch logic with regex for efficiency and correctness.
- Enhanced JTD type guards for robust schema validation.
- Extracted credential storage to shared @oh-my-pi/pi-ai package with AuthCredentialStore and AuthStorage classes.
- Consolidated UI formatting logic from ToolUIKit class into standalone utility functions across render-utils and output-meta modules.
- Moved utility functions (parseCommandArgs, substituteArgs, expandPath, normalizeUnicode) to dedicated modules for improved code reuse.
- Extracted JTD type definitions and type guards to jtd-utils module for shared use across schema conversion tools.
- Updated Claude model pricing and added cache read costs in models.json for accurate billing calculations.
- Refactored agent-storage to delegate credential management to AuthCredentialStore instead of direct SQLite operations.
- Added GitLab Duo provider with support for Claude, GPT-5, and Duo Chat models via GitLab AI Gateway.
- Added OAuth authentication for GitLab Duo with automatic token refresh, PKCE security, and 25-minute token caching.
- Added 16 new GitLab Duo models including Claude Opus/Sonnet/Haiku and GPT-5 variants with reasoning and multimodal support.
- Added `isOAuth` option to Anthropic provider for OAuth bearer token authentication mode.
- Exported `streamGitLabDuo`, `getGitLabDuoModels`, and `clearGitLabDuoDirectAccessCache` functions for GitLab Duo integration.
- Consolidated truncation and output utilities from tools/truncate.ts and tools/output-utils.ts into session/streaming-output.ts with improved UTF-8 boundary handling.
- Renamed formatSize() to formatBytes() across codebase for consistency and clarity in byte-level formatting.
- Refactored OutputSink to use windowed byte truncation instead of full-buffer encoding, improving memory efficiency on large outputs.
- Migrated from Buffer to Uint8Array in web scrapers for better cross-platform compatibility and native browser support.
- Added getArtifactManager() lazy-initialization method to ToolSession for deferred artifact manager instantiation.
- Simplified API surface with wildcard exports from tools and session modules, reducing import complexity.
- Exported readModelCache and writeModelCache functions for SQLite-backed model cache access.
- Migrated model cache storage from per-provider JSON files to unified SQLite database (models.db).
- Renamed cachePath option to cacheDbPath in ModelManagerOptions for database-backed storage.
- Improved non-authoritative cache handling with 5-minute retry backoff instead of per-startup retries.
- Added peekApiKey method to AuthStorage for non-blocking API key retrieval during model discovery.
- Added <turn_aborted> guidance marker as synthetic user message for aborted/errored assistant messages.
- Improved OAuth token refresh error messages to include provider-specific error details from API responses.
- Separated rate limit and usage limit error handling in OpenAI response handler with distinct error messages and retry timing.
- Enhanced error message propagation in OAuth refresh flow to preserve original error reasons for better debugging.
- Fixed regex pattern in auth storage to use word boundaries for accurate HTTP status code matching.
- Added `includeDisabled` parameter to `listAuthCredentials()` to optionally retrieve disabled credentials.
- Added `disableAuthCredential()` method for soft-deleting auth credentials while preserving database records.
- Changed auth credential removal to use soft-delete (disable) instead of hard-delete when OAuth refresh fails, keeping credentials in database for audit purposes.
- Added `disabled` column to auth_credentials table schema with automatic migration from v3 to v4.
- Added prepared statements for querying active (non-disabled) credentials with optional provider filtering.
- Added streamed tool intent display in working message to show real-time intent tracking during agent execution.
- Changed intent tracing field name from `$intent` to `_intent` across tool schemas and agent core for consistency.
- Added support for file deletion and renaming operations in hashline edit mode.
- Renamed hashline edit operation fields: `set` to `target`/`new_content`, `set_range` to `first`/`last`/`new_content`, `insert` to `inserted_lines`.
- Added 'unable to connect' to transient error patterns in retry logic to properly handle connection failures.
- Updated retry error detection in coding-agent to match ai package pattern for consistency.
- Reorganized import statements in openai-compat.ts for consistency.
- Consolidated multi-line ternary expression into single line in openai-compat.ts.
- Simplified AgentBusyError instantiation to use default message.
- Updated test assertion to check error type instead of message content.