- Changed abort() method signature to return Promise<void> instead of void, making it async-compatible.
- Added bash executor fallback to one-shot shell execution when persistent sessions fail to respond to cancellation.
- Fixed bash execution timeout handling to prevent subsequent commands from hanging after hard timeouts.
- Extracted abort token management into ShellAbortState for thread-safe cancellation handling across shell sessions.
- Added SessionManager.close() method for proper cleanup of persistent writers and session resources.
- Added `close()` method to SessionManager and AuthStorage for proper resource cleanup and finalization of prepared statements.
- Added `initiatorOverride` option support in OpenAI and Anthropic providers for message attribution control.
- Fixed resource leaks in RpcClient timeout handling by centralizing timeout creation with unref() and adding explicit clearTimeout() calls.
- Fixed AgentSession disposal to call SessionManager's `close()` method for guaranteed resource cleanup instead of fallback flush.
- Updated all test suites to properly dispose AuthStorage instances in cleanup hooks to prevent resource leaks between tests.
Three early-return paths in #runAutoCompaction emitted auto_compaction_end
with result=undefined, aborted=false, and no errorMessage when compaction
was legitimately not needed (no model selected, no candidate models
available, or nothing to compact yet). The event-controller had no way to
distinguish these benign skips from a genuine failure, and fell through to
show the warning:
'Auto context-full maintenance failed; continuing without maintenance'
This was visible after a successful compaction: the next threshold check
would find nothing new to compact (prepareCompaction returns null), emit a
soft-skip event, and trigger the false warning.
Add skipped?: boolean to the auto_compaction_end event type and set it on
the three soft-skip paths. Update the event-controller to treat skipped
events as silent no-ops. Propagate the field through the extension and
hook AutoCompactionEndEvent interfaces so extensions can observe the
distinction.
- Simplified TruncationResult interface by making maxLines, maxBytes, and other derived fields optional, reducing redundancy.
- Refactored noTruncResult helper to auto-compute totalLines and totalBytes, eliminating repetitive parameter passing.
- Removed truncatedBy null checks and conditional formatting logic in truncation notice functions for cleaner output.
- Consolidated internal URL handling to use noTruncResult and removed unused displayMode variable.
- Clarified documentation in read.md to distinguish filesystem output from text output formatting.
- Changed eager todo reminder message role from 'developer' to 'custom' with customType field for better message categorization.
- Removed userRequest parameter from eager todo prelude generation to simplify prompt template rendering.
- Updated eager todo prompt to avoid redundant todo_write calls unless task state materially changed.
- Modified eager todo reminder message to use string content with display: false property instead of array format.
- Refactored eager todo injection from recursive prompt call to prepended message pattern.
- Removed state fields for todo injection tracking and consolidated logic into message composition.
- Added prependMessages option to promptWithMessage for composing messages before main prompt.
- Updated system prompt to require todo creation before substantive work on user requests.
- Fixed race condition where outer prompt would continue after user abort during eager todo's inner prompt by checking generation counter before proceeding.
- Added 'todo.eager' configuration setting to automatically create a comprehensive todo list after the first user message.
- Added 'buildNamedToolChoice' utility function to build provider-aware tool choice constraints for named tools.
- Modified tool choice resolution to support per-turn tool choice overrides via consumeNextToolChoiceOverride() method.
- Implemented eager todo enforcement mechanism that injects a synthetic prompt to encourage todo creation when conditions are met.
- Extracted tool choice building logic into reusable utility module for better code organization.
- Added comprehensive test coverage for eager todo enforcement functionality in AgentSession.
writeTerminalBreadcrumb used void Bun.write() — discarding the promise.
When parallel tasks write the same breadcrumb file concurrently on
Windows, Bun.write fails with EBUSY. The discarded promise rejects
unhandled, crashing the process.
Replace void with .catch(() => {}) to properly swallow best-effort
failures.
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Preserved text signature metadata (id and phase) when building OpenAI native history during session compaction.
- Added test case validating that codex assistant text signature metadata is correctly preserved in remote compaction history.
When the OpenAI responses stream stalls (e.g. with github-copilot/
gpt-5.4), the error message "stream stalled while waiting for the
next event" is now recognized as a retryable transient error.
Two locations updated:
- packages/ai/src/utils/retry.ts: add "stream stall" to
TRANSIENT_MESSAGE_PATTERN (used by provider-level retry logic)
- packages/coding-agent/src/session/agent-session.ts: add
"stream stall" to #isRetryableErrorMessage() regex (used by
agent-session retry loop for both thrown errors and error
AssistantMessages)
Fixes#348
Co-authored-by: GitHub User <user@example.com>
- Removed Kagi Universal Summarizer integration from fetch tool and YouTube scraper.
- Removed `fetch.useKagiSummarizer` configuration setting from settings schema.
- Simplified renderHtmlToText() and renderUrl() functions by removing Kagi summarization fallback logic.
- Fixed indentation inconsistencies in test files from tabs to spaces.
* fix(session): bypass user-prompt pipeline in handoff
handoff() was calling #promptWithMessage, which gates on an API key
check before reaching this.agent.prompt(). That gate is appropriate for
user-facing prompts but has no place in an internal document-generation
call: it blocked the test spy on agent.prompt and required callers to
carry real credentials just to run the handoff path.
Fix: call #promptAgentWithIdleRetry directly (preserving the
busy-wait behaviour and #promptInFlightCount tracking) and skip the
user-prompt pipeline (API key validation, bash/python flushes, file
mention expansion, plan messages, extension events) entirely. handoff
creates a fresh session immediately after, so none of that setup
applies.
Tests now reach agent.prompt with no stub on modelRegistry.getApiKey.
* fix(patch): HASHLINE_PREFIX_RE strips comment lines with word: pattern
The regex used [0-9a-zA-Z]{1,16} for the hash ID segment, which matched
common comment patterns like '# Note:', '# TODO:', '# FIXME:'. When a
single-line replacement contained such a comment, nonEmpty===1 and
hashPrefixCount===1, triggering stripping and eating the comment prefix.
Actual hashline IDs are always exactly 2 chars from ZPMQVRWSNKTXJBYH.
Constrain the regex to that exact alphabet so no English word can match.
Also update tests that used fake IDs (AB, CD, EF) not in the real alphabet.
* Revert "fix(patch): HASHLINE_PREFIX_RE strips comment lines with word: pattern"
This reverts commit 112ad083de956d4ed8b78a7e6e9af2c061befbd5.
---------
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
* fix: strip invalid thinking signatures from aborted/errored messages
When a stream is interrupted mid-response, thinking blocks may have
empty or partial cryptographic signatures. These get persisted to
session history and sent on the next API call, causing:
'Invalid signature in thinking block'
transformMessages() now detects aborted/errored assistant messages and
clears thinkingSignature fields so they are treated as unsigned thinking
(converted to text by the serializer).
Also protect truncateForPersistence from corrupting signatures — clear
them entirely instead of truncating, since a partial signature is always
invalid.
* fix: disable thinking when tool_choice forces tool use on Bedrock
Bedrock rejects requests that combine extended thinking with forced
tool_choice (any or specific tool). The Anthropic provider already had
a guard (disableThinkingIfToolChoiceForced) but the Bedrock provider
was missing the equivalent check.
Also fix thinking block serialization: when a thinking block has no
valid signature (e.g., from an aborted stream), convert it to plain
text instead of sending it as reasoningContent without a signature.
The API requires the signature field on all reasoning blocks for models
that support it.
Add thinking block diagnostics to error messages for signature/thinking
related failures to aid debugging.
- Added skipPostPromptRecoveryWait option to HandoffOptions for deferring recovery work in handoff operations.
- Added deferred auto-compaction scheduling for threshold-triggered handoffs via post-prompt task queue.
- Extracted handoff document template to dedicated system prompt file for improved maintainability and reusability.
- Changed handoff prompt generation to use template rendering with custom focus instructions support.
- Refactored prompt-in-flight tracking from boolean flag to counter for proper nested operation handling.
The context fullness gauge was driven by output token count, causing
erratic jumps between turns (e.g. 84% -> 64%) with no compaction.
Status bar and estimateContextTokens now use calculatePromptTokens()
which returns input + cacheRead + cacheWrite — the actual input context
size. Previously both used a formula that included the final output token
count, which fluctuates with response length and is not part of the
context window for the current request.
isContextOverflow's usage-based fallback (z.ai silent overflow) was
also missing cacheWrite (cache_creation_input_tokens). Per Anthropic
docs the threshold is input + cache_read + cache_creation — all three.
Ref: https://platform.claude.com/docs/en/about-claude/pricing#long-context-pricing
google.ts and google-vertex.ts were double-counting cached tokens.
Gemini's promptTokenCount already includes cachedContentTokenCount, so
assigning input = promptTokenCount and cacheRead = cachedContentTokenCount
overcounted by cachedContentTokenCount on every cached request. Fixed
by subtracting first, matching the OpenAI convention:
input = promptTokenCount - cachedContentTokenCount
cacheRead = cachedContentTokenCount
=> input + cacheRead = promptTokenCount (total prompt, no double-count)
Ref: https://ai.google.dev/api/generate-content#v1beta.GenerateContentResponse.UsageMetadata
All other providers validated: amazon-bedrock (inputTokens is uncached
by API contract), openai-completions/responses/azure (already subtract
cached), kimi/gitlab-duo (delegate to correct implementations), cursor
(API exposes output tokens only — input stays 0 by design).
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Added `disabledCause` parameter to credential deletion methods to track reason credentials are disabled.
- Changed credential disabling mechanism from boolean `disabled` flag to `disabled_cause` text field for better auditability.
- Fixed credential purging to respect disabled credentials during email deduplication operations.
- Refactored `replaceAuthCredentialsForProvider()` to update matching credentials instead of deleting all, preserving credential history.
- Added incremental history mode to OpenAI responses .
- Changed OpenAI Codex to exclusively use websockets v2 protocol with fatal error detection for automatic SSE fallback.
- Fixed Gemini model parsing to strip `-preview` suffix for consistent model identification across API calls.
- Improved websocket error handling to extract and report detailed error messages from error events.
- Removed deprecated BETA_RESPONSES_WEBSOCKETS constant and websocket v2 feature flag branching logic.
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Added serviceTier option to OpenAI providers for controlling processing priority and cost across agent, completions, responses, and codex APIs.
- Added providerPayload field to messages for transport-native history reconstruction in OpenAI Responses and Codex APIs.
- Added /fast slash command and serviceTier setting to coding-agent for toggling OpenAI priority mode with fast mode indicator.
- Added remote compaction support with encrypted reasoning preservation for OpenAI models in coding-agent.
- Removed usage caching layer across all providers and refactored UsageFetchContext to eliminate cache and now dependencies.
- Fixed OpenAI Codex streaming service_tier inclusion, provider retry logic with exponential backoff, and email-based credential deduplication.
* idiomatic rust fixes
* idiomatic rust fixes
* display an image if we are fetching an image
* MIME type strictness
* codex nagging me
* codex nagging
* handoff instead of compaction as context filled strategy and surfacing
* handoff instead of compaction as context filled strategy and surfacing p2
* handoff instead of compaction as context filled strategy and surfacing p3
* handoff instead of compaction as context filled strategy and surfacing p4
* handoff instead of compaction as context filled strategy and surfacing p5
* handoff instead of compaction as context filled strategy and surfacing, fixes
* failing fetch test from the fetch tool updates
* handoff focus prompt skeleton
* handoff focus prompt skeleton p2
* fetch bugs
* further codex improvements
* further codex improvements
---------
Co-authored-by: Brit <lol@no.com>
- Removed static thinking mode constant exports (THINKING_LEVELS, ALL_THINKING_LEVELS, ALL_THINKING_MODES, THINKING_MODE_DESCRIPTIONS, THINKING_MODE_LABELS) in favor of dynamic function-based API.
- Renamed formatThinking() to getThinkingMetadata() with return type changed from string to structured ThinkingMetadata object containing value, label, and description.
- Renamed getAvailableThinkingLevel() to getAvailableThinkingLevels() and getAvailableThinkingEffort() to getAvailableThinkingEfforts() with added default parameters for runtime flexibility.
- Updated all consumer modules to use new function-based API instead of static constants, enabling dynamic thinking mode configuration.
- Fixed provider session state not being cleared when branching or navigating tree history, preventing resource leaks with codex provider sessions.
- Added calls to `#closeCodexProviderSessionsForHistoryRewrite()` in branch and navigateTree methods to ensure proper cleanup.
- Added test coverage for provider session cleanup during history branching and tree navigation.
- Updated option interfaces, param builders, and stream functions to use unified `reasoning` field.
- Added `resolveOpenAiReasoningEffort()` to centralize xhigh clamping logic.
- Replaced type casts with `castApi()` helper and fixed test option names.
- Updated coding-agent and benchmark references for consistency.
- Extracted thinking module with ThinkingEffort, ThinkingLevel, and ThinkingMode types to centralize reasoning configuration across packages.
- Migrated ThinkingLevel type from pi-agent-core to pi-ai package with new validation functions parseThinkingLevel() and getAvailableThinkingLevel().
- Consolidated thinking level constants and descriptions into reusable exports (ALL_THINKING_LEVELS, THINKING_MODE_DESCRIPTIONS) for consistent UI display.
- Removed local thinking-effort-label utility and replaced formatThinkingEffortLabel() with centralized formatThinking() function from pi-ai.
- Refactored thinking mode handling to distinguish ThinkingSelector (user-facing with 'off' option) from ThinkingEffort (provider-level).
parseModelString now extracts valid thinking level suffixes (e.g.,
"anthropic/claude-opus-4-6:high") instead of treating them as part of
the model ID. This enables per-role thinking levels in config:
modelRoles:
slow: anthropic/claude-opus-4-6:high
default: anthropic/claude-opus-4-6:low
smol: google/gemini-3-flash:medium
The thinking level is applied at startup, in SDK fallback resolution,
and during Ctrl+P role cycling. The original config string is preserved
on role cycle so the suffix round-trips correctly.
* feat(mcp): resource notifications, subscriptions, and read_resource builtin tool
- Add MCP resource subscription lifecycle (subscribe/unsubscribe on connect/disconnect)
- Wire mcp.notifications setting with live toggle support
- Add debounced followUp injection for resource change notifications
- Add global read_resource builtin tool with server resolution by URI/template scheme
- Add MCP prompt commands (buildMCPPromptCommands) with array content support
- Add server instructions injection into system prompt with attribution
- Add mcp.notificationDebounceMs configurable setting
Client (client.ts):
listResources, listResourceTemplates, readResource with pagination
subscribeToResources, unsubscribeFromResources
listPrompts, getPrompt, serverSupportsPrompts
serverSupportsResources, serverSupportsResourceSubscriptions
Manager (manager.ts):
Notification dispatch with subscribed-URI guard
Concurrent refresh deduplication via pending promise map
setNotificationsEnabled with subscribe/unsubscribe toggle
Tests:
client-resources.test.ts (31 tests)
client-prompts.test.ts (20 tests)
mcp-read-resource.test.ts (13 tests)
* fix(mcp): address PR review - eager prompt init and stale subscription cleanup
P1: Make setOnPromptsChanged eagerly fire for servers that already
have prompts loaded. The callback is registered after MCP discovery
has already loaded prompts and fired the hook, so without this the
handler is never called on the common startup path. The fix is in
the manager itself (not the caller), eliminating the race condition
regardless of when the callback is wired.
P2: Unsubscribe removed resource URIs on resource refresh.
refreshServerResources was subscribing to the new URI set and
overwriting #subscribedResources without unsubscribing URIs that
were previously subscribed but no longer present, leaving stale
subscriptions active on the server.
* fix(mcp): add resources and prompts to /mcp help text and subcommand completions
* feat(mcp): add /mcp notifications command
Shows per-server notification capabilities with subscription state:
- Lists supported notification types (tools/list_changed, resources/list_changed,
prompts/list_changed) with check marks
- Shows resources/subscribe status with active subscription count
- Lists subscribed URIs with green ticks when notifications are enabled
- Displays overall enabled/disabled state (mcp.notifications setting)
* fix(mcp): address PR review comments on race conditions and stale state
- Await subscribe/unsubscribe in refreshServerResources so the refresh
promise doesn't resolve before subscriptions are settled, preventing
a second refresh from racing and overwriting tracking state (P2 #3)
- Guard setNotificationsEnabled subscribe .then() against a disable
that happens while the subscribe request is in-flight (P2 #5)
- Re-check mcp.notifications setting inside debounce setTimeout
callback so toggling off mid-window actually suppresses the
follow-up message (P2 #4)
- Fire onToolsChanged and onPromptsChanged callbacks in
disconnectServer so stale slash commands and tool registrations
are cleaned up when a server is removed (P2 #2)
---------
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Add `rate-limit-utils.ts` to classify 429/503 errors (Quota, Rate Limit, Capacity)
- Fix `google-gemini-cli` to fail-fast on 429s instead of getting stuck in internal retries
- Update `isUsageLimitErrorMessage` regex in `AgentSession` to catch all Google-specific error variations
- Implement smart backoff timings (30m for quota, 30s for rate limit, 45s+jitter for capacity)
- Remove 0% hiding logic in `/usage` to always display account rotation pool
- Add comprehensive unit tests for rate limit parsing and provider behavior
- Enforced tool decision in plan mode--agent now requires calling either `ask` or `exit_plan_mode` when a turn ends without a required tool call.
- Fixed cancellation behavior of `ask` tool to abort the current turn instead of returning a normal cancelled selection, while timeout-driven auto-cancel still returns without aborting.
- Added plan-mode-tool-decision-reminder system prompt to guide agent when required tools are not called.
- Improved agent_end event handling to use fallback assistant message when #lastAssistantMessage is unavailable.
- Added checkpoint and rewind tools to create context checkpoints before exploratory work and rewind to replace exploration messages with concise reports.
- Added checkpoint.enabled setting to control availability of checkpoint and rewind tools in agent sessions.
- Added getCheckpointState() and setCheckpointState() methods to agent session API for checkpoint state management.
- Implemented checkpoint state tracking with message count, entry ID, and timestamp to enable context cost optimization during investigations.