- Preserved text signature metadata (id and phase) when building OpenAI native history during session compaction.
- Added test case validating that codex assistant text signature metadata is correctly preserved in remote compaction history.
- Extracted OpenAI Responses API stream processing logic into shared module with 424 lines of reusable utilities.
- Consolidated text signature encoding, tool call normalization, and message conversion functions into openai-responses-shared module.
- Refactored openai-responses.ts and azure-openai-responses.ts to delegate stream processing to shared processResponsesStream() helper.
- Removed 473 lines of duplicated stream event handling and utility functions across OpenAI provider implementations.
- Added `onPayload` callback option to intercept and transform provider request payloads before transmission across agent and AI packages.
- Added structured text signature metadata with phase information to OpenAI and Azure OpenAI providers for enhanced response tracking.
- Added `before_provider_request` extension event to coding-agent for chaining payload transformations across multiple handlers.
- Improved error messages in `response.failed` events with detailed error codes, messages, and incomplete reasons from provider responses.
When the OpenAI responses stream stalls (e.g. with github-copilot/
gpt-5.4), the error message "stream stalled while waiting for the
next event" is now recognized as a retryable transient error.
Two locations updated:
- packages/ai/src/utils/retry.ts: add "stream stall" to
TRANSIENT_MESSAGE_PATTERN (used by provider-level retry logic)
- packages/coding-agent/src/session/agent-session.ts: add
"stream stall" to #isRetryableErrorMessage() regex (used by
agent-session retry loop for both thrown errors and error
AssistantMessages)
Fixes#348
Co-authored-by: GitHub User <user@example.com>
- Added `modelId` parameter to `getApiKey()` for model-specific credential selection.
- Implemented plan-based account prioritization for OpenAI Codex Pro models, preferring Pro plan accounts.
- Added `AuthApiKeyOptions` type to support model-specific authentication configuration.
- Added helper functions for OpenAI Codex Pro model detection and plan-based ranking logic.
- Added background model discovery with provider status tracking and 24-hour model caching to improve startup performance.
- Changed model discovery timeout from 3000ms to 250ms and deferred blocking refresh to background operations for faster initialization.
- Fixed model discovery to preserve cached models when providers are unavailable or unauthenticated.
- Added validation in Kitty key formatting to reject unsupported modifiers and improved error handling.
- Reorganized CI workflow to run TypeScript linting in check job and conditional Rust checks in native job matrix.
- Updated system prompt and tool guidance to recommend combined investigation-and-edit workflows and maintainable code practices.
- Removed provider parameter from web search tool schema; provider selection now handled internally.
- Removed deprecated no_fallback option from search parameters; fallback behavior is now automatic.
- Renamed SearchParams type to SearchToolParams and introduced SearchQueryParams for CLI queries.
- Updated executeSearch() to accept SearchQueryParams with optional provider selection.
- Introduced SubmittedUserInput type to track submission state including cancelled and started flags. Added startPendingSubmission, cancelPendingSubmission, markPendingSubmissionStarted, and finishPendingSubmission methods to manage the lifecycle of user input submissions. Enhanced escape key handling to prioritize canceling pending optimistic submissions before aborting active sessions.
- Extracted loading animation lifecycle management into centralized ensureLoadingAnimation() method.
- Consolidated duplicate loading animation setup across event-controller and interactive-mode into single reusable method.
- Updated showError() to properly clean up loading animation state when errors occur.
- Added comprehensive test coverage for InputController escape key behavior with optimistic submission.
- Imported buildSessionContext from session-manager module.
- Replaced inline session context object with buildSessionContext factory call in test assertion.
The manual /compact command path in executeCompaction() called
rebuildChatFromMessages() but never added the compactionSummary
message to the chat. The auto-compaction path in event-controller
explicitly adds it, but the manual path was missing this step.
Capture the CompactionResult returned by session.compact() and
render the summary via addMessageToChat(), matching the behavior
of the auto-compaction flow.
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Reapplied #applyHardcodedModelPolicies method to enforce gpt-5.4 context window policy across all model loading paths.
- Updated model registry to apply hardcoded policies after both initial load and dynamic discovery.
- Adjusted test expectations to reflect the restored policy enforcement behavior.
- Added Tavily web search provider with API key authentication and credential discovery from environment or database.
- Integrated Tavily as highest-priority search provider in fallback chain with structured response mapping and error handling.
- Added Tavily OAuth login flow in CLI and auth-storage with manual API key input and validation.
- Added comprehensive test suite for Tavily provider covering registration, response mapping, error handling, and credential validation.
Fixes#313
- Removed Kagi Universal Summarizer integration from fetch tool and YouTube scraper.
- Removed `fetch.useKagiSummarizer` configuration setting from settings schema.
- Simplified renderHtmlToText() and renderUrl() functions by removing Kagi summarization fallback logic.
- Fixed indentation inconsistencies in test files from tabs to spaces.
findAnthropicAuth() only checked OAuth credentials in agent.db, missing
api_key type credentials entirely. Users authenticated via stored API key
(not env var, not OAuth) got null from isAvailable(), causing the provider
chain to skip Anthropic.
Added tier 4 (api_key in agent.db) between OAuth check and env var
fallback. Refactored store lifecycle so tiers 3-4 share one instance.
Also fixed ExaProvider.isAvailable() which unconditionally returned true,
ignoring exa.enabled/exa.enableSearch settings and never checking for
credentials. It now respects both settings and requires an actual API key.
* fix(session): bypass user-prompt pipeline in handoff
handoff() was calling #promptWithMessage, which gates on an API key
check before reaching this.agent.prompt(). That gate is appropriate for
user-facing prompts but has no place in an internal document-generation
call: it blocked the test spy on agent.prompt and required callers to
carry real credentials just to run the handoff path.
Fix: call #promptAgentWithIdleRetry directly (preserving the
busy-wait behaviour and #promptInFlightCount tracking) and skip the
user-prompt pipeline (API key validation, bash/python flushes, file
mention expansion, plan messages, extension events) entirely. handoff
creates a fresh session immediately after, so none of that setup
applies.
Tests now reach agent.prompt with no stub on modelRegistry.getApiKey.
* fix(patch): HASHLINE_PREFIX_RE strips comment lines with word: pattern
The regex used [0-9a-zA-Z]{1,16} for the hash ID segment, which matched
common comment patterns like '# Note:', '# TODO:', '# FIXME:'. When a
single-line replacement contained such a comment, nonEmpty===1 and
hashPrefixCount===1, triggering stripping and eating the comment prefix.
Actual hashline IDs are always exactly 2 chars from ZPMQVRWSNKTXJBYH.
Constrain the regex to that exact alphabet so no English word can match.
Also update tests that used fake IDs (AB, CD, EF) not in the real alphabet.
* Revert "fix(patch): HASHLINE_PREFIX_RE strips comment lines with word: pattern"
This reverts commit 112ad083de956d4ed8b78a7e6e9af2c061befbd5.
---------
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Resolved symlinked paths before passing to brush shell to keep `pwd` output aligned with canonical Git worktree paths.
- Added `resolveShellCwd` helper that safely resolves symlinks and falls back to original path on error.
- Added test case verifying symlinked directories are canonicalized before execution.
* fix: strip invalid thinking signatures from aborted/errored messages
When a stream is interrupted mid-response, thinking blocks may have
empty or partial cryptographic signatures. These get persisted to
session history and sent on the next API call, causing:
'Invalid signature in thinking block'
transformMessages() now detects aborted/errored assistant messages and
clears thinkingSignature fields so they are treated as unsigned thinking
(converted to text by the serializer).
Also protect truncateForPersistence from corrupting signatures — clear
them entirely instead of truncating, since a partial signature is always
invalid.
* fix: disable thinking when tool_choice forces tool use on Bedrock
Bedrock rejects requests that combine extended thinking with forced
tool_choice (any or specific tool). The Anthropic provider already had
a guard (disableThinkingIfToolChoiceForced) but the Bedrock provider
was missing the equivalent check.
Also fix thinking block serialization: when a thinking block has no
valid signature (e.g., from an aborted stream), convert it to plain
text instead of sending it as reasoningContent without a signature.
The API requires the signature field on all reasoning blocks for models
that support it.
Add thinking block diagnostics to error messages for signature/thinking
related failures to aid debugging.
* fix(patch): HASHLINE_PREFIX_RE strips comment lines with word: pattern
The regex used [0-9a-zA-Z]{1,16} for the hash ID segment, which matched
common comment patterns like '# Note:', '# TODO:', '# FIXME:'. When a
single-line replacement contained such a comment, nonEmpty===1 and
hashPrefixCount===1, triggering stripping and eating the comment prefix.
Actual hashline IDs are always exactly 2 chars from ZPMQVRWSNKTXJBYH.
Constrain the regex to that exact alphabet so no English word can match.
Also update tests that used fake IDs (AB, CD, EF) not in the real alphabet.
* test(patch): add regression tests for comment line prefix stripping bug
Three new tests in hashlineParseContent describe block:
- hashlineParseText preserves '# Word:' comment lines (unit)
- full pipeline: replacing '# Note:' comment line preserves prefix
- full pipeline: replacing '# TODO:' comment line preserves prefix
These would have caught the HASHLINE_PREFIX_RE bug where [0-9a-zA-Z]{1,16}
matched comment words, causing stripNewLinePrefixes to eat the '# Note:'
prefix when a single comment line was the sole replacement entry.
---------
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
#applyHardcodedModelPolicies forced contextWindow=1_000_000 on any model
with id 'gpt-5.4' regardless of provider. This was wrong in every case:
- github-copilot/gpt-5.4: inflated from bundled 400K (and overwrote the
~274K that Copilot's live /models discovery returns via capabilities.limits)
- openai/gpt-5.4: deflated from bundled 1_050_000 to 1_000_000
- openai-codex, opencode, opencode-zen: same downgrade from 1_050_000
The method was called twice — after static load and after runtime discovery —
so it reliably clobbered the correct provider-specific value both times.
No documented rationale exists for the override. The bundled models.json
values are already correct. Remove the method and its two call sites.
fixes#332
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
- Extracted OAuth identifier logic into public functions extractOAuthCredentialIdentifiers and extractOAuthTokenIdentifiers.
- Replaced single credentialIdentity string with multi-identifier resolveCredentialIdentifiers returning string[] for flexible matching.
- Changed credential deduplication from email-based to accountId-based matching in replaceAuthCredentialsForProvider.
- Updated auth-storage tests to verify accountId-prioritized deduplication behavior across soft-disable and hard-delete scenarios.
- Added documentation comments in coding-agent modules explaining partial JSON preservation for streaming tool previews.
- Documented streaming tool preview requirements and render paths in AGENTS.md.
- Added `env` parameter to bash tool for safe environment variable passing without shell re-parsing.
- Added support for rendering partial environment variable assignments in command preview during streaming.
- Updated bash tool prompt to recommend `env` parameter for multiline, quote-heavy, and untrusted values.
- Refactored tool execution component to conditionally merge partial JSON arguments during streaming.
- Added helper functions for environment variable normalization, escaping, and formatting.
- Fixed WebSocket stream fallback logic to safely replay buffered output over SSE when WebSocket fails after partial content has been streamed.
- Added tracking flag to prevent unsafe replays of tool calls and terminal events during fallback transitions.
- Enhanced error recovery to reset output state when replaying buffered content over SSE connection.
- Added docs.rs scraper for extracting Rust crate documentation from rustdoc JSON, supporting modules, functions, structs, traits, enums, and other Rust items with intelligent caching.
- Implemented rustdoc JSON parsing with type rendering for complex Rust types including generics, lifetimes, trait bounds, and qualified paths.
- Added caching layer for rustdoc JSON with date-based versioning for 'latest' releases to reduce repeated fetches.
- Added fallback search strategies in librarian and explore agent prompts for handling empty results.
- Clarified task completion priority in system prompt by prohibiting premature tool call cessation.
- Consolidated and simplified 'Giving Up' guidance in subagent prompt with clearer uncertainty handling.
- Removed duplicate instructions and redundant phrasing to improve prompt clarity and conciseness.
- Added skipPostPromptRecoveryWait option to HandoffOptions for deferring recovery work in handoff operations.
- Added deferred auto-compaction scheduling for threshold-triggered handoffs via post-prompt task queue.
- Extracted handoff document template to dedicated system prompt file for improved maintainability and reusability.
- Changed handoff prompt generation to use template rendering with custom focus instructions support.
- Refactored prompt-in-flight tracking from boolean flag to counter for proper nested operation handling.
- Moved llms.txt endpoint discovery to fallback strategy when rendered page content is low quality, prioritizing page-specific content over site-wide files.
- Enhanced llms.txt endpoint detection to scope candidates to the requested URL path, searching section-specific files before site-wide ones.
- Replaced getOrigin() with buildLlmEndpointCandidates() to generate path-scoped endpoint candidates with depth-based fallback strategy.
- Updated tryLlmEndpoints() to accept full URL and return endpoint metadata alongside content for better fallback tracking.
- Added 2 integration tests validating section-scoped llms.txt discovery and preference for rendered content over site-wide files.
- Clarified contextual pattern mode behavior in ast-grep and ast-edit documentation to explain that results target the selected node, not the outer wrapper.
- Enhanced TypeScript pattern examples to include class method matching syntax with `class $_ { method(...) }` wrapper pattern.
- Removed redundant class method examples that were superseded by improved documentation of contextual pattern mode.
- Renamed parameter names in ast-grep and ast-edit tools from `patterns`/`selector` to `pat`/`sel` for brevity across schema, implementation, and tests.
- Expanded ast-grep and ast-edit tool documentation with 12+ new usage guidelines, examples, and critical notes on pattern syntax, metavariable placement, and error handling.
- Updated CHANGELOG.md to document parameter renames and expanded tool guidance for AST pattern syntax and metavariable usage.
- Reformatted test assertions and type annotations across ast-edit and ast-grep test files for improved readability.
- Added `glob` parameter to `ast_grep`, `ast_edit`, and `grep` tools for filtering files relative to `path`.
- Implemented `combineSearchGlobs()` utility to merge glob patterns from multiple sources instead of throwing errors.
- Changed `grep` tool to combine glob patterns when both `path` and `glob` parameters are provided.
- Updated tool documentation to recommend pairing `path`, `glob`, and `lang` for language-scoped search in mixed repositories.
- Added comprehensive test coverage for combined path and glob parameter handling across grep, ast_grep, and ast_edit tools.
- Added automatic Ollama model capability detection via /api/show endpoint to discover reasoning and input modality support.
- Improved Kagi API error handling with structured error parsing for JSON and plain text response formats.
- Fixed Cerebras streaming compatibility by omitting stream_options.include_usage parameter.
- Simplified API key credential storage to always replace credentials instead of merging for non-minimax providers.
- Updated Kagi Search API key format from 'kagi_...' to 'KG_...' and clarified beta access requirement in provider description.
Fixes#326.
Fixes#321.
Fixes#298.
The context fullness gauge was driven by output token count, causing
erratic jumps between turns (e.g. 84% -> 64%) with no compaction.
Status bar and estimateContextTokens now use calculatePromptTokens()
which returns input + cacheRead + cacheWrite — the actual input context
size. Previously both used a formula that included the final output token
count, which fluctuates with response length and is not part of the
context window for the current request.
isContextOverflow's usage-based fallback (z.ai silent overflow) was
also missing cacheWrite (cache_creation_input_tokens). Per Anthropic
docs the threshold is input + cache_read + cache_creation — all three.
Ref: https://platform.claude.com/docs/en/about-claude/pricing#long-context-pricing
google.ts and google-vertex.ts were double-counting cached tokens.
Gemini's promptTokenCount already includes cachedContentTokenCount, so
assigning input = promptTokenCount and cacheRead = cachedContentTokenCount
overcounted by cachedContentTokenCount on every cached request. Fixed
by subtracting first, matching the OpenAI convention:
input = promptTokenCount - cachedContentTokenCount
cacheRead = cachedContentTokenCount
=> input + cacheRead = promptTokenCount (total prompt, no double-count)
Ref: https://ai.google.dev/api/generate-content#v1beta.GenerateContentResponse.UsageMetadata
All other providers validated: amazon-bedrock (inputTokens is uncached
by API contract), openai-completions/responses/azure (already subtract
cached), kimi/gitlab-duo (delegate to correct implementations), cursor
(API exposes output tokens only — input stays 0 by design).
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>