- Migrated test mocking from bun:test mock API to vitest spyOn pattern across 9 test files.
- Extracted mock setup logic into reusable helper functions with Symbol.dispose cleanup pattern.
- Replaced manual beforeEach/afterEach and try-finally blocks with TypeScript 5.2 using declarations.
- Removed 178 lines of boilerplate mock initialization and restoration code from test suite.
- Simplified null/empty checks across TypeScript codebase using optional chaining operator (?.) for improved readability.
- Replaced explicit null checks in validation logic with optional chaining in oauth-discovery, gemini-cli, claude, zai, and lsp modules.
- Updated error handling in Rust command invocation to use double question mark operator (??) for cmd_result.
- Consolidated null validation patterns across tools (bash-skill-urls, browser, gemini-image, resolve) and keybindings using optional chaining.
- Added contract system for validating benchmark commands, metrics, scope paths, constraints, and off-limits paths.
- Contract validation enforces matching initialization parameters against autoresearch.md before init_experiment.
- Segment fingerprinting detects configuration drift and warns when metrics are not directly comparable.
- Added pending run detection and recovery to resume incomplete experiments from .autoresearch/runs/.
- Run directories organize artifacts with benchmark logs and optional checks logs for traceability.
- Extended experiment state to track run number, command, scope, off-limits, constraints, and fingerprint.
- Restructured hashline edit schema from flat op/pos/end/lines fields to nested loc/content objects with discriminated union types.
- Replaced operation names (replace_line, replace_range, append_at, prepend_at) with unified loc object patterns supporting line, block, append, and prepend anchors.
- Updated content field to accept array of strings or null instead of lines parameter in edit entries.
- Refactored edit resolution logic to dispatch on loc shape instead of op enum, simplifying location handling.
- Added boundary duplication warning to detect off-by-one range errors in replace_range and replace_line operations.
- Updated hashline tool documentation with boundary duplication trap guidance to prevent closing delimiter duplication.
- Added git branch isolation for autoresearch sessions with automatic branch creation, reuse, and worktree safety checks.
- Added scope definition sections (Files in Scope, Off Limits, Constraints) to autoresearch template for explicit session boundaries.
- Added keybinding matcher utilities for consistent escape/cancel key handling across interactive components.
- Added ASI metadata validation (hypothesis and rollback context) in log_experiment tool for experiment tracking.
- Refactored keybinding logic across 13 components to use centralized matcher functions instead of inline key checks.
- Renamed hashline operation types for clarity: append->append_at, prepend->prepend_at, append_eof->append_file, prepend_bof->prepend_file.
- Updated all operation type references in patch implementation, tests, and documentation to reflect new naming convention.
- Restructured hashline tool documentation with hierarchical sections and simplified examples for improved clarity.
- Consolidated validation rules and added explicit warning about invalid anchors and operation/field combinations.
- Added `defaultInactive` property to ToolDefinition for conditional tool registration and activation control.
- Added dynamic tool activation/deactivation API `setActiveTools()` for managing experiment tools in autoresearch mode.
- Replaced single `command-start.md` workflow with separate `command-initialize.md` and `command-resume.md` prompts for autoresearch initialization and session resumption.
- Added interactive intent dialog for autoresearch optimization goals with automatic session resumption detection based on autoresearch.md presence.
- Refactored autoresearch command handler to distinguish resume vs initialize flows and dynamically activate/deactivate experiment tools based on mode and session state.
- Added retry mechanism for benchmark tasks with separate system and retry prompt templates to improve edit success rates.
- Introduced autocorrect tracking metrics including autocorrect-free success rate and edit autocorrect counts in task and benchmark summaries.
- Refactored prompt building into modular functions (buildBenchmarkSystemPrompt, buildInitialBenchmarkPrompt, buildRetryBenchmarkPrompt) with BenchmarkPromptDelivery type for distinguishing initial and follow-up messages.
- Added session management with cache-keyed provider session IDs using xxHash64 and centralized RPC argument building via prepareBenchmarkSessionSetup.
- Refactored hashline edit operations from implicit `replace` to explicit `replace_line` and `replace_range` with mandatory `end` parameter for ranges.
- Added file-level operations `append_eof` and `prepend_bof` for boundary insertions, making `pos` required for anchor-based operations.
- Enforced stricter anchor validation per operation type and simplified edit application logic by separating file-level from anchor-based operations.
- Updated hashline edit application to preserve duplicated boundary lines without auto-correction, changing previous behavior.
- Added autoresearch extension with autonomous experiment loop supporting init, run, and log experiment tools for metric-driven optimization.
- Added widget placement system enabling extensions to position UI components above or below the editor via ExtensionWidgetOptions.
- Added dashboard controller with interactive overlay for viewing experiment results, metrics, and progress with keyboard navigation.
- Removed auto-correction logic for off-by-one range edits in hashline editor to preserve user intent in patch operations.
- Added state reconstruction utilities to parse autoresearch.jsonl logs and rebuild experiment state across sessions.
- Added comprehensive type definitions and helper utilities for metric parsing, ASI validation, and process management.
- Added createTestToolContext helper to construct AgentToolContext instances for tool tests.
- Updated 2 test cases to use createTestToolContext instead of inline object literals for consistency.
- Added imports for AgentToolContext and SessionManager to support new test helper.
- Changed bash interceptor configuration from boolean flags to customizable pattern-based rules array.
- Clarified hashline range replace semantics: end parameter is now strictly exclusive boundary.
- Fixed bash interceptor to apply built-in default rules when no custom patterns are configured.
- Updated hashline range validation and calculations to enforce exclusive end semantics throughout.
- Added renderInlineMarkdown() utility function to support inline markdown rendering with optional base color styling.
- Refactored ask tool to render questions and option labels with markdown formatting for improved text styling.
- Updated hook-input and hook-selector components to render titles as markdown with theme-aware styling.
- Implemented recursive token processing for nested markdown elements including bold, italic, code, links, and strikethrough.
Fixes#491
- Fixed rate-limit-utils to recognize and classify 'usage limit' errors as QUOTA_EXHAUSTED instead of transient.
- Added isUsageLimitError() utility function for unified detection of persistent quota limit errors across providers.
- Fixed Codex provider to return immediately on usage-limit errors instead of retrying, preventing unnecessary 5-minute delays.
- Removed usage.?limit pattern from TRANSIENT_MESSAGE_PATTERN to prevent misclassification of persistent quota errors.
- Added ACP (Agent Client Protocol) mode for headless agent operation via --mode acp flag.
- Integrated Agent Client Protocol SDK with session management, streaming communication, and event mapping.
- Added ensureOnDisk() method to SessionManager for immediate session persistence without requiring assistant messages.
- Changed session persistence to use atomic file rewrite for unflushed sessions.
- Implemented AcpAgent class with session management, prompt handling, MCP server configuration, and event streaming.
- Fixed thinking configuration format by replacing `levels` array with `minLevel`/`maxLevel` properties across 100+ model definitions.
- Corrected GPT-5.4 mini/nano context window from 400000 to 272000 tokens for accurate token limit reporting.
- Normalized GPT-5.4 variant priority handling to use parsed variant instead of raw model IDs for consistent behavior.
- Added "mini" variant support to OpenAI model parsing regex and updated thinking mode configuration for Claude models.
- Fixed test robustness by replacing exact string matching with numeric range comparison to handle BSD seq notation on macOS.
- Corrected model generation script execution order to apply policy overrides before promotion target linking.
- Restructured model policy application logic to improve readability and reduce nesting.
- Extracted model override application into dedicated function calls for consistency.
- Added auto-reconnect capability for MCP servers with SSE stream monitoring and exponential retry backoff.
- Added tool-level reconnect handling for retriable connection errors (ECONNREFUSED, ECONNRESET, 404/502/503).
- Added `/mcp reconnect <name>` command for manual MCP server recovery.
- Improved reconnect robustness by aborting retries when MCP configuration changes via epoch checking.
- Extended transport reconnect handling to all transport types (stdio, HTTP/SSE) with unified onClose logic.
- Added comprehensive test coverage for MCPManager reconnect behavior and tool-level abort propagation.
- Replaced custom withAbort() implementation with untilAborted() utility from pi-utils for consistent abort handling.
- Extracted normalizeToolArgs() helper to consolidate argument normalization logic across renderCall(), renderResult(), and execute() methods.
- Added MCPToolCallParams type import to improve type safety for tool parameter handling.
- Fixed STT Alt+H mic cursor rendering to measure the actual microphone glyph width, preventing one-column TUI overflow crashes when the active symbol preset uses a wide icon (#484).
- Extracted mic cursor styling into dedicated method to consolidate color and width calculation logic.
Fixes#484
- Refactored MCP manager reconnection logic to track configuration changes via reconnectEpoch parameter.
- Consolidated error pattern matching in tool-bridge to use lowercase normalization for consistent comparison.
- Extracted retry provider variables to eliminate repeated ternary expressions in error handling.
- Reformatted code across multiple files for improved readability with consistent multi-line formatting.
- Added test coverage for transport reconnection scenarios including connection reuse after reconnection.
- Updated explore agent thinking level from off to med for improved reasoning.
- Simplified explore agent output schema: consolidated file references into single `ref` field with optional line ranges instead of separate `path`, `line_start`, `line_end` fields.
- Removed `code` section from explore agent output (critical code excerpts no longer extracted).
- Removed `dependencies`, `risks`, and `start_here` sections from explore agent output.
The 8e3e0ebf9 refactor centralized thinking inference but dropped
Bedrock-specific handling:
- inferAnthropicSupportedEfforts gated xhigh on anthropic-messages API,
excluding Bedrock models
- inferThinkingControlMode returned 'budget' for all Bedrock models
instead of 'anthropic-adaptive' for 4.6+
- parseAnthropicModel regex captured Bedrock date suffixes (20250819)
as version components, causing parseSemVer to return null
Fixes:
- Remove API guard from inferAnthropicSupportedEfforts so Opus 4.6+
gets xhigh regardless of API transport
- Add Anthropic version checks to Bedrock case in
inferThinkingControlMode (4.6+ adaptive, 4.5+ budget-effort)
- Narrow version regex from \d+ to \d{1,2} so date suffixes in
Bedrock model IDs are not absorbed
- Update four stale Bedrock Opus 4.6 entries in models.json
- Add thinkingBudgets.xhigh (default 32768) to settings schema
* feat: auto-reconnect MCP servers on connection loss
When an HTTP SSE stream drops (server restart, network interruption),
the transport fires onClose, and the manager proactively reconnects
with retry backoff (500ms, 1s, 2s, 4s). Tools are kept in the registry
during reconnection so they remain selected and available to the agent.
If proactive reconnection fails, stale tools stay registered. When the
agent calls one, the tool bridge detects the retriable connection error
(ECONNREFUSED, ECONNRESET, stale session 404/502/503, etc.), triggers
reconnectServer on the manager, and retries the call once on the fresh
connection. Concurrent reconnect attempts for the same server are deduped.
connectToServer now always installs a default onRequest handler for ping
and roots/list (using getProjectDir()), so all connections -- including
short-lived test/probe ones -- properly respond to server-initiated
requests during initialization.
Post-connection setup (resources, prompts, subscriptions) is extracted
into a shared #loadServerResourcesAndPrompts method used by both initial
connection and reconnection paths.
Add /mcp reconnect <name> command for manual recovery after extended
outages where both proactive and reactive reconnection have failed.
* docs: add changelog entry for MCP auto-reconnect
* fix: address P1 review findings in MCP reconnection
- Save server configs before connection attempt so deferred tools can
reconnect even when the initial connection timed out (P1-1)
- Make waitForConnection() and getConnectionStatus() aware of in-flight
reconnections so callers wait instead of failing immediately (P1-2)
- Add epoch counter incremented on disconnectAll() and checked in
connectAndWireServer() to invalidate stale reconnect attempts that
outlive a manager reset/reload (P1-3)
- Skip servers with pending reconnections in connectServers() to prevent
parallel connection attempts for the same server
* fix: deferred tool reconnect and non-blocking transport teardown
- DeferredMCPTool.execute now reconnects when getConnection() fails
("MCP server not connected"), not only on network errors from
callTool. Servers that missed the startup window can now be woken
by the first tool call against their cached tools. (P1-4)
- #doReconnect fire-and-forgets the old transport close instead of
awaiting it. HttpTransport.close() sends a DELETE with 30s timeout;
blocking here delayed the first reconnect attempt by that amount
on every server restart. (P1-5)
* fix: abort-aware reconnect waits and preserve tool selection on reconnect
- Wrap all reconnect() awaits with withAbort(signal) so user
cancellation (Esc) interrupts the reconnect backoff loop instead
of blocking for up to 7.5s. Applies to MCPTool (1 site) and
DeferredMCPTool (2 sites). (P2-1)
- Remove activateDiscoveredMCPTools call from /mcp reconnect handler.
refreshMCPTools already preserves the user's prior MCP tool
selection; the extra activation was silently opting into all
server tools including ones the user had not enabled. (P2-2)
* fix: rebind MCPTool connection after reconnect, add stdio retriable error
- MCPTool.connection is now mutable; after a successful reconnect retry,
this.connection is rebound to the fresh connection so subsequent calls
on the same instance (e.g. batched tool calls) use it instead of
triggering another reconnect cycle. (P2-3)
- Add "Transport closed" to RETRIABLE_PATTERNS. StdioTransport rejects
pending requests with this message when the subprocess dies, which
should trigger the reconnect path just like HTTP transport errors. (P2-4)
---------
Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
Apply user modelOverrides after hardcoded defaults so contextWindow
and other fields from models.json take precedence.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Simplified plan mode enforcement by removing tool retrieval and restoration logic.
- Replaced tool existence checks with registry lookup to reduce intermediate variables.