Commit Graph
2739 Commits
Author SHA1 Message Date
can1357 df369dd7fd refactor(coding-agent): migrated test mocks to vitest spyOn with Symbol.dispose cleanup
- Migrated test mocking from bun:test mock API to vitest spyOn pattern across 9 test files.
- Extracted mock setup logic into reusable helper functions with Symbol.dispose cleanup pattern.
- Replaced manual beforeEach/afterEach and try-finally blocks with TypeScript 5.2 using declarations.
- Removed 178 lines of boilerplate mock initialization and restoration code from test suite.
2026-03-26 16:17:27 +01:00
can1357 4ff0f309a1 chore: bump models.json 2026-03-26 15:33:58 +01:00
daanddenandGitHub 6ea9504b03 Make temporary model selector keybinding configurable (#539)
Fixes #533
2026-03-26 15:33:26 +01:00
Cheol KangandGitHub f6af42018e Allow overriding the Codex web search model (#516)
* Allow codex web search model override

* Handle blank codex web search model
2026-03-26 15:29:54 +01:00
978f56121a Normalize pasted image formats before attach (#543)
Co-authored-by: iter <itertoolz@gmail.com>
2026-03-26 15:28:04 +01:00
daanddenandGitHub b8c423e319 Fix stale OpenAI Responses replay across session boundaries (#534)
* Fix stale OpenAI Responses replay across session boundaries

Fixes #505

* Fix CI tests for session replay change

* Harden session reload and switch rollback

* Guard session switch snapshots

* fix(coding-agent): preserve responses replay snapshots
2026-03-26 15:27:58 +01:00
can1357 b8b1c75d20 chore: bump version to 13.15.2 2026-03-26 14:51:40 +01:00
can1357 80d7a8f413 chore: bump version to 13.15.1 2026-03-26 14:09:28 +01:00
can1357 e31f2ef6f0 refactor: simplified null checks using optional chaining across TypeScript and Rust modules
- Simplified null/empty checks across TypeScript codebase using optional chaining operator (?.) for improved readability.
- Replaced explicit null checks in validation logic with optional chaining in oauth-discovery, gemini-cli, claude, zai, and lsp modules.
- Updated error handling in Rust command invocation to use double question mark operator (??) for cmd_result.
- Consolidated null validation patterns across tools (bash-skill-urls, browser, gemini-image, resolve) and keybindings using optional chaining.
2026-03-26 14:09:14 +01:00
can1357 e5b10b876f test(coding-agent): removed macOS fallback test case from theme detection
- Removed test case for macOS fallback behavior inside Zellij.
2026-03-26 14:06:21 +01:00
zamorakpdsandGitHub 870232104c Fix/resume timestamp mutation (#528)
* fix(coding-agent): kept resumed sessions from reordering

* fix(coding-agent): corrected resume state restoration
2026-03-25 12:40:08 +01:00
can1357 4adaee02da chore: bump version to 13.15.0 2026-03-23 05:58:44 +01:00
can1357 2c93655796 feat(autoresearch): added auto-resume, path validation, and security guards
- Added auto-resume mechanism with state tracking to automatically resume pending experiment runs and prevent duplicate resumptions.
- Added contract path validation to reject unsafe path specifications with absolute paths and parent directory traversal attempts.
- Added secondary metrics input to autoresearch setup flow for specifying tradeoff metrics alongside primary objectives.
- Enhanced command parsing with shell operator detection to reject piped, redirected, or chained autoresearch.sh commands.
- Added prototype pollution guards in object cloning functions to prevent injection via __proto__, constructor, and prototype keys.
- Fixed boundary duplication warnings in hashline detection to properly report multiple overlapping hashline references.
2026-03-23 02:16:39 +01:00
can1357 003f46f42c feat(coding-agent/autoresearch): added contract validation and run tracking system
- Added contract system for validating benchmark commands, metrics, scope paths, constraints, and off-limits paths.
 - Contract validation enforces matching initialization parameters against autoresearch.md before init_experiment.
 - Segment fingerprinting detects configuration drift and warns when metrics are not directly comparable.
 - Added pending run detection and recovery to resume incomplete experiments from .autoresearch/runs/.
 - Run directories organize artifacts with benchmark logs and optional checks logs for traceability.
 - Extended experiment state to track run number, command, scope, off-limits, constraints, and fingerprint.
2026-03-23 01:49:03 +01:00
can1357 9e824d9235 feat(coding-agent): restructured hashline edit schema to nested loc/content objects
- Restructured hashline edit schema from flat op/pos/end/lines fields to nested loc/content objects with discriminated union types.
- Replaced operation names (replace_line, replace_range, append_at, prepend_at) with unified loc object patterns supporting line, block, append, and prepend anchors.
- Updated content field to accept array of strings or null instead of lines parameter in edit entries.
- Refactored edit resolution logic to dispatch on loc shape instead of op enum, simplifying location handling.
2026-03-23 01:45:35 +01:00
can1357 f09c5bd67c feat(patch): added boundary duplication detection to prevent off-by-one range errors
- Added boundary duplication warning to detect off-by-one range errors in replace_range and replace_line operations.
- Updated hashline tool documentation with boundary duplication trap guidance to prevent closing delimiter duplication.
2026-03-22 22:42:42 +01:00
can1357 cac628ccce feat(coding-agent): added git isolation and keybinding utilities for autoresearch
- Added git branch isolation for autoresearch sessions with automatic branch creation, reuse, and worktree safety checks.
- Added scope definition sections (Files in Scope, Off Limits, Constraints) to autoresearch template for explicit session boundaries.
- Added keybinding matcher utilities for consistent escape/cancel key handling across interactive components.
- Added ASI metadata validation (hypothesis and rollback context) in log_experiment tool for experiment tracking.
- Refactored keybinding logic across 13 components to use centralized matcher functions instead of inline key checks.
2026-03-22 21:47:48 +01:00
can1357 012e0c90b9 feat(coding-agent): renamed hashline operation types for clarity
- Renamed hashline operation types for clarity: append->append_at, prepend->prepend_at, append_eof->append_file, prepend_bof->prepend_file.
- Updated all operation type references in patch implementation, tests, and documentation to reflect new naming convention.
- Restructured hashline tool documentation with hierarchical sections and simplified examples for improved clarity.
- Consolidated validation rules and added explicit warning about invalid anchors and operation/field combinations.
2026-03-22 21:14:31 +01:00
can1357 66a478e6e6 feat(coding-agent): implemented dynamic tool activation for autoresearch lifecycle management
- Added `defaultInactive` property to ToolDefinition for conditional tool registration and activation control.
- Added dynamic tool activation/deactivation API `setActiveTools()` for managing experiment tools in autoresearch mode.
- Replaced single `command-start.md` workflow with separate `command-initialize.md` and `command-resume.md` prompts for autoresearch initialization and session resumption.
- Added interactive intent dialog for autoresearch optimization goals with automatic session resumption detection based on autoresearch.md presence.
- Refactored autoresearch command handler to distinguish resume vs initialize flows and dynamically activate/deactivate experiment tools based on mode and session state.
2026-03-22 21:14:17 +01:00
can1357 5d2e2cef1f feat(react-edit-benchmark): added retry mechanism and autocorrect tracking to benchmarks
- Added retry mechanism for benchmark tasks with separate system and retry prompt templates to improve edit success rates.
- Introduced autocorrect tracking metrics including autocorrect-free success rate and edit autocorrect counts in task and benchmark summaries.
- Refactored prompt building into modular functions (buildBenchmarkSystemPrompt, buildInitialBenchmarkPrompt, buildRetryBenchmarkPrompt) with BenchmarkPromptDelivery type for distinguishing initial and follow-up messages.
- Added session management with cache-keyed provider session IDs using xxHash64 and centralized RPC argument building via prepareBenchmarkSessionSetup.
2026-03-22 21:13:42 +01:00
can1357 7fb18faf4c fix(coding-agent): backported pi-mono changes (1feccfed..b21b42d0)
packages/ai:
- feat: expose provider responseId on AssistantMessage
- feat: lazy-load provider modules for faster startup
- fix: hash foreign Responses API tool call IDs exceeding 64-char limit
- fix: ignore null chunks in openai-completions streams
- fix: keep image tool results inline for Gemini 3+ and Antigravity
- fix: correct Bedrock Claude 4.6 context window to 200k
- fix: support prompt caching for Bedrock application inference profiles
- fix: add OpenRouter reasoning payload format
- fix: ignore placeholder Vertex API keys
- fix: skip AJV validation in restricted runtimes
- fix: Anthropic OAuth client injection and responseId extraction
- fix: Codex incomplete/failed response status handling

packages/agent:
- fix: defer steering until after tool execution completes

packages/tui:
- feat: namespaced keybinding IDs with KeybindingsManager conflict detection
- feat: configurable select list column sizing (#2154 by @markusylisiurunen)
- fix: stream truncateToWidth for large strings
- fix: skip Termux height redraws
- fix: stop evicting unrelated default keybindings
- fix: resolve raw backspace ambiguity on Windows Terminal
- fix: clear stale scrollback on session switch (#2155 by @Perlence)
- fix: remove trailing markdown block spacing (#2152 by @markusylisiurunen)

packages/coding-agent:
- feat: add resizable share sidebar (#2435 by @dmmulroy)
- feat: emit OSC 133 command-executed marker
- feat: reload custom themes from disk watcher
- feat: add --fork session flag
- feat: file mutation queue for serialized writes
- feat: initial message consolidation utility
- fix: keybindings migrated to namespaced IDs
- fix: resolve waitForRetry() race when auto-retry produces tool calls
- fix: handle slash-delimited /model refs
- fix: refresh active model after provider updates
- fix: extended transient error patterns for retry
2026-03-22 18:28:40 +01:00
can1357 c9a7bb0c96 feat(patch): restructured hashline edits with explicit replace_line/replace_range and boundary operations
- Refactored hashline edit operations from implicit `replace` to explicit `replace_line` and `replace_range` with mandatory `end` parameter for ranges.
- Added file-level operations `append_eof` and `prepend_bof` for boundary insertions, making `pos` required for anchor-based operations.
- Enforced stricter anchor validation per operation type and simplified edit application logic by separating file-level from anchor-based operations.
- Updated hashline edit application to preserve duplicated boundary lines without auto-correction, changing previous behavior.
2026-03-22 17:17:13 +01:00
can1357 2771c2399d feat(autoresearch): introduced autonomous experiment loop with metric-driven optimization
- Added autoresearch extension with autonomous experiment loop supporting init, run, and log experiment tools for metric-driven optimization.
- Added widget placement system enabling extensions to position UI components above or below the editor via ExtensionWidgetOptions.
- Added dashboard controller with interactive overlay for viewing experiment results, metrics, and progress with keyboard navigation.
- Removed auto-correction logic for off-by-one range edits in hashline editor to preserve user intent in patch operations.
- Added state reconstruction utilities to parse autoresearch.jsonl logs and rebuild experiment state across sessions.
- Added comprehensive type definitions and helper utilities for metric parsing, ASI validation, and process management.
2026-03-22 16:50:13 +01:00
can1357 c0f5711588 revert: exclusive end 2026-03-22 16:49:49 +01:00
can1357 c11cd9023a test(tools): added createTestToolContext helper for consistent tool testing
- Added createTestToolContext helper to construct AgentToolContext instances for tool tests.
- Updated 2 test cases to use createTestToolContext instead of inline object literals for consistency.
- Added imports for AgentToolContext and SessionManager to support new test helper.
2026-03-22 13:57:02 +01:00
can1357 d4b6e869cb feat(coding-agent): implemented pattern-based bash interceptor rules + exclusive range semantics
- Changed bash interceptor configuration from boolean flags to customizable pattern-based rules array.
- Clarified hashline range replace semantics: end parameter is now strictly exclusive boundary.
- Fixed bash interceptor to apply built-in default rules when no custom patterns are configured.
- Updated hashline range validation and calculations to enforce exclusive end semantics throughout.
2026-03-22 12:53:16 +01:00
can1357 855d89cc5e feat(coding-agent): added inline markdown rendering with theme-aware styling
- Added renderInlineMarkdown() utility function to support inline markdown rendering with optional base color styling.
- Refactored ask tool to render questions and option labels with markdown formatting for improved text styling.
- Updated hook-input and hook-selector components to render titles as markdown with theme-aware styling.
- Implemented recursive token processing for nested markdown elements including bold, italic, code, links, and strikethrough.

Fixes #491
2026-03-22 12:38:07 +01:00
can1357 1bf48ca208 fix(modes): corrected setHookWidget to preserve undefined/null values
- Corrected setHookWidget to preserve undefined/null values instead of converting them to string literals.

Fixes #496
2026-03-22 01:39:05 +01:00
can1357 43977259cf fix(ai): corrected quota exhaustion detection to prevent unnecessary retries
- Fixed rate-limit-utils to recognize and classify 'usage limit' errors as QUOTA_EXHAUSTED instead of transient.
- Added isUsageLimitError() utility function for unified detection of persistent quota limit errors across providers.
- Fixed Codex provider to return immediately on usage-limit errors instead of retrying, preventing unnecessary 5-minute delays.
- Removed usage.?limit pattern from TRANSIENT_MESSAGE_PATTERN to prevent misclassification of persistent quota errors.
2026-03-22 01:36:57 +01:00
can1357 b64b6b8a79 feat(modes): introduced ACP mode for headless agent operation with session management and event streaming
- Added ACP (Agent Client Protocol) mode for headless agent operation via --mode acp flag.
- Integrated Agent Client Protocol SDK with session management, streaming communication, and event mapping.
- Added ensureOnDisk() method to SessionManager for immediate session persistence without requiring assistant messages.
- Changed session persistence to use atomic file rewrite for unflushed sessions.
- Implemented AcpAgent class with session management, prompt handling, MCP server configuration, and event streaming.
2026-03-22 01:35:20 +01:00
can1357 7be29ab70f chore: bump version to 13.14.2 2026-03-21 17:34:47 +01:00
can1357 a4026c588e fix(ai): corrected thinking config format and model context windows across 100+ definitions
- Fixed thinking configuration format by replacing `levels` array with `minLevel`/`maxLevel` properties across 100+ model definitions.
- Corrected GPT-5.4 mini/nano context window from 400000 to 272000 tokens for accurate token limit reporting.
- Normalized GPT-5.4 variant priority handling to use parsed variant instead of raw model IDs for consistent behavior.
- Added "mini" variant support to OpenAI model parsing regex and updated thinking mode configuration for Claude models.
- Fixed test robustness by replacing exact string matching with numeric range comparison to handle BSD seq notation on macOS.
- Corrected model generation script execution order to apply policy overrides before promotion target linking.
2026-03-21 17:34:38 +01:00
can1357 b89631541a chore: bump version to 13.14.1 2026-03-21 17:09:04 +01:00
lukeandGitHub 47dc03b835 fix: prevent TUI freeze on massive bash output and fix spinner rendering (#500)
- Sync OutputSink.push(): eliminate promise chain per chunk, buffer
  management and onChunk run inline, file writes deferred via queue
- 64KB native read buffer (was 4KB): reduces chunk count ~16x
- chunkThrottleMs in OutputSink: gate onChunk to every 50ms
- BashExecutionComponent streaming throttle: gate + 100-line cap
- Remove requestRender from chunk callbacks: spinner drives renders
- Remove double sanitization in appendOutput (already done by OutputSink)
- Inline SEGMENT_RESET in TUI doRender buffer writes: eliminates O(N)
  string allocations per frame from #applyLineResets
- Cache header Text in BashExecutionComponent (created once, reused)
- Gate sixel mask computation behind protocol + passthrough check
- Fix spinner: #spinnerFrame made optional, interval calls #updateDisplay
- Remove pendingChunks promise chains from bash-executor and bash-interactive
2026-03-21 16:05:42 +01:00
can1357 32e6379fbe chore: bump version to 13.14.0 2026-03-20 23:53:34 +01:00
can1357 73ceeae59c refactor(coding-agent): restructured model policy application logic for clarity
- Restructured model policy application logic to improve readability and reduce nesting.
- Extracted model override application into dedicated function calls for consistency.
2026-03-20 23:53:24 +01:00
can1357 4f7b61ab62 feat(coding-agent): added MCP server auto-reconnect with SSE monitoring
- Added auto-reconnect capability for MCP servers with SSE stream monitoring and exponential retry backoff.
- Added tool-level reconnect handling for retriable connection errors (ECONNREFUSED, ECONNRESET, 404/502/503).
- Added `/mcp reconnect <name>` command for manual MCP server recovery.
- Improved reconnect robustness by aborting retries when MCP configuration changes via epoch checking.
- Extended transport reconnect handling to all transport types (stdio, HTTP/SSE) with unified onClose logic.
- Added comprehensive test coverage for MCPManager reconnect behavior and tool-level abort propagation.
2026-03-20 23:48:38 +01:00
can1357 9a37cc8096 refactor(coding-agent/mcp): restructured tool-bridge abort handling and argument normalization
- Replaced custom withAbort() implementation with untilAborted() utility from pi-utils for consistent abort handling.
- Extracted normalizeToolArgs() helper to consolidate argument normalization logic across renderCall(), renderResult(), and execute() methods.
- Added MCPToolCallParams type import to improve type safety for tool parameter handling.
2026-03-20 23:42:02 +01:00
can1357 138e7e1d91 fix(interactive-mode): corrected mic cursor width measurement to prevent TUI overflow
- Fixed STT Alt+H mic cursor rendering to measure the actual microphone glyph width, preventing one-column TUI overflow crashes when the active symbol preset uses a wide icon (#484).
- Extracted mic cursor styling into dedicated method to consolidate color and width calculation logic.

Fixes #484
2026-03-20 23:40:48 +01:00
can1357 c718104e8b fix(patch): handled plus-prefixed hashline IDs in diff-style output
Fixes #485
2026-03-20 23:37:04 +01:00
can1357 51228e9811 refactor(coding-agent): restructured MCP reconnection tracking and error handling
- Refactored MCP manager reconnection logic to track configuration changes via reconnectEpoch parameter.
- Consolidated error pattern matching in tool-bridge to use lowercase normalization for consistent comparison.
- Extracted retry provider variables to eliminate repeated ternary expressions in error handling.
- Reformatted code across multiple files for improved readability with consistent multi-line formatting.
- Added test coverage for transport reconnection scenarios including connection reuse after reconnection.
2026-03-20 23:32:07 +01:00
can1357 c60607bfb0 feat(coding-agent/prompts): enabled thinking and streamlined explore agent output schema
- Updated explore agent thinking level from off to med for improved reasoning.
- Simplified explore agent output schema: consolidated file references into single `ref` field with optional line ranges instead of separate `path`, `line_start`, `line_end` fields.
- Removed `code` section from explore agent output (critical code excerpts no longer extracted).
- Removed `dependencies`, `risks`, and `start_here` sections from explore agent output.
2026-03-20 23:32:07 +01:00
RzNmKXandGitHub 31c3ec55e7 fix(ai): restore xhigh thinking level for Bedrock Opus 4.6 (#481)
The 8e3e0ebf9 refactor centralized thinking inference but dropped
Bedrock-specific handling:

- inferAnthropicSupportedEfforts gated xhigh on anthropic-messages API,
  excluding Bedrock models
- inferThinkingControlMode returned 'budget' for all Bedrock models
  instead of 'anthropic-adaptive' for 4.6+
- parseAnthropicModel regex captured Bedrock date suffixes (20250819)
  as version components, causing parseSemVer to return null

Fixes:
- Remove API guard from inferAnthropicSupportedEfforts so Opus 4.6+
  gets xhigh regardless of API transport
- Add Anthropic version checks to Bedrock case in
  inferThinkingControlMode (4.6+ adaptive, 4.5+ budget-effort)
- Narrow version regex from \d+ to \d{1,2} so date suffixes in
  Bedrock model IDs are not absorbed
- Update four stale Bedrock Opus 4.6 entries in models.json
- Add thinkingBudgets.xhigh (default 32768) to settings schema
2026-03-20 23:29:12 +01:00
8b028224e3 feat: auto-reconnect MCP servers on connection loss (#482)
* feat: auto-reconnect MCP servers on connection loss

When an HTTP SSE stream drops (server restart, network interruption),
the transport fires onClose, and the manager proactively reconnects
with retry backoff (500ms, 1s, 2s, 4s). Tools are kept in the registry
during reconnection so they remain selected and available to the agent.

If proactive reconnection fails, stale tools stay registered. When the
agent calls one, the tool bridge detects the retriable connection error
(ECONNREFUSED, ECONNRESET, stale session 404/502/503, etc.), triggers
reconnectServer on the manager, and retries the call once on the fresh
connection. Concurrent reconnect attempts for the same server are deduped.

connectToServer now always installs a default onRequest handler for ping
and roots/list (using getProjectDir()), so all connections -- including
short-lived test/probe ones -- properly respond to server-initiated
requests during initialization.

Post-connection setup (resources, prompts, subscriptions) is extracted
into a shared #loadServerResourcesAndPrompts method used by both initial
connection and reconnection paths.

Add /mcp reconnect <name> command for manual recovery after extended
outages where both proactive and reactive reconnection have failed.

* docs: add changelog entry for MCP auto-reconnect

* fix: address P1 review findings in MCP reconnection

- Save server configs before connection attempt so deferred tools can
  reconnect even when the initial connection timed out (P1-1)
- Make waitForConnection() and getConnectionStatus() aware of in-flight
  reconnections so callers wait instead of failing immediately (P1-2)
- Add epoch counter incremented on disconnectAll() and checked in
  connectAndWireServer() to invalidate stale reconnect attempts that
  outlive a manager reset/reload (P1-3)
- Skip servers with pending reconnections in connectServers() to prevent
  parallel connection attempts for the same server

* fix: deferred tool reconnect and non-blocking transport teardown

- DeferredMCPTool.execute now reconnects when getConnection() fails
  ("MCP server not connected"), not only on network errors from
  callTool. Servers that missed the startup window can now be woken
  by the first tool call against their cached tools. (P1-4)
- #doReconnect fire-and-forgets the old transport close instead of
  awaiting it. HttpTransport.close() sends a DELETE with 30s timeout;
  blocking here delayed the first reconnect attempt by that amount
  on every server restart. (P1-5)

* fix: abort-aware reconnect waits and preserve tool selection on reconnect

- Wrap all reconnect() awaits with withAbort(signal) so user
  cancellation (Esc) interrupts the reconnect backoff loop instead
  of blocking for up to 7.5s. Applies to MCPTool (1 site) and
  DeferredMCPTool (2 sites). (P2-1)
- Remove activateDiscoveredMCPTools call from /mcp reconnect handler.
  refreshMCPTools already preserves the user's prior MCP tool
  selection; the extra activation was silently opting into all
  server tools including ones the user had not enabled. (P2-2)

* fix: rebind MCPTool connection after reconnect, add stdio retriable error

- MCPTool.connection is now mutable; after a successful reconnect retry,
  this.connection is rebound to the fresh connection so subsequent calls
  on the same instance (e.g. batched tool calls) use it instead of
  triggering another reconnect cycle. (P2-3)
- Add "Transport closed" to RETRIABLE_PATTERNS. StdioTransport rejects
  pending requests with this message when the subprocess dies, which
  should trigger the reconnect path just like HTTP transport errors. (P2-4)

---------

Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
2026-03-20 23:27:30 +01:00
8013be90c9 fix(ai): guard resumed OpenAI Responses replay (fixes #488) (#489)
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
2026-03-20 23:24:37 +01:00
汐andGitHub ed2d10cf59 fix(upstream): allow to using extension dir in "pi.extension" of packages.json (#480)
ref
https://github.com/badlogic/pi-mono/blob/8a8e2a8049fba194d999cf7f3ec4b247910657ee/packages/coding-agent/src/core/package-manager.ts#L435-L438
2026-03-20 23:16:53 +01:00
91f7417975 fix(coding-agent): respect user model overrides in hardcoded policies (#483)
Apply user modelOverrides after hardcoded defaults so contextWindow
and other fields from models.json take precedence.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 23:14:20 +01:00
can1357 09a243444b refactor(coding-agent/session): simplified plan mode enforcement logic
- Simplified plan mode enforcement by removing tool retrieval and restoration logic.
- Replaced tool existence checks with registry lookup to reduce intermediate variables.
2026-03-19 20:18:12 +01:00
can1357 b00d01493d fix: guard model access in formatSessionAsText for sessions without model 2026-03-19 06:49:08 +01:00
can1357 999d066d71 chore: bump version to 13.13.2 2026-03-18 23:21:20 +01:00