* Add MCP tool discovery search and live refresh
* Fix MCP discovery review feedback
* Address remaining MCP discovery review comments
* feat: compact MCP discovery search results
* fix: align MCP discovery search contract
* feat: add MCP server tool counts to discovery hints
* fix(agent): corrected stale toolChoice validation against active tools
- Fixed stale forced toolChoice passed to provider after mid-turn tool refresh by validating against active tools.
- Added refreshToolChoiceForActiveTools() to filter invalid tool choices when available tools change.
- Changed getToolChoice config to use computed function instead of static property for dynamic validation.
- Fixed MCP tool selection tracking in coding-agent to distinguish between discovery-enabled and non-discovery sessions.
- Updated search_tool_bm25 to filter already-selected tools before applying limit parameter.
---------
Co-authored-by: can1357 <me@can.ac>
- Extracted OpenAI Responses API stream processing logic into shared module with 424 lines of reusable utilities.
- Consolidated text signature encoding, tool call normalization, and message conversion functions into openai-responses-shared module.
- Refactored openai-responses.ts and azure-openai-responses.ts to delegate stream processing to shared processResponsesStream() helper.
- Removed 473 lines of duplicated stream event handling and utility functions across OpenAI provider implementations.
- Added `onPayload` callback option to intercept and transform provider request payloads before transmission across agent and AI packages.
- Added structured text signature metadata with phase information to OpenAI and Azure OpenAI providers for enhanced response tracking.
- Added `before_provider_request` extension event to coding-agent for chaining payload transformations across multiple handlers.
- Improved error messages in `response.failed` events with detailed error codes, messages, and incomplete reasons from provider responses.
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Extracted tool result emission logic into dedicated emitToolResult function to eliminate duplication.
- Consolidated tool execution result handling to emit results immediately after execution rather than deferring to post-processing loop.
- Simplified post-execution loop by delegating result emission to emitToolResult and removing redundant message construction.
- Exported `finalizeSubprocessOutput()` function and `SubmitResultItem` interface for subprocess output finalization with submit_result validation.
- Added automatic reminders (up to 3) when subagent stops without calling submit_result tool, aborting with exit code 1 after final reminder.
- Extracted subprocess output finalization logic into dedicated `finalizeSubprocessOutput()` function for improved testability and reusability.
- Added comprehensive test coverage for subagent warning injection, reminder behavior, and abort handling in executor module.
- Added streamed tool intent display in working message to show real-time intent tracking during agent execution.
- Changed intent tracing field name from `$intent` to `_intent` across tool schemas and agent core for consistency.
- Added support for file deletion and renaming operations in hashline edit mode.
- Renamed hashline edit operation fields: `set` to `target`/`new_content`, `set_range` to `first`/`last`/`new_content`, `insert` to `inserted_lines`.
- Added optional `intent` field to `ToolCall` interface for capturing harness-level intent metadata.
- Added `intentTracing` configuration option to enable intent goal extraction from tool calls with automatic `$intent` field injection and argument stripping.
- Implemented intent injection and extraction logic in agent-loop to populate tool call intent metadata when intentTracing is enabled.
- Added `tools.intentTracing` setting to coding-agent configuration schema with environment variable override support.
- Added ModelManager API with createModelManager() factory for managing bundled and dynamically discovered models with configurable refresh strategies.
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints with provider-specific model manager configuration helpers.
- Renamed public API functions for clarity: getModel() -> getBundledModel(), getModels() -> getBundledModels(), getProviders() -> getBundledProviders().
- Added on-disk model caching with TTL-based invalidation and resolveProviderModels() function for runtime model resolution with source precedence.
- Refactored model discovery script to dynamically fetch models from Codex, Cursor, and Antigravity using OAuth credentials instead of hardcoded lists.
- Migrated timing measurements from `performance.now()` to `Bun.nanoseconds()` for higher precision benchmarking across all benchmark and timing-sensitive code.
- Updated elapsed time calculations to convert nanoseconds to milliseconds using division by 1e6 to maintain consistent time units.
- Increased benchmark precision from 4 to 6 decimal places for per-operation timing measurements.
- Refactored grep benchmark to use averaged timing values instead of storing individual iteration times, reducing memory overhead.
- Added `concurrency` option to `AgentTool` interface supporting 'shared' (default, parallel execution) and 'exclusive' (solo execution) modes.
- Implemented parallel execution of shared tools within a single agent turn with ordered result emission.
- Refactored tool execution to support concurrent scheduling with interrupt handling via steering messages and AbortSignal.
- Added steering message checking mechanism that interrupts tool execution when user messages arrive during tool processing.
- Removed HTTP proxy setup code from stream.ts that was conditionally setting up undici global dispatcher.
- Added persistent shell session support for bash tool with environment variable preservation across commands.
- Added shellForceBasic setting to force bash/sh even if user's default shell is different (default: true).
- Added OMP_SHELL_PERSIST environment variable to control persistent shell behavior (set to 0 to disable).
- Restructured system prompt with coordinator-specific guidance for parallel task delegation.
- Extracted delta event batching and throttling logic into AssistantMessageEventStream base class to eliminate duplication across test mocks.
- Changed access modifiers from private to protected in EventStream class to allow subclass customization of queue, waiting, and completion handling.
- Simplified MockAssistantStream implementations across five test files by removing custom event handling logic and delegating to AssistantMessageEventStream.
- Added protected methods deliver() and endWaiting() to EventStream to support subclass event delivery patterns.
- Implemented delta event type guard and merging utilities to consolidate event batching logic in one place.
- Removed Prettier configuration files (.prettierignore and .prettierrc) and migrated formatting to Biome.
- Updated Biome configuration from version 2.3.11 to 2.3.12 and changed arrowParentheses rule from 'always' to 'asNeeded'.
- Pinned @biomejs/biome dependency to exact version 2.3.12 in package.json and bun.lock.
- Applied consistent arrow function formatting across 489 files by removing unnecessary parentheses around single parameters.
- Removed blank lines after comment blocks and reorganized imports for consistency across the codebase.
- Removed WASM generation script; use Bun `wasm?raw` loader for imports.
- Added bunfig.toml with loaders for `.md`, `.py`, and `.wasm?raw` text imports.
- Added types/assets/index.d.ts for global TypeScript module declarations.
- Unified TypeScript configuration with tsgo-based checking across monorepo.
- Removed build and WASM steps from install and publish pipelines.
- converted relative imports to path aliases ($c/*, $ai/*, $tui/*, etc.) across all packages
- added per-package tsconfig.json with complete path mappings for runtime resolution
- set importModuleSpecifier to non-relative for IDE auto-import preferences
- updated dev script to run from monorepo root for consistent path resolution
- Added ToolCallContext interface with batch execution metadata including batchId, index, and total.
- Enhanced getToolContext() to accept optional ToolCallContext parameter for batch-aware tool execution.
- Implemented LSP batching support to coalesce formatting and diagnostics operations on parallel edits.
- Added comprehensive tests for tool call batch context and LSP writethrough batching functionality.
- Added thinkingLevel field to agent frontmatter allowing subagents to override thinking level with clamping to model capabilities.
- Added structured output schema to reviewer agent ensuring consistent review output format with findings, correctness verdict, and confidence.
- Expanded system prompt with defensive reasoning guidance including assumption checks, edge case awareness, and completion reflex inhibition.
- Consolidated expandPath, parseFrontmatter, and parseAgentFields utilities into discovery/helpers module eliminating code duplication across loaders.
- Replaced silent error handling with logger.warn calls across config parsing, migrations, extension loading, and OAuth refresh failures.
- Port pi-ai from upstream with Bun-first approach (source files, no dist)
- Add cli_ prefix to tool names in OAuth mode to avoid collisions
- Improve beta header handling with proper deduplication
- Convert codex-instructions.md to Bun text embed
- Strip .js extensions from all imports
- Changed isError property from required to optional across ToolResultMessage and tool event interfaces.
- Added SpinnerType variants with distinct animation frames per symbol preset.
- Fixed spinner animation crash when frames array is empty by adding guard clause.
- Refactored agent-loop tool execution to consolidate toolCallId, toolName, and isError into details object.
- Renamed npm package scope from @mariozechner to @oh-my-pi across all packages for consistent branding.
- Removed packages/pods directory and all related pod management functionality.
- Updated import paths and module declarations throughout codebase to reflect new package names.
- Reformatted code with consistent trailing comma removal and ternary operator alignment.
- agentLoop now accepts AgentMessage[] instead of single message
- agent.prompt() accepts AgentMessage | AgentMessage[]
- Emits message_start/end for each message in the array
- AgentSession.prompt() builds array with hook message + user message
- TUI now receives events for before_agent_start injected messages
- Renamed AppMessage to AgentMessage throughout
- New agent-loop.ts with AgentLoopContext, AgentLoopConfig
- Removed transport abstraction, Agent now takes streamFn directly
- Extracted streamProxy to proxy.ts utility
- Removed agent-loop from pi-ai (now in agent package)
- Updated consumers (coding-agent, mom) for AgentMessage rename
- Tests updated but some consumers still need migration
Known issues:
- AgentTool, AgentToolResult not exported from pi-ai
- Attachment not exported from pi-agent-core
- ProviderTransport removed but still referenced
- messageTransformer -> convertToLlm migration incomplete
- CustomMessages declaration merging not working properly
- Add agentLoopContinue() to pi-ai for resuming from existing context
- Add Agent.continue() method and transport.continue() interface
- Simplify AgentSession compaction to two cases: overflow (auto-retry) and threshold (no retry)
- Remove proactive mid-turn compaction abort
- Merge turn prefix summary into main summary
- Add isCompacting property to AgentSession and RPC state
- Block input during compaction in interactive mode
- Show compaction count on session resume
- Rename RPC.md to rpc.md for consistency
Related to #128
Previously, errors in turn_end events (e.g., from OpenRouter Auto Router)
were not captured in agent.state.error, making failed requests appear as
successful completions.
Fixes#6
Tool results now use content blocks and can include both text and images.
All providers (Anthropic, Google, OpenAI Completions, OpenAI Responses)
correctly pass images from tool results to LLMs.
- Update ToolResultMessage type to use content blocks
- Add placeholder text for image-only tool results in Google/Anthropic
- OpenAI providers send tool result + follow-up user message with images
- Fix Anthropic JSON parsing for empty tool arguments
- Add comprehensive tests for image-only and text+image tool results
- Update README with tool result content blocks API
- Added string-width library for proper terminal column width calculation
- Fixed wrapLine() to split by newlines before wrapping (like Text component)
- Fixed Loader interval leak by stopping before container removal
- Changed loader message from 'Loading...' to 'Working...'