- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
Two boundary defects on the mirrored-todo path.
1. The Agent error drain snapshotted #cursorToolResultBuffer without
awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
async cursorOnToolResult still running when the provider errored
patched an entry the catch path had already detached, so the
pre-transform payload was persisted. A provider error is exactly when
a transform is most likely to be in flight.
2. The todo renderer interpolated mirrored provider text straight into
terminal output. A Cursor snapshot carries model-authored task
content, phase names and summary text verbatim, so a label holding
ANSI/C0 sequences rewrote the terminal on every render and replay.
sanitizeText alone is not enough - it preserves tabs, which punch
holes in bordered output - so every display path now funnels through
one forDisplay() helper: task labels, blocker notes, phase headers,
the zero-task fallback, and the streaming renderCall preview. Raw
values are untouched; content and phase name are the identity keys
the local list is looked up by and what gets persisted.
Three orphan paths, same failure mode: the assistant block is marked
kCursorExecResolved before the work runs, so agent-loop.ts emits no
placeholder for it, and any path that produces no toolResult leaves the
call unpaired — buildSessionContext then strips the whole interaction
from every rebuilt transcript.
1. resolveExecHandler returned no toolResult on three exits (no handler
installed, handler produced nothing, handler threw). Each now pairs a
result carrying the same text the server sees in execResult, routed
through onToolResult like a real one. `pairing` is a required
parameter so a new callsite cannot silently recreate the orphan.
2. Agent only installed its result-buffer sink when cursorExecHandlers
or cursorOnToolResult was set. Both are optional, so a bare SDK host
dropped the provider result on the floor. Installed unconditionally;
a non-Cursor provider never calls it.
3. A todo completion frame with no tool_call (the field is optional)
skipped settlement entirely. It now settles as "nothing to mirror".
Also fixes an empty update_todos with a nonzero total_count being
mirrored as an authoritative clear: the length guard added earlier
skipped the mismatch check for empty responses, so a partial or
size-limited merge response deleted every local task at once.
An async cursorOnToolResult that resolved after the message_end drain
had its rewrite silently discarded: the reservation kept the call from
dangling, but the late patch mutated a buffer entry the drain had
already detached, so the persisted message kept the pre-transform
payload.
Each entry now records the in-flight transformer promise, and
#emitCursorSplitAssistantMessage awaits any that are still pending
before appending and emitting. This matches the exec-channel paths,
which already await onToolResult (cursor.ts:1461).
A rejecting transformer is swallowed per-entry, so a failing hook can
neither take the turn down nor cost the reserved result.
The previous test asserted the old limitation (late rewrite NOT
persisted) and its premise is now unreachable, so it is replaced by the
rejection contract. Stale limitation notes in agent.ts and types.ts are
updated.
Two review findings on the native todo sync.
- TodoItem.dependencies is a graph the local model cannot store: rows
are keyed by content, carry no id, and hold no edges. An imported
dependent row files as plain pending and nextActionableTask then
offers work the server considers blocked. Refuse snapshots with an
edge pointing at an unfinished row; edges whose blockers already
finished constrain nothing and still mirror.
- The todo failure warning interpolated the provider error verbatim.
Collapse and truncate it at the render boundary.
Also documents two known, unfixed defects: an async cursorOnToolResult
transformer resolving after the buffer drain, and the todo card
lifecycle race. Emitting a synthetic tool_execution_start for the
latter was measured and rejected -- the completion deletes the entry it
creates, so the late streamed block adds a second card.
The previous commit paired every server-resolved todo block with a
result, but built that result in the provider from the flat snapshot.
`todoToolRenderer.renderResult` reconstructs the list exclusively from
`details.phases`, so the block survived the dangling-strip only to replay
as `Todo 0 tasks`.
Only the host computes that grouping -- the provider sees a flat list --
so `todoSync` now returns the result it already assembled and the
provider persists it verbatim. A refused snapshot never reaches the host,
so the provider's summary-only fallback still covers that path, and
exactly one result is emitted either way.
Separately, `Agent`'s Cursor buffering wrapper pushed its entry only
after awaiting the optional `cursorOnToolResult` transformer. The
provider dispatches decoded messages with `void handleServerMessage(...)`,
so a `message_end` from the same chunk could drain the buffer while a
transformer was still pending, dropping the result. The entry is now
reserved synchronously and patched in place when the transformer
resolves, keeping buffer order and still applying the customization.
Production is unaffected -- `sdk.ts` sets no transformer -- but the
option is supported and its contract returns a Promise.
Tests: a delayed-transformer case that loses the result without the
buffering change, and a replay case driving the persisted result through
`buildSessionContext` and asserting `details.phases` rebuilds a non-empty
list -- the id-pair assertion alone did not catch the empty render.
API-level refusals now stay visible as terminal errors without being sent back as assistant dialogue on the next provider request. Added core and coding-agent conversion coverage for Anthropic refusal metadata.
Fixes#3592
- Validated queued toolChoice against active tools in agent and coding-agent sessions.
- Rejected queued forced choices with reason "unavailable" when selected tools were inactive.
- Dropped provider toolChoice payloads when requested function tools were not offered.
- Probed Tokio worker-thread support and fell back to current-thread runtime creation.
- Re-polled steering at the loop yield boundary and included it in the pre-stop pending batch so late messages are processed immediately.
- Added session-side draining for stranded queued messages, scheduling an auto-continue when a prompt settles and follow-ups or steers remain.
- Added a regression test for late steering injection at yield and updated mid-turn collab prompt handling to keep steering messages in the pending display queue until consumed.
Added AgentLoopConfig.getDisableReasoning so the agent loop refreshes disableReasoning on every model call, matching getReasoning. Mid-run thinking-level transitions in and out of off now propagate to the next request instead of sticking on the value captured at prompt start.\n\nFixes #2239
Propagated explicit thinking-off state through the agent loop so provider requests receive disableReasoning instead of an undefined effort. Added Ollama and agent-session regressions for the :off path.\n\nFixes #2239
- Added `/tan` slash command registration and interactive handling.
- Added TanCommandController validation and async task scheduling for `/tan` dispatch.
- Added session cloning that suppresses breadcrumbs, copies artifacts, and handles abort cleanup.
- Added `promptCacheKey` support in Agent and inherited `providerPromptCacheKey` in session creation.
- Emitted content-filter blocks as normal assistant error lifecycle events.
- Prevented interactive clients from silently dropping the streaming bubble.
- Finalized an existing assistant stream when blocked mid-stream.
F3 (agent): mutate state.messages and state.pendingToolCalls in place on
appendMessage/popMessage/clearMessages/reset/tool_execution_start/_end
instead of allocating a fresh array/Set on every transition. Subscribers
that capture state.messages by reference now observe updates directly.
Public type signature unchanged.
F5 (ai): add parseStreamingJsonThrottled to utils/json-parse — a per-delta
wrapper around parseStreamingJson that skips the re-parse until the
buffer has grown by minGrowthBytes (default 256). Wired into every
provider's tool-call argument accumulator (anthropic, amazon-bedrock,
openai-completions, openai-codex-responses, openai-responses-shared) so
per-delta cost becomes O(N) in total buffer length instead of O(N²).
Every provider's toolcall_end still runs a final unthrottled parse, so
the published block.arguments is unchanged.
F8 (coding-agent): drop the per-delta structuredClone of streaming tool
arguments in ToolExecutionComponent.updateArgs. event-controller.ts and
ui-helpers.ts already spread their input into a fresh object on each
delta, so cloning here was dead work on the rendering hot path. Added a
reference-equality short-circuit so repeat calls with the same args
object skip the preview-diff and display refresh.
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
- Replaced custom MockAssistantStream helpers with createMockModel streams across agent tests.
- Removed manual queueMicrotask stream-event scripting in favor of scripted mock responses.
- Consolidated helper fixtures by deleting local aliases and reusing shared user-message/model helpers.
- Updated test assertions to use mock.calls and mock.model metadata for call and context validation.
Anthropic counts sessions by metadata.user_id. Without this fix, OMP
generated fresh random entropy on every API request, inflating the
session count and preventing backend attribution to the authenticated
account.
Changes:
packages/ai:
- resolveAnthropicMetadataUserId() now accepts JSON-format user_id
matching real Claude Code's getAPIMetadata shape
({ session_id, account_uuid, ... }). Previously only the legacy
cloaking format was accepted on OAuth, causing stable caller-supplied
values to be silently discarded.
- AnthropicOAuthFlow.exchangeToken() and refreshAnthropicToken() now
populate OAuthCredentials.{accountId, email} from the token response
account block, removing the need for a separate /api/oauth/profile
round-trip.
- AuthStorage.getOAuthAccountId(provider, sessionId) returns the OAuth
accountId for the session-sticky credential, used to build
account_uuid in metadata.user_id. Guards against misattribution for
API-key, runtime-override, env-key, and fallback-resolver paths that
do not record a session credential.
packages/agent:
- Agent.metadataForProvider(provider) resolves request metadata for
the given provider via the installed resolver, or returns the static
metadata value. The plain metadata getter now returns only the static
value; provider-aware resolution is explicit.
- Agent.setMetadataResolver(fn) installs a (provider: string) resolver
evaluated per LLM request in agent-loop, after getApiKey records the
session-sticky credential, so account_uuid reflects the credential
actually used.
- AgentLoopConfig.metadataResolver is called with config.model.provider
after getApiKey, overriding the static metadata field.
packages/coding-agent:
- AgentSession.#syncAgentSessionId installs a metadata resolver that
builds { user_id: JSON.stringify({ session_id, account_uuid? }) },
matching the Anthropic session attribution format. account_uuid is
only included for provider="anthropic" to avoid leaking the OAuth
identity to third-party Anthropic-format-compatible providers.
- sessionId getter prefers providerSessionId when supplied via
AgentSessionConfig so all API paths (getApiKey, direct calls,
metadata resolver) share the same provider-facing session ID.
- prepareSimpleStreamOptions stamps session metadata on direct calls
(runEphemeralTurn, compaction, branch summary, title generation) so
they share the same session bucket as Agent.prompt requests.
- generateBranchSummary and generateSessionTitle accept a
(provider: string) metadata resolver evaluated after their own
getApiKey call for correct credential attribution.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Added an optional `getReasoning` callback to `AgentLoopConfig` to resolve reasoning effort dynamically for each LLM call.
- Updated the agent loop to resolve reasoning via `getReasoning` and use it in place of static `reasoning` when provided.
- Added a test confirming a run re-reads the thinking level between consecutive model calls when it changes mid-run.
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
* Add MCP tool discovery search and live refresh
* Fix MCP discovery review feedback
* Address remaining MCP discovery review comments
* feat: compact MCP discovery search results
* fix: align MCP discovery search contract
* feat: add MCP server tool counts to discovery hints
* fix(agent): corrected stale toolChoice validation against active tools
- Fixed stale forced toolChoice passed to provider after mid-turn tool refresh by validating against active tools.
- Added refreshToolChoiceForActiveTools() to filter invalid tool choices when available tools change.
- Changed getToolChoice config to use computed function instead of static property for dynamic validation.
- Fixed MCP tool selection tracking in coding-agent to distinguish between discovery-enabled and non-discovery sessions.
- Updated search_tool_bm25 to filter already-selected tools before applying limit parameter.
---------
Co-authored-by: can1357 <me@can.ac>
- Extracted OpenAI Responses API stream processing logic into shared module with 424 lines of reusable utilities.
- Consolidated text signature encoding, tool call normalization, and message conversion functions into openai-responses-shared module.
- Refactored openai-responses.ts and azure-openai-responses.ts to delegate stream processing to shared processResponsesStream() helper.
- Removed 473 lines of duplicated stream event handling and utility functions across OpenAI provider implementations.
- Added `onPayload` callback option to intercept and transform provider request payloads before transmission across agent and AI packages.
- Added structured text signature metadata with phase information to OpenAI and Azure OpenAI providers for enhanced response tracking.
- Added `before_provider_request` extension event to coding-agent for chaining payload transformations across multiple handlers.
- Improved error messages in `response.failed` events with detailed error codes, messages, and incomplete reasons from provider responses.
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
- Added ModelManager API with createModelManager() factory for managing bundled and dynamically discovered models with configurable refresh strategies.
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints with provider-specific model manager configuration helpers.
- Renamed public API functions for clarity: getModel() -> getBundledModel(), getModels() -> getBundledModels(), getProviders() -> getBundledProviders().
- Added on-disk model caching with TTL-based invalidation and resolveProviderModels() function for runtime model resolution with source precedence.
- Refactored model discovery script to dynamically fetch models from Codex, Cursor, and Antigravity using OAuth credentials instead of hardcoded lists.
- Extracted delta event batching and throttling logic into AssistantMessageEventStream base class to eliminate duplication across test mocks.
- Changed access modifiers from private to protected in EventStream class to allow subclass customization of queue, waiting, and completion handling.
- Simplified MockAssistantStream implementations across five test files by removing custom event handling logic and delegating to AssistantMessageEventStream.
- Added protected methods deliver() and endWaiting() to EventStream to support subclass event delivery patterns.
- Implemented delta event type guard and merging utilities to consolidate event batching logic in one place.
- Removed Prettier configuration files (.prettierignore and .prettierrc) and migrated formatting to Biome.
- Updated Biome configuration from version 2.3.11 to 2.3.12 and changed arrowParentheses rule from 'always' to 'asNeeded'.
- Pinned @biomejs/biome dependency to exact version 2.3.12 in package.json and bun.lock.
- Applied consistent arrow function formatting across 489 files by removing unnecessary parentheses around single parameters.
- Removed blank lines after comment blocks and reorganized imports for consistency across the codebase.
- Removed WASM generation script; use Bun `wasm?raw` loader for imports.
- Added bunfig.toml with loaders for `.md`, `.py`, and `.wasm?raw` text imports.
- Added types/assets/index.d.ts for global TypeScript module declarations.
- Unified TypeScript configuration with tsgo-based checking across monorepo.
- Removed build and WASM steps from install and publish pipelines.
- converted relative imports to path aliases ($c/*, $ai/*, $tui/*, etc.) across all packages
- added per-package tsconfig.json with complete path mappings for runtime resolution
- set importModuleSpecifier to non-relative for IDE auto-import preferences
- updated dev script to run from monorepo root for consistent path resolution
- Port pi-ai from upstream with Bun-first approach (source files, no dist)
- Add cli_ prefix to tool names in OAuth mode to avoid collisions
- Improve beta header handling with proper deduplication
- Convert codex-instructions.md to Bun text embed
- Strip .js extensions from all imports