- Added `hasSteeringMessages` to `executeToolCalls` polling so queued steering is detected without dequeuing.
- Retained fallback to `getSteeringMessages` by consuming messages when no peek callback exists.
- Added optional `hasSteeringMessages` config hook and limited steering checks to boundaries.
- Fixed interrupted tool-batch steering by keeping queued messages until boundary handling.
- Added idle text and image submissions to steer queueing when no input waiter exists.
- Auto-continued resumable sessions after queued steering and preserved submit metadata.
- Added `renderSnapcompactFrames()` and `snapcompactFrameCount()` to @oh-my-pi/snapcompact for paging arbitrary text into PNG image blocks without dim-marker bookkeeping.
- Widened the agent loop's `transformProviderContext` hook to `(context, model) => Context` so per-request transforms can gate on the dispatch model's capabilities.
- Added `SnapcompactInlineTransformer` rendering the system prompt and large historical tool results as snapcompact frames on vision models: vision gate, per-provider image budgets, 3k-token floor, savings-margin gate, skip-last rule, and hash-keyed render caches swept to live tool calls.
- Added default-off `snapcompact.systemPrompt` and `snapcompact.toolResults` settings under a new Context → Experimental group, composed after secret obfuscation in `sdk.ts` so frames are built per-request and never persisted to session.jsonl.
- Added prompt stubs (`snapcompact-system-stub.md`, `snapcompact-system-frames-note.md`, `snapcompact-toolresult-note.md`) and unit tests covering frame paging, no-mutate guarantees, budget caps, gates, and render caching.
- Mapped Codex `end_turn:false` terminal events to `pause_turn` stop details in response stream parsing.
- Updated `agent-loop` to re-sample `pause_turn` turns, reset on tool calls, and cap continuations at 8.
- Added coverage for pause-turn mapping and continuation-capping in agent and AI stream tests.
- Extended `AgentTool.concurrency` to accept per-call resolver functions and resolved concurrency mode from each tool call, falling back to exclusive on resolver errors.
- Updated BashTool to schedule non-PTY calls as shared and PTY calls as exclusive so non-interactive bash calls can run in parallel within one message.
- Tracked in-use persistent shell sessions in the bash executor and routed overlapping calls on the same session key to isolated one-shot shells while preserving owner session availability.
Interrupting mid-tool execution (e.g. Enter with a pending steer) drained the steering queue into the dying run — it landed in history without a response and the post-abort resume saw an empty queue, so the agent stopped instead of continuing. Steering/follow-up/aside queue polls in runLoopBody and the post-tool-call check in executeToolCalls are now skipped once the run's abort signal fires, leaving the queue intact for Agent.continue().
Anthropic rejects tool_result blocks when is_error is true and content is
empty after trimming whitespace. Fill a placeholder at encode time and in
coerceToolResult so wedged sessions recover on the next request.
Added AgentLoopConfig.getDisableReasoning so the agent loop refreshes disableReasoning on every model call, matching getReasoning. Mid-run thinking-level transitions in and out of off now propagate to the next request instead of sticking on the value captured at prompt start.\n\nFixes #2239
- Kept committed tool-call blocks on TTSR/user-interrupt aborts, pairing them with a labeled placeholder result.
- Dropped incomplete tool calls only on anonymous `abort()` where partial args are unsafe to replay.
- Stopped aborted runs from waiting on provider iterator cleanup.
- Ran afterToolCall for completed executions after a run aborts.
- Ignored unknown Anthropic content blocks while preserving known ones.
- Kept sampling params and tool-result caching for disabled thinking.
- Updated `docs/ttsr-injection-lifecycle.md` to document `astCondition` behavior, including registration rules and tool-stream matching.
- Extended `packages/natives/test/native.test.ts` with new `astMatch` coverage for Smart matching, metavariable consistency, parse errors, and empty-language rejection.
- Deferred non-interrupting asides to the outer drain when a turn was finishing, preventing an extra model turn from firing before queued follow-ups.
- Coerced post-hook tool results before emission and only marked interrupted tool calls as skipped when execution failed, preserving completed call outcomes.
- Reworked Anthropic CCH patching with native `Buffer.indexOf`, and logged a warning while sending the body unchanged when the billing placeholder could not be anchored.
- Removed `maxToolCallsPerTurn` from `AgentOptions`, `AgentLoopConfig`, and config serialization.
- Removed stream-loop cap enforcement, including the `toolcall_end` abort path and capped assistant messages.
- Removed Anthropic Opus 4.8 batch-cap resolver and agent-session sync logic from coding-agent.
- Updated tests and changelogs to align with uncapped tool-call behavior and dropped cap-specific cases.
- Introduced `AsideMessage` as a message-or-thunk union so aside providers can defer injection decisions.
- Updated agent loop handling to resolve aside thunks at injection time and skip entries that returned `null`, then switched the session yield queue to `drainLazy` for deferred message building.
- Added tests validating lazy aside evaluation and staleness-aware dropping when everything becomes stale after dequeueing.
- Added a new aside-message source on `Agent` and exposed it in `AgentLoopConfig` as `getAsideMessages`.
- Updated the agent loop to poll aside messages after tool batches and before yielding, merging them with follow-up messages before continuing.
- Changed coding-agent session and yield queue handling to pull queued background messages via `drainMessages` at step boundaries instead of streaming-only injection.
- Updated `executeToolCalls` to always pass the active `toolSignal` into `tool.execute` rather than bypassing it for non-abortable tools.
- Removed the `nonAbortable` option from `AgentTool` and from the read, write, and edit tools so they can no longer opt out of abort handling.
- Documented the cancellation behavior change in the affected tool docs and package changelogs.
- Added optional `reason` parameters to `Agent.abort` and `AgentSession.abort` APIs.
- Passed abort reasons through interrupt flows into underlying agent cancellation.
- Replaced hard-coded abort text with `resolveAbortLabel` for streaming and replayed messages.
- Fell back to generic `Request was aborted` text when no abort reason was provided.
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
When a tool call (most visibly `write` with >~1000 lines of content) is
truncated by `stop_reason: length` — e.g. OpenCode Zen's claude-3-5-haiku
with its 8192 `max_tokens` cap — the agent loop correctly refuses to
execute the call (its streamed arguments are mid-string) but used to
attach the same generic placeholder result it uses for non-runnable
non-tool turns: "Tool call was not executed because the assistant ended
its turn." The auto-continue loop re-prompted, the model re-emitted the
same oversized payload, and the user saw the file write fail again and
again — perceived as a "write tool crash" with the target file lost.
`createAbortedToolResult` now takes a `length` reason that names
`stop_reason: length` and tells the model to split the work into multiple
smaller tool calls (write the first chunk, append the rest with `edit`
insert ops, or break the file into multiple `write` targets). The skip
path in agentLoop forwards the real stop reason so the hint reaches the
model on the very next continuation. Tool execution is still guarded —
the truncated args never run.
Regression test in packages/agent/test/agent-loop.test.ts verifies the
synthetic `write` call is skipped AND the resulting tool-result message
carries the length-specific guidance.
Fixes#1785
- Added `maxToolCallsPerTurn` support to `AgentOptions` and `AgentLoopConfig`, with Agent getter/setter and serialized state wiring.
- Implemented stream-loop cap handling by normalizing bad values and halting after `toolcall_end` reaches the limit.
- Added `ANTHROPIC_TOOL_CALL_BATCH_CAP`=8 and wired session cap sync on init, model changes, and restore.
- Added tests that truncated a 10-call stream to 8 tool calls, and verified non-Claude models resolve no cap.
- Fixed agent loop abandoning tool_use blocks on `stop`/`end_turn` turns; only `length` (truncation) now skips execution.
- Verified against live Anthropic API: stop_reason is never replayed on wire and doesn't gate continuation validity.
- Added tests pinning the wire-safety contract: thinking blocks with stale/missing signatures are downgraded to text before replay.
- Updated the agent loop to execute tool calls only when the assistant stop reason was `toolUse`.
- Added skipped placeholder `tool_result` messages for leftover `toolCall` blocks when a turn ended without `toolUse`.
- Promoted OpenAI/Ollama `stop` tool-call turns to `toolUse` and stripped thinking signatures on abandoned tool-use turns during message transforms.
- yieldIfDue() uses compensated sleep (sleepAtLeast): retries Bun.sleep()
until the requested wall-clock duration has elapsed. This is necessary
because napi callbacks (uv_async_send) can wake the event loop
prematurely, causing Bun.sleep(N) to return after only ~1-2ms.
- ExponentialYield for bash-executor: starts at 20ms, doubles to 10s.
Closes#1384
- Exported `normalizeTools` so `AppendOnlyContext` uses the same tool normalization as the agent loop.
- Added `BuildOptions.intentTracing` to `build()`/`reset()`/`takeSnapshot()` so intent injection is consistent and included in the prefix fingerprint.
- Improved `#computeDigest` to cover tool_calls, tool_call_id, name, and id fields to catch in-place mutations.
- Fixed `#unsubscribeAppendOnly` leak and added no-op guard in `#syncAppendOnlyContext`.
ImmutablePrefix caches system prompt + tool specs after first build()
so subsequent turns reuse identical byte sequences. AppendOnlyLog
converts messages once via syncMessages() and only appends deltas
on further turns — prior-turn bytes stay stable.
- New module: packages/agent/src/append-only-context.ts
StablePrefix, AppendOnlyLog, AppendOnlyContextManager
- AppendOnlyContextManager added to AgentLoopConfig
- Wired into streamAssistantResponse in agent-loop.ts
- Toggleable via provider.appendOnlyContext setting (auto/on/off)
- Default auto enables for deepseek provider
- 38 tests covering prefix, log, sync, compaction handling
- /session info surfaces current active state
- Added optional `onBeforeYield` configuration and `setOnBeforeYield` in Agent, executed before follow-up checks.
- Added `YieldQueue` to `AgentSession`, with setup/teardown and streaming/idle flush via `setOnBeforeYield`.
- Replaced immediate async-result follow-up dispatch with queued batch entries, including stale-state suppression.
- Added MCP follow-up queueing in SDK, deduplicating updates by `serverName` and `uri`.
- Added changelog entries for `onBeforeYield`, async-result batching, MCP dedupe, and `display.shimmer` modes.
- Added yield queue unit tests for streaming emission, debounced idle batches, stale filtering, and error isolation.
- Added MockResponse metadata fields and invoked onResponse pre-stream with lowercased headers, status fallback, and requestId.
- Wrapped request onResponse in agent-loop, captured response headers, and forwarded them with baseUrl to finish/fail span handling.
- Added detectGatewayFromHeaders export and pi.gen_ai.gateway.* span attributes via header-based gateway detection.
- Extended telemetry event/span payloads with responseHeaders and validated detection-priority and onResponse forwarding in tests.
- Skipped tool double-count: runTool's interrupt early-return no longer records the skipped tool inline; the post-batch tail sweep now handles accounting once per record so a single queued steering cancellation no longer logs N+1 skips.
- Aborted/errored assistant messages with embedded tool calls now record a collector orphan with status 'aborted' or 'error', so coverage.toolsInvoked and tools counters reflect them.
- run-collector chat record now stores inputTokens = input + cacheRead + cacheWrite, matching ChatUsageEvent and the public AgentRunSummary contract.
- onSpanStart and onSpanEnd hook invocations are wrapped in safeOnSpanStart/safeOnSpanEnd; thrown user callbacks surface via onTelemetryWarning (on_span_start_failed / on_span_end_failed) instead of leaking through finishChatSpan/finishExecuteToolSpan/finishInvokeAgentSpan/recordHandoff.
- summarizeTelemetryValue gained a depth+ancestor guard for arrays (matching the existing object recursion guard); cycles return '[Circular]' and over-depth returns the bounded {kind:'array',length} sentinel.
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
- Swapped deprecated telemetry keys for OpenAIAttr/PiGenAIAttr constants, including tool intent keys.
- Normalized provider names via mapProviderNameToOtel and emitted OTEL pi.gen_ai fields on chat/request/response spans.
- Revised usage and aggregate telemetry to include cache-token totals plus failed and skipped step counts in run summaries.
- Updated OTEL and run-summary tests to use z.object schemas and renamed GenAI/PiGenAI attribute assertions.
- Added `onChatUsage` telemetry configuration and `ChatUsageEvent` payload so each chat step with usage now emits a usage event without requiring a cost estimator.
- Updated chat span completion paths to await async chat usage emission, and emitted `on_chat_usage_failed` warnings when callbacks or hooks reject.
- Awaited `finishChatSpan` in assistant-stream completion and abort flows so telemetry callbacks run before returning from chat completion.
- Added `runEnded`/`markRunEnded()` tracking and only fired `telemetry.onRunEnd` once per run.
- Added `agent_end` telemetry support for per-run `telemetry`/`coverage` and `agentLoopDetailed()` with `detailed()`.
- Added `aggregateAgentRunSummaries`/`aggregateAgentRunCoverage` and mapped `execute_tool` outcomes to `blocked` and `skipped`.
- Updated `finishInvokeAgentSpan` to derive failure `error.type`/status text from run status and exception state.
- Added run-summary test helpers covering `agent_end`, aggregation, and `onRunEnd` warning/compatibility scenarios.
- Added `run-collector` exports, `agentLoopDetailed`, and `agentLoopContinueDetailed` APIs.
- Expanded `agent_end` event payloads with optional `telemetry` and `coverage` fields.
- Added `AgentRunCollector` span tracking with typed chat/tool records and summary/coverage builders.
- Fixed telemetry totals to include interrupted, skipped, and failed tool/chat paths via `failChatSpan` and skip recording.
- Added opt-in telemetry configuration to Agent and session APIs, including Agent#setTelemetry mutator.
- Implemented OpenTelemetry spans for invoke_agent, chat, execute_tool, and handoff paths with metadata and step tracking.
- Added a telemetry helper module, OpenTelemetry request/usage types, and dependency wiring with no-op behavior when tracer SDK is absent.
- Added OTEL end-to-end tests and fixed coding-agent OutputSink realignment and artifact-link newline output issues.
Mirrors the pi-mono API surface so apps can preflight tool execution
(block or mutate validated args) and post-process tool results
(override content/details/isError) without wrapping tools.
- `AgentLoopConfig.beforeToolCall` runs after argument validation. Return
`{ block: true, reason }` to short-circuit with a tool-error result.
Mutations to `context.args` are forwarded to `tool.execute` without
revalidation, matching pi-mono semantics.
- `AgentLoopConfig.afterToolCall` runs after execution and before
`tool_execution_end` / tool-result message emission. Returned fields
override the executed result; omitted fields fall through. Hook
exceptions surface as tool errors and do not abort the batch.
- `Agent` exposes both hooks as public, reassignable fields so extension
reloads can swap implementations mid-session.
Compatibility: fully additive. Both hooks default to undefined and the
loop behaves identically when neither is set. The internal
`executeToolCalls` signature was collapsed to `(context, message,
signal, stream, config)` -- it is not exported, so this is not a public
API change. Pi-mono's `terminate` field on `AfterToolCallResult` is
omitted because our `AgentToolResult` has no batch-level early-stop
contract.
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
- Stored a failure counter during single-path edit execution and set isError on aggregate results when any entry edit failed.
- Set streaming-edit handling to always evaluate auto-generated-file checks, but only primed the file cache when edit.streamingAbort was enabled.
- Replaced `===== ... =====` eval cell headers with `*** Begin ` / `*** End ` markers; legacy format remains renderable in HTML exports.
- Replaced hashline patch grammar with `*** Begin Patch` / `*** End Patch` envelope; old inputs without the envelope are still accepted.
- Extracted `sniffEvalLanguage` into a shared `sniff.ts` module reused by the parser and tool.
- Added `docs/ERRATA-GPT5-HARMONY.md` and `scripts/session-stats/harmony_backtest.py` documenting and backtesting the GPT-5 Harmony-header leak defect.
Anthropic counts sessions by metadata.user_id. Without this fix, OMP
generated fresh random entropy on every API request, inflating the
session count and preventing backend attribution to the authenticated
account.
Changes:
packages/ai:
- resolveAnthropicMetadataUserId() now accepts JSON-format user_id
matching real Claude Code's getAPIMetadata shape
({ session_id, account_uuid, ... }). Previously only the legacy
cloaking format was accepted on OAuth, causing stable caller-supplied
values to be silently discarded.
- AnthropicOAuthFlow.exchangeToken() and refreshAnthropicToken() now
populate OAuthCredentials.{accountId, email} from the token response
account block, removing the need for a separate /api/oauth/profile
round-trip.
- AuthStorage.getOAuthAccountId(provider, sessionId) returns the OAuth
accountId for the session-sticky credential, used to build
account_uuid in metadata.user_id. Guards against misattribution for
API-key, runtime-override, env-key, and fallback-resolver paths that
do not record a session credential.
packages/agent:
- Agent.metadataForProvider(provider) resolves request metadata for
the given provider via the installed resolver, or returns the static
metadata value. The plain metadata getter now returns only the static
value; provider-aware resolution is explicit.
- Agent.setMetadataResolver(fn) installs a (provider: string) resolver
evaluated per LLM request in agent-loop, after getApiKey records the
session-sticky credential, so account_uuid reflects the credential
actually used.
- AgentLoopConfig.metadataResolver is called with config.model.provider
after getApiKey, overriding the static metadata field.
packages/coding-agent:
- AgentSession.#syncAgentSessionId installs a metadata resolver that
builds { user_id: JSON.stringify({ session_id, account_uuid? }) },
matching the Anthropic session attribution format. account_uuid is
only included for provider="anthropic" to avoid leaking the OAuth
identity to third-party Anthropic-format-compatible providers.
- sessionId getter prefers providerSessionId when supplied via
AgentSessionConfig so all API paths (getApiKey, direct calls,
metadata resolver) share the same provider-facing session ID.
- prepareSimpleStreamOptions stamps session metadata on direct calls
(runEphemeralTurn, compaction, branch summary, title generation) so
they share the same session bucket as Agent.prompt requests.
- generateBranchSummary and generateSessionTitle accept a
(provider: string) metadata resolver evaluated after their own
getApiKey call for correct credential attribution.
- Added a boundary normalizer that validated tool output shape and replaced malformed responses with a fallback text-only result.
- Updated tool execution to coerce both streaming partial updates and final tool results through that normalizer.
- Set the tool call error state when a malformed result was detected by the coercer.