- Added T co-signal requirement for tool_arg surface so legitimate edits carrying the marker are never hard-aborted.
- Extended detectHarmonyLeakInAssistantMessage with optional toolArgParseEnd resolver; agent loop omits it, keeping tool_arg inert.
- Updated tests to inject a boundary-at-0 helper for corpus cases and added T-gate unit tests.
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
- Replaced flat `protectedTools: string[]` with `ProtectedToolMatcher[]` supporting predicate functions.
- Regular file/URL `read` calls are now eligible for pruning and shake compaction.
- `read` calls whose `path` starts with `skill://` remain protected like native `skill` results.
- Added `collectToolCallsById` to correlate tool results with their originating call arguments.
F3 (agent): mutate state.messages and state.pendingToolCalls in place on
appendMessage/popMessage/clearMessages/reset/tool_execution_start/_end
instead of allocating a fresh array/Set on every transition. Subscribers
that capture state.messages by reference now observe updates directly.
Public type signature unchanged.
F5 (ai): add parseStreamingJsonThrottled to utils/json-parse — a per-delta
wrapper around parseStreamingJson that skips the re-parse until the
buffer has grown by minGrowthBytes (default 256). Wired into every
provider's tool-call argument accumulator (anthropic, amazon-bedrock,
openai-completions, openai-codex-responses, openai-responses-shared) so
per-delta cost becomes O(N) in total buffer length instead of O(N²).
Every provider's toolcall_end still runs a final unthrottled parse, so
the published block.arguments is unchanged.
F8 (coding-agent): drop the per-delta structuredClone of streaming tool
arguments in ToolExecutionComponent.updateArgs. event-controller.ts and
ui-helpers.ts already spread their input into a fresh object on each
delta, so cloning here was dead work on the rendering hot path. Added a
reference-equality short-circuit so repeat calls with the same args
object skip the preview-diff and display refresh.
- Added `maxToolCallsPerTurn` support to `AgentOptions` and `AgentLoopConfig`, with Agent getter/setter and serialized state wiring.
- Implemented stream-loop cap handling by normalizing bad values and halting after `toolcall_end` reaches the limit.
- Added `ANTHROPIC_TOOL_CALL_BATCH_CAP`=8 and wired session cap sync on init, model changes, and restore.
- Added tests that truncated a 10-call stream to 8 tool calls, and verified non-Claude models resolve no cap.
- Fixed agent loop abandoning tool_use blocks on `stop`/`end_turn` turns; only `length` (truncation) now skips execution.
- Verified against live Anthropic API: stop_reason is never replayed on wire and doesn't gate continuation validity.
- Added tests pinning the wire-safety contract: thinking blocks with stale/missing signatures are downgraded to text before replay.
- Updated the agent loop to execute tool calls only when the assistant stop reason was `toolUse`.
- Added skipped placeholder `tool_result` messages for leftover `toolCall` blocks when a turn ended without `toolUse`.
- Promoted OpenAI/Ollama `stop` tool-call turns to `toolUse` and stripped thinking signatures on abandoned tool-use turns during message transforms.
- Centralized compaction stop-reason error throws via createSummarizationError().
- Set compaction thrown errors to copy response.errorStatus into Error.status.
- Expanded compaction auth detection to treat HTTP 401/403 as auth failures with regex fallback preserved.
- Added regression tests for 401/403 status propagation and compaction fallback auth behavior.
- Documented both package fixes in Unreleased Fixed changelog entries.
- Replaced the `keepaliveWhile` Promise wrapper with a new `EventLoopKeepalive` class that registers and disposes an interval timer through `Symbol.dispose`.
- Updated `Agent` to instantiate `EventLoopKeepalive` via `using` during prompt execution instead of manually managing an interval.
- Wrapped interactive mode's await path with the new helper and removed redundant `keepaliveWhile` usage from the CLI entrypoint.
The previous commit accidentally dropped the try/catch in #emit
because the local agent.ts was based on an older main that lacked
the listener-isolation code. Restore it to match main.
EventLoopKeepalive is not exported from yield.ts on main.
Use setInterval + unref() directly, matching the pattern in
keepaliveWhile(). Biome import order also fixed (type imports
before value imports within the same group).
Biome organizeImports sorts type imports before value imports within
the same relative-path group. Move EventLoopKeepalive import after
all type imports.
Bun 1.3.x event loop busy-waits when the only pending work is an
unresolved Promise. Agent.prompt() sets #runningPromise via
Promise.withResolvers() which stays unresolved during the entire
agent loop execution (LLM calls + tool iterations), causing ~100%
CPU even when the process is idle.
PR #1419 added keepaliveWhile() to getUserInput() in main.ts, but
session.prompt() callers (interactive mode, resume, etc.) still
await the unresolved #runningPromise, bypassing the keepalive.
Install EventLoopKeepalive directly in Agent.prompt() so all
callers are covered. Dispose in the finally block after the agent
loop completes.
The thinking-level fix (e07b47ee4) added SummaryOptions.thinkingLevel
and threaded it from agent-session.ts into compact(), but the
field-by-field rebuild of summaryOptions inside compact() (and a
second inline rebuild for generateShortSummary) silently dropped it.
Effect on every call site that fans through compact():
generateSummary, generateTurnPrefixSummary, generateShortSummary all
see options?.thinkingLevel === undefined => resolveCompactionEffort
falls back to Effort.High => user's /model :off selection is
silently overridden, and xai-oauth/grok-build still trips on the
unsupported-effort path even though fix#2 strips it at the wire
layer of the openai-responses mapper.
Add `thinkingLevel: options?.thinkingLevel` at both rebuild sites in
compact(): the summaryOptions literal feeding generateSummary and
generateTurnPrefixSummary, and the inline options literal feeding
generateShortSummary.
Extend compaction-thinking-level.test.ts with four compact()-level
cases driving isSplitTurn:true so all three summarizers fire:
- Off -> every fan-out call gets reasoning=undefined
- Low -> every fan-out call gets reasoning="low"
- <unset> -> every fan-out call gets reasoning="high" (default)
- grok-build + High -> every fan-out call gets reasoning=undefined (clamp)
TDD red-green verified: stashing the source fix flips the Off and Low
cases to fail with received="high" (exactly the reviewer's prediction);
restoring the fix returns all four to green. Suite: 131 pass / 0 fail
(baseline 127 + 4 new). biome + tsgo --noEmit clean.
Op: correct
Restores: spec:compaction-honors-session-thinking-level
(cherry picked from commit 9b501e369b820cb992345ceb3f2a207ccc9ee233)
Triple-stacked failure on the same axis (thinking effort) produced the
user-visible
Error: Compaction failed: Thinking effort high is not supported by
xai-oauth/grok-build.
Supported efforts:
(empty list after the colon) whenever the active model was a curated
xAI catalog entry with compat.supportsReasoningEffort: false.
Three defects lined up. (1) Behavior: compaction at four call sites
in packages/agent/src/compaction/compaction.ts hardcoded
reasoning: Effort.High and never threaded session.thinkingLevel —
the user's /model :off selection (and any explicit low/medium) was
silently overridden. On every other model this was invisible.
(2) Validation: requireSupportedEffort threw at the openai-flavored
mapper layer before the wire-side omitReasoningEffort gate in
providers/xai-responses.ts ever ran; two contradictory guards on the
same wire param. (3) Message: when getSupportedEfforts returned [],
the rendered error tail was 'Supported efforts: ' with nothing after
the colon — disappears as a side-effect of fix#2.
Fix#1 — thread ThinkingLevel | undefined end-to-end. Add
SummaryOptions.thinkingLevel and HandoffOptions.thinkingLevel.
Convert via a single exhaustive switch (effortFromThinkingLevel) in
the new resolveCompactionEffort helper:
- Off → undefined (omit reasoning entirely)
- undefined/Inherit → Effort.High → clamp per model (preserves the
historical default for users
who never touched the dial)
- explicit Effort → respect user → clamp per model
resolveCompactionEffort lives in compaction.ts; all four call sites
(generateSummary, generateHandoff, generateShortSummary,
generateTurnPrefixSummary) route through it. agent-session.ts threads
this.thinkingLevel into all three production compaction entry points
(manual /compact at L6201, auto-compaction at L6458 — the most-fired
path, originally missed in plan review — and direct generateHandoff
at L5465). The audit-gate test
(test/agent-session-compaction-thinking-threading.test.ts) scans the
file with a brace-balanced extractor and refuses any unthreaded site.
Fix#2 — silent-clamp at the openai-flavored mapper layer. Extract
exported modelOmitsReasoningEffort(model) in model-thinking.ts as the
single source of truth for compat.supportsReasoningEffort: false on
openai-responses* APIs. getSupportedEfforts now calls it instead of
inlining the check (pure refactor — observable behavior preserved).
resolveOpenAiReasoningEffort in stream.ts early-returns undefined
when the predicate is true, so the wire-side omitReasoningEffort
gate (providers/xai-responses.ts:78) becomes the single source of
truth for the actual strip — no redundant throw.
Three regression tests pin the contract:
- packages/ai/test/xai-oauth-effort-strip.test.ts (5 tests):
modelOmitsReasoningEffort returns true for grok-build and
grok-4.20-0309-reasoning, false for grok-4.3 / Anthropic /
openai-completions.
- packages/agent/test/compaction-thinking-level.test.ts (5 tests):
every ThinkingLevel outcome through generateHandoff — Off stays
undefined (not coerced to High), Low stays Low, Inherit / undefined
default to High, grok-build clamps to undefined regardless of
requested level. Covers the Codex-caught Off-vs-not-provided
distinction.
- packages/coding-agent/test/agent-session-compaction-thinking-threading.test.ts
(2 tests): brace-balanced source scan asserts every direct
compact() / generateHandoff() in agent-session.ts threads
'thinkingLevel: this.thinkingLevel'; floor of 3 threaded sites.
TDD red-green verified for fix#1: temporarily reverted the handoff
call-site back to hardcoded Effort.High → compaction-thinking-level
went 2 pass / 3 fail (Off coerced, Low overridden, grok-build throws);
restored → 5 pass / 0 fail.
Verified:
- packages/agent: 127 pass / 0 fail
- packages/ai: 1061 pass / 337 skip / 0 fail
- packages/coding-agent (focused): 179 pass / 5 skip / 0 fail
- biome + tsgo --noEmit clean across all three packages
Out of scope (follow-ups):
- branch-summarization.ts:307 already passes no reasoning — no edit.
- The empty-list error message at model-thinking.ts:296 is now
structurally unreachable from the openai-responses path.
- modelOmitsReasoningEffort and grokSupportsReasoningEffort
(xai-responses.ts:22) overlap; collapse into a single predicate
in a future commit.
Op: correct
Restores: spec:compaction-honors-session-thinking-level
Restores: spec:xai-oauth-grok-build-compaction-no-throw
(cherry picked from commit e07b47ee46769053c658819437e2478389a4cee0)
- Imported isPromise from node:util/types in four modules.
- Replaced four instanceof Promise checks with isPromise calls for more reliable promise detection.
- Removed the EventLoopKeepalive class and replaced it with a direct setInterval call.
- Added an unref call to the interval timer to prevent blocking the process exit.
Two TS errors in CI:
1. `agent-session.ts:1317` — `error TS1345: An expression of type 'void'
cannot be tested for truthiness`. The listener type is
`(event: AgentSessionEvent) => void`, so the returned value can't be
directly tested. Same shape in `agent.ts:1079`.
2. `test/session/emit-listener-isolation.test.ts:19` — the test fixture
for `AgentEvent.tool_execution_start` was missing the required `args`
field.
Cast the return to `unknown` and check `instanceof Promise` instead of
duck-typing `.then` — type-safe and matches what async functions actually
return. Add `args: {}` to the test fixture.
Project convention (AGENTS.md) prohibits ReturnType<> — use the
concrete type name instead. NodeJS.Timeout matches the existing
pattern used throughout the codebase (e.g. interactive-mode.ts).
Root cause: Bun 1.3.x (JavaScriptCore) busy-waits when the only
pending work is an unresolved Promise. A setInterval keepalive
keeps the event loop in epoll_wait instead of userspace spinning.
- EventLoopKeepalive: setInterval-based keepalive (re-arms after each
firing, addressing the bot review concern about setTimeout expiry)
- keepaliveWhile(): wrapper to await a Promise with keepalive active
- Applied to getUserInput() in main.ts
- Retains yieldIfDue() and ExponentialYield from #1396
Idle CPU drops from ~100% to ~0% (wchan=do_epoll_wait).
Both `AgentSession.#emit` (session/agent-session.ts) and `Agent.#emit`
(packages/agent/src/agent.ts) iterated listeners with no error isolation.
A synchronous throw in any subscriber aborted the for-loop, so later
subscribers (TUI rendering, ACP bridge, task executor progress,
hindsight) silently missed events. Many listeners — see
`modes/controllers/event-controller.ts:141` and
`modes/controllers/input-controller.ts:576` — are registered as
`async (event) => { await this.handleEvent(event); }`; the returned
Promise was dropped, so any rejection became an unhandled rejection.
Wrap each listener invocation in try/catch and attach a `.catch` to any
returned thenable. Errors are logged via `logger.warn` (already imported
in agent-session.ts) and `console.error` (agent.ts has no logger
dependency, keep it that way).
Test: new `test/session/emit-listener-isolation.test.ts` registers two
listeners on both classes; first listener throws (or returns a rejecting
Promise); asserts the second listener still receives the event AND no
`unhandledRejection` fires. 4 cases (sync+async × Agent+AgentSession).
All fail on current main; all pass with the fix.
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
- Replaced Bun.sleep with scheduler.wait for Node-compatible cancellable sleeps.
- Added module-level timestamp gate to skip yields within 50ms of the last one.
- Threaded AbortSignal through ExponentialYield.sleep to cancel losing timers in race.
- Added tests covering gate behaviour and stray-timer cancellation.
Fire Pass support:
- New provider with login command 'omp /login firepass'
- Hand-seeded kimi-k2.6-turbo model with Fire Pass router wire id
- pi-ai CLI --help lists firepass
AI provider fixes surfaced during the Fire Pass review:
- service_tier whitelist restored for openai/openai-codex only (was leaking to
Fireworks, Firepass, OpenRouter, Azure OpenAI Responses)
- Anthropic tool schema normalizer collapses {} -> true for additionalProperties
- anthropic.prepareParams fires onPayload after drop helpers so callers see the
real wire body
- isServiceTier type guard narrowed to ResolvedServiceTier
- transformMessages stops dropping orphan tool_result when all pending tool
calls have already resolved
- zodToWireSchema preserves null for non-scalar nullable() inner schemas
- isEmptyObject / isJsonObjectEmpty use Object.keys().length === 0 instead of
prototype-walking for...in
- Telemetry records resolved service_tier (priority) instead of scoped
placeholder (openai-only/claude-only)
- Robomp dirty-state reminder no longer asserts a formatter-failure premise
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.
- yieldIfDue() uses compensated sleep (sleepAtLeast): retries Bun.sleep()
until the requested wall-clock duration has elapsed. This is necessary
because napi callbacks (uv_async_send) can wake the event loop
prematurely, causing Bun.sleep(N) to return after only ~1-2ms.
- ExponentialYield for bash-executor: starts at 20ms, doubles to 10s.
Closes#1384
- Exported `normalizeTools` so `AppendOnlyContext` uses the same tool normalization as the agent loop.
- Added `BuildOptions.intentTracing` to `build()`/`reset()`/`takeSnapshot()` so intent injection is consistent and included in the prefix fingerprint.
- Improved `#computeDigest` to cover tool_calls, tool_call_id, name, and id fields to catch in-place mutations.
- Fixed `#unsubscribeAppendOnly` leak and added no-op guard in `#syncAppendOnlyContext`.
ImmutablePrefix caches system prompt + tool specs after first build()
so subsequent turns reuse identical byte sequences. AppendOnlyLog
converts messages once via syncMessages() and only appends deltas
on further turns — prior-turn bytes stay stable.
- New module: packages/agent/src/append-only-context.ts
StablePrefix, AppendOnlyLog, AppendOnlyContextManager
- AppendOnlyContextManager added to AgentLoopConfig
- Wired into streamAssistantResponse in agent-loop.ts
- Toggleable via provider.appendOnlyContext setting (auto/on/off)
- Default auto enables for deepseek provider
- 38 tests covering prefix, log, sync, compaction handling
- /session info surfaces current active state
- Added optional `onBeforeYield` configuration and `setOnBeforeYield` in Agent, executed before follow-up checks.
- Added `YieldQueue` to `AgentSession`, with setup/teardown and streaming/idle flush via `setOnBeforeYield`.
- Replaced immediate async-result follow-up dispatch with queued batch entries, including stale-state suppression.
- Added MCP follow-up queueing in SDK, deduplicating updates by `serverName` and `uri`.
- Added changelog entries for `onBeforeYield`, async-result batching, MCP dedupe, and `display.shimmer` modes.
- Added yield queue unit tests for streaming emission, debounced idle batches, stale filtering, and error isolation.
The OpenAI Responses compactor previously rewrote any empty tool result
into "(see attached image)" regardless of whether the result actually
contained an image block. Empty file reads now persist as empty strings.
- Migrated per-object caches (chat/tool starts, model fingerprints, validation contexts, provider indexes, render IDs) from WeakMap to Symbol-keyed properties on the objects themselves.
- Rewrote SSE debug tee as a single-pass inline parser, eliminating the body.tee() + readSseEvents re-parse pipeline.
- Refactored MockModel from a factory function + external WeakMap state into a self-contained class.
- Added FIFO memoization caches for heuristic candidate expansion and namespace suffix lookups.
- Added optional AgentTelemetry to summary, handoff, branch-summary, and compact option types.
- Replaced one-shot `completeSimple` usage with `instrumentedCompleteSimple` across compaction, summary, and branch-summary calls and passed `oneshotKind`.
- Added `PiGenAIAttr.OneshotKind`, `InstrumentedChatSpanOptions`, and response-header forwarding in telemetry span lifecycle.
- Added `resolveTelemetry` propagation in coding-agent session and inspect-image paths to pass request-scoped telemetry.
- Added compaction telemetry test harness and span assertions for success, no-telemetry, and error cases.
- Added MockResponse metadata fields and invoked onResponse pre-stream with lowercased headers, status fallback, and requestId.
- Wrapped request onResponse in agent-loop, captured response headers, and forwarded them with baseUrl to finish/fail span handling.
- Added detectGatewayFromHeaders export and pi.gen_ai.gateway.* span attributes via header-based gateway detection.
- Extended telemetry event/span payloads with responseHeaders and validated detection-priority and onResponse forwarding in tests.