RESCUE_SHAKE_CONFIG spreads AGGRESSIVE_SHAKE_CONFIG, so the new 4k manual
tail leaked into dead-end recovery and could block eliding the very
result that caused the dead end. Override protectTokens back to 0 in the
rescue preset and add the coding-agent changelog entry for the manual
/shake behavior change.
Addresses review on #8067.
Manual /shake used protectTokens: 0, stripping every eligible tool result
including the ones the agent is still working from. Keep a small 4k-token
recent window (matching the automatic shake mechanism, at a quarter of its
16k budget) so the full escape hatch stays aggressive without destroying
the live tail. Two matcher-focused tests that implicitly relied on the
zero window now pin protectTokens: 0 explicitly.
Fixes#7776
- Reused the Responses replay lifecycle policy for V1 and V2 compaction input.
- Covered persisted native history and converted assistant replay items.
Fixes#7742
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
A peer-IRC interrupt (e.g. a subagent message) aborts only interruptible
waits and leaves already-running non-interruptible foreground work alone.
But runTool's `interruptState.triggered` early-return skipped every
not-yet-started tool regardless of source, so a non-interruptible tool
queued behind an interruptible wait in the same batch (a batched todo/write
after `hub wait`) was dropped with "Skipped due to pending peer interrupt".
Exclude non-interruptible tools from the early skip on the IRC path; user
and system steering still preempt all queued work.
Fixes#7493
Publish terminal daemon completions to the session that started the
process so idle agents can resume without polling hub status.
Persist every unacknowledged generation with a stable completion ID and
immutable snapshot. Replay the collection after reconnect or broker
recovery, and clear each event only after the owning client acknowledges
it.
Signed-off-by: Christian Stewart <christian@aperture.us>
- Added shared Python call and literal serialization utilities with multiline verbatim support.
- Standardized tool inventories to format as an OpenAI-Harmony functions namespace using TypeScript declarations.
- Updated tool normalization and rendering functions to accept options objects and default to Python-syntax examples.
- Refactored Gemini dialect rendering to leverage shared serialization functions directly.
- Added `isStreamEnvelopeErrorText` to packages/ai/src/error/flags.ts to recognize stream envelope truncation errors.
- Updated `streamAnthropicOnce` in packages/ai/src/providers/anthropic.ts to throw an envelope error when streams die mid-generation without a terminal frame.
- Updated `recoverTransientErrorToolTurn` in packages/agent/src/agent-loop.ts to recognize Anthropic stream envelope truncation errors for tool call salvage.
- Applied response.metadata turn-state/models-etag refreshes while draining Codex V2 compaction WebSocket events.
- Deferred the update until the WebSocket attempt succeeds so a discarded attempt cannot leak turn state.
- Added coverage for a mid-turn compaction refreshing x-codex-turn-state.
Fixes#7198
Track entry into tool.execute separately from tool event emission. Never-started skips retain SyntheticToolResultDetails with executed:false; in-flight aborts now use distinct interrupted metadata with execution:started so consumers do not assume no partial work occurred.
Keep both interrupt states neutral in the TUI and cover the agent metadata boundary plus rendering behavior.
Fixes#7199
- Buffered WebSocket compaction events until terminal completion.
- Discarded buffered events when transport failure triggers an SSE replay.
- Added coverage for failure after a partial compaction output item.
Fixes#7198
- Reused the live Codex provider session for WebSocket-first V2 compaction.
- Fell back to SSE V2 on WebSocket transport failure before the existing V1 fallback.
- Propagated the configured WebSocket preference through manual, automatic, and advisor compaction paths.
- Added transport reuse and fallback regression coverage.
Fixes#7198
The summary budget is floor(0.8 * reserveTokens) and the effective reserve is at
least 15% of the declared context window, so window size alone decided how long
a summary was allowed to be: a 1M-token window authorized roughly 120k tokens of
summary. At that size the model copies the conversation instead of compressing
it, on the slowest and most expensive token class, so compaction got worse as the
window got bigger.
Clamp with MAX_SUMMARY_TOKENS, set to DEFAULT_RESERVE_TOKENS so no new tuning
constant is introduced. Reserves below the cap are unaffected.
- Added `prepareToolCallDispatch` and `PreparedToolCall` to handle argument validation and `beforeToolCall` before message snapshotting.
- Implemented `preparedDispatchByMessage` WeakMap to store pre-dispatch results for streamed messages.
- Updated `executeToolCalls` to consume pre-computed dispatch preparation results.
- Updated documentation and changelog to specify `beforeToolCall` timing on the streamed path.
- Added a prepareToolCall phase to the agent loop running before tool scheduling for validation and hooks.
- Updated BeforeToolCallContext and result types to support argument replacement instead of in-place mutation.
- Updated coding-agent extension handling and runner to track emitted tool calls and re-evaluate approvals on input revisions.
- Added comprehensive test coverage for argument replacement, concurrency resolution, and schema validation.
- Cleared the retained soft-requirement lifecycle alongside the deferred
hard choice: clearDeferredToolDirectives() owns both, is called from
clearAllQueues/reset and session-scoped tool-state cleanup, with a
regression covering reminder re-injection after a queue clear.
- Allowed void-returning pre-model gates via the named AgentBeforeModelCall
type and normalized gate results in the loop and Agent dispatcher.
- Documented that the first gate installed mid-run applies from the next
run; corrected the onToolChoiceRejected contract docs; documented the
cross-run lifetime of ToolChoiceQueue's in-flight claim.
- Removed the unused addBeforeModelContextBuild hook.
- Relocated both packages' changelog entries out of the released 17.1.4
sections into Unreleased with PR attribution, folding the never-shipped
Fixed bullet into Added and noting the input-event timing change.
A pre-model gate can defer a claimed hard tool choice for the next call. Branch transitions cleared the coding-agent queue but left that agent-owned value alive, allowing an obsolete forced tool to cross into the replacement transcript.
Expose the narrow deferred-choice reset at the Agent owner and invoke it from the shared session-scoped tool-state cleanup used by both branch paths. Failed session switches retain their existing rollback behavior.
Signed-off-by: Christian Stewart <christian@aperture.us>
A Harmony retry keeps its logical turn open while the next provider call is prepared and gated. An abort during that gate previously ended the agent stream directly, leaving observers with an unmatched turn_start event.
Route the aborted gate through the existing pre-model stop owner so it emits the synthetic aborted message and closes the open turn before ending the stream. Fresh turns retain their existing no-provider-call cancellation behavior.
Signed-off-by: Christian Stewart <christian@aperture.us>
The agent loop had no place to refuse a provider request. A host that needs to
act on the assembled context before it is billed, checking that the prompt still
fits the window, that a budget boundary has not been crossed, or that the
session should hand off instead of spending, could only observe the request
after the fact, when the tokens were already committed.
Add `AgentLoopConfig.beforeModelCall`, asked once per turn beside the deadline
check and before `turn_start` is emitted. A `stop` result ends the stream with
no turn open, so nothing has to synthesise a cancellation event and no consumer
is left holding a half-open turn. Placing it there also keeps `turn_end`'s
contract intact: that event carries the assistant message for a completed turn,
and a gated stop has no assistant message to report.
`syncContextBeforeModelCall` keeps its existing void contract and its job of
refreshing prompt and tool state, so implementations typed as returning void are
unaffected.
`Agent.setBeforeModelCall` installs the host's callback, and `addBeforeModelCall`
registers an additional callback without displacing the host's, returning a
disposer so an extension can attach and detach independently. A supplied
`reason` is logged where the loop stops.
Signed-off-by: Christian Stewart <christian@aperture.us>
Made finalized terminal content optional and retained the delta-reconstructed assistant content when older proxy servers omit it.
Covered both legacy done and error events so terminal cleanup remains safe and the stream resolves.
Fixes#6703
Made proxy done and error events carry finalized assistant content so provider-only blocks survive event reconstruction.
Covered Anthropic native web-search history restored after live events omitted the opaque blocks.
Fixes#6703
Two boundary defects on the mirrored-todo path.
1. The Agent error drain snapshotted #cursorToolResultBuffer without
awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
async cursorOnToolResult still running when the provider errored
patched an entry the catch path had already detached, so the
pre-transform payload was persisted. A provider error is exactly when
a transform is most likely to be in flight.
2. The todo renderer interpolated mirrored provider text straight into
terminal output. A Cursor snapshot carries model-authored task
content, phase names and summary text verbatim, so a label holding
ANSI/C0 sequences rewrote the terminal on every render and replay.
sanitizeText alone is not enough - it preserves tabs, which punch
holes in bordered output - so every display path now funnels through
one forDisplay() helper: task labels, blocker notes, phase headers,
the zero-task fallback, and the streaming renderCall preview. Raw
values are untouched; content and phase name are the identity keys
the local list is looked up by and what gets persisted.
Three orphan paths, same failure mode: the assistant block is marked
kCursorExecResolved before the work runs, so agent-loop.ts emits no
placeholder for it, and any path that produces no toolResult leaves the
call unpaired — buildSessionContext then strips the whole interaction
from every rebuilt transcript.
1. resolveExecHandler returned no toolResult on three exits (no handler
installed, handler produced nothing, handler threw). Each now pairs a
result carrying the same text the server sees in execResult, routed
through onToolResult like a real one. `pairing` is a required
parameter so a new callsite cannot silently recreate the orphan.
2. Agent only installed its result-buffer sink when cursorExecHandlers
or cursorOnToolResult was set. Both are optional, so a bare SDK host
dropped the provider result on the floor. Installed unconditionally;
a non-Cursor provider never calls it.
3. A todo completion frame with no tool_call (the field is optional)
skipped settlement entirely. It now settles as "nothing to mirror".
Also fixes an empty update_todos with a nonzero total_count being
mirrored as an authoritative clear: the length guard added earlier
skipped the mismatch check for empty responses, so a partial or
size-limited merge response deleted every local task at once.