- Added compaction.asyncEnabled (Async Compaction, default on): when
context enters the pre-threshold band [threshold - lead, threshold)
with lead = clamp(threshold * 0.125, 8192, 32000), maintenance
speculatively summarizes in the background off a branch snapshot
(first configured LLM-backed method: remote, handoff, or soft) using
a side session id isolated from the live turn. Crossing the threshold
splices the armed result in instantly instead of blocking on a
summarization round-trip. Armed results are invalidated by branch
changes, reset boundaries, model switches that strand provider-native
replay payloads, and context growth past keepRecentTokens (which
re-speculates); extensions registering session_before_compact keep
exact blocking semantics (speculation disabled).
- Reworked handoff to commit in place: /handoff and the auto handoff
method now write the generated document as a regular compaction entry
on the current session (summary = document + <files> tag, cut from
prepareCompaction) instead of starting a new session. SessionHandoff
shrank to a document generator; session_before_switch/session_switch
no longer fire with reason "handoff"; mid-turn maintenance no longer
suppresses the handoff preference; overflow recovery can apply an
armed handoff result.
- Extracted the shared auto-compaction commit tail
(#commitAutoCompactionResult / #commitCompactionEntry) used by the
blocking production path, the armed speculative apply, manual
compaction, and manual handoff.
- Status line pulses the auto-compact icon while a speculation runs and
holds it in accent once a result is armed.
- Exported remotePreserveReusable from pi-agent-core/compaction for
apply-time validation of speculative remote results.
- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.
buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.
Fixes#8789
prepareCompaction walked the branch from the last compaction and ignored reset_boundary markers, so /compact (and auto-compaction) resurrected pre-/clear turns into the summary even though buildSessionContext already starts the model context after the boundary.
Model reset_boundary as a first-class agent-core session entry and start the summarization window after the latest boundary, dropping the superseded pre-reset compaction summary. A boundary before the last compaction stays superseded by it.
Fixes#8718
Each is a single, side-effect-free completion whose result is parsed after
it resolves, so one transient provider failure previously aborted the
whole operation - for /compact that left the user's context full.
Adds SummaryOptions.oneshotRetry, because both compaction paths call the
same generateSummary and the policy therefore cannot be a constant inside
it. Manual /compact has no outer loop and gets retry by default;
auto-compaction passes false because session-maintenance already retries
the whole attempt, and nesting would multiply the budget (10 outer x 3
inner) while stacking each outer wait on an inner backoff.
Cursor bash/grep frames wrote optional kwargs as present-undefined, and
tools.format gemini projectors dropped kCursorExecResolved so settled
calls ran twice.
Co-authored-by: Cursor <cursoragent@cursor.com>
Manual /shake used protectTokens: 0, stripping every eligible tool result
including the ones the agent is still working from. Keep a small 4k-token
recent window (matching the automatic shake mechanism, at a quarter of its
16k budget) so the full escape hatch stays aggressive without destroying
the live tail. Two matcher-focused tests that implicitly relied on the
zero window now pin protectTokens: 0 explicitly.
Fixes#7776
- Reused the Responses replay lifecycle policy for V1 and V2 compaction input.
- Covered persisted native history and converted assistant replay items.
Fixes#7742
- Union changelog merges interleaved stale pre-17.2.5 PR-branch entries into
released sections; released bodies are restored byte-for-byte from the
pre-merge main state.
- [Unreleased] now carries exactly the entries for PR #7080 and the nine
merged fixes (#7495, #7460, #7466, #7468, #7473, #7481, #7477, #7368, #7453).
A peer-IRC interrupt (e.g. a subagent message) aborts only interruptible
waits and leaves already-running non-interruptible foreground work alone.
But runTool's `interruptState.triggered` early-return skipped every
not-yet-started tool regardless of source, so a non-interruptible tool
queued behind an interruptible wait in the same batch (a batched todo/write
after `hub wait`) was dropped with "Skipped due to pending peer interrupt".
Exclude non-interruptible tools from the early skip on the IRC path; user
and system steering still preempt all queued work.
Fixes#7493
- Added shared Python call and literal serialization utilities with multiline verbatim support.
- Standardized tool inventories to format as an OpenAI-Harmony functions namespace using TypeScript declarations.
- Updated tool normalization and rendering functions to accept options objects and default to Python-syntax examples.
- Refactored Gemini dialect rendering to leverage shared serialization functions directly.
- Added `isStreamEnvelopeErrorText` to packages/ai/src/error/flags.ts to recognize stream envelope truncation errors.
- Updated `streamAnthropicOnce` in packages/ai/src/providers/anthropic.ts to throw an envelope error when streams die mid-generation without a terminal frame.
- Updated `recoverTransientErrorToolTurn` in packages/agent/src/agent-loop.ts to recognize Anthropic stream envelope truncation errors for tool call salvage.
Track entry into tool.execute separately from tool event emission. Never-started skips retain SyntheticToolResultDetails with executed:false; in-flight aborts now use distinct interrupted metadata with execution:started so consumers do not assume no partial work occurred.
Keep both interrupt states neutral in the TUI and cover the agent metadata boundary plus rendering behavior.
Fixes#7199
A tool call aborted mid-batch to service queued steering/peer input emits a synthetic placeholder result with isError:true so the model retries it. The TUI keyed all error styling (red ✘, red frame/text) off that flag, so a normal steering skip rendered identically to a real tool failure.
Mark the skip placeholder with the existing SyntheticToolResultDetails discriminator (source: "interrupt_skipped", executed:false) and render benign skips through the neutral generic card (info glyph, dim text, neutral background), bypassing any bespoke error frame. Genuine failures keep their error styling.
Fixes#7199
- Reused the live Codex provider session for WebSocket-first V2 compaction.
- Fell back to SSE V2 on WebSocket transport failure before the existing V1 fallback.
- Propagated the configured WebSocket preference through manual, automatic, and advisor compaction paths.
- Added transport reuse and fallback regression coverage.
Fixes#7198