- Added `CompactionSummaryMessageOptions` interface to support optional fields and metadata in compaction summary messages.
- Included compaction method labels and before-after amount badge rendering in compaction views.
- Updated session maintenance and compaction methods to track and compute post-compaction token counts.
- Updated calls across agent and coding-agent packages to use the new compaction options object.
- Added compaction.asyncEnabled (Async Compaction, default on): when
context enters the pre-threshold band [threshold - lead, threshold)
with lead = clamp(threshold * 0.125, 8192, 32000), maintenance
speculatively summarizes in the background off a branch snapshot
(first configured LLM-backed method: remote, handoff, or soft) using
a side session id isolated from the live turn. Crossing the threshold
splices the armed result in instantly instead of blocking on a
summarization round-trip. Armed results are invalidated by branch
changes, reset boundaries, model switches that strand provider-native
replay payloads, and context growth past keepRecentTokens (which
re-speculates); extensions registering session_before_compact keep
exact blocking semantics (speculation disabled).
- Reworked handoff to commit in place: /handoff and the auto handoff
method now write the generated document as a regular compaction entry
on the current session (summary = document + <files> tag, cut from
prepareCompaction) instead of starting a new session. SessionHandoff
shrank to a document generator; session_before_switch/session_switch
no longer fire with reason "handoff"; mid-turn maintenance no longer
suppresses the handoff preference; overflow recovery can apply an
armed handoff result.
- Extracted the shared auto-compaction commit tail
(#commitAutoCompactionResult / #commitCompactionEntry) used by the
blocking production path, the armed speculative apply, manual
compaction, and manual handoff.
- Status line pulses the auto-compact icon while a speculation runs and
holds it in accent once a result is armed.
- Exported remotePreserveReusable from pi-agent-core/compaction for
apply-time validation of speculative remote results.
- Implement a live status board utility for transient multi-line CLI status displays with TTY fallback support.
- Add progress callback support and forward subagent progress events to runner hooks.
- Update cleanse execution flow to track checker runs, agent progress, and status rendering.
- Add comprehensive unit tests for live board repainting and cleanse progress assertions.
- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.
buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.
Fixes#8789
- Kept the shared cursor/interaction-query module as the single handler and deleted the duplicate local implementation in cursor.ts
- Added the named webFetchRequestQuery approval case (field 9 is named under the regenerated proto)
- Preserved the deliberate no-fake-VM-success semantics for setupVmEnvironmentArgs (review of #8047)
- Updated the field-9 regression test to assert the named decode of the raw same-field reply, which also pins the LEN-prefix wire framing
The absolute 16,384-token floor plus the carried summary and output
reserves exceeds windows below ~58k outright, and overflow recovery then
bailed at the very floor that caused the rejection, leaving compaction
unusable on small-context models. Scale the floor to window/8 (min 1k)
and use the same floor in overflow recovery.
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.
Three independent defects:
1. `generateSummary` serialized the entire span into one prompt with no budget
check. It now plans windows that fit the summarizer's context and folds them
with the update prompt that iterative compaction already uses, so a stranded
boundary is recovered instead of rejected. A provider that rejects a window
the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
beta-gated to 200k on OAuth credentials) halves what was actually sent and
re-plans, because only the rejection knows the real cap.
2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
appends to its own errors classified a deterministic 400 as a transient 503.
Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.
3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
an input that does not fit never fits; both layers now fail fast to the next
candidate.
The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.
Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
Retry: widened agent dequeue-hook deadline budgets from 25ms to 1s — the run loop checks the deadline before invoking dequeue hooks, so a cold or CPU-starved mock roundtrip expired the deadline first and the hooks never ran (deterministic failure in isolation, flaky under CI parallel load).
prepareCompaction walked the branch from the last compaction and ignored reset_boundary markers, so /compact (and auto-compaction) resurrected pre-/clear turns into the summary even though buildSessionContext already starts the model context after the boundary.
Model reset_boundary as a first-class agent-core session entry and start the summarization window after the latest boundary, dropping the superseded pre-reset compaction summary. A boundary before the last compaction stays superseded by it.
Fixes#8718
- Retained tool-search server calls and opaque results in signed assistant history across direct streams, gateways, and custom-endpoint projection.
- Added replay regressions for interleaved thinking and client tool continuations.
Fixes#8559
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
Each is a single, side-effect-free completion whose result is parsed after
it resolves, so one transient provider failure previously aborted the
whole operation - for /compact that left the user's context full.
Adds SummaryOptions.oneshotRetry, because both compaction paths call the
same generateSummary and the policy therefore cannot be a constant inside
it. Manual /compact has no outer loop and gets retry by default;
auto-compaction passes false because session-maintenance already retries
the whole attempt, and nesting would multiply the budget (10 outer x 3
inner) while stacking each outer wait on an inner backoff.
instrumentedCompleteSimple is the single funnel for every oneshot LLM
call in the agent, so retry belongs here rather than in a try/catch above
each caller - the failure arrives as a resolved AssistantMessage.
Opt-in rather than default-on: oneshotKind is free-form and callers may
pass arbitrary ctx.tools, so the funnel cannot itself prove a request is
replay-safe.
Response headers are captured per attempt and cleared between attempts,
so a stale retry-after can never be reused for a later failure.
- Refactored and condensed numerous system prompts, agent instructions, and tool documentation files across packages.
- Streamlined workflow rules, formatting constraints, and execution guidelines for improved clarity and brevity.
- Updated discovery rules, recommendation criteria, and syntax standards in prompt templates.
RESCUE_SHAKE_CONFIG spreads AGGRESSIVE_SHAKE_CONFIG, so the new 4k manual
tail leaked into dead-end recovery and could block eliding the very
result that caused the dead end. Override protectTokens back to 0 in the
rescue preset and add the coding-agent changelog entry for the manual
/shake behavior change.
Addresses review on #8067.