feat(coding-agent): added async compaction and in-place handoff

- Added compaction.asyncEnabled (Async Compaction, default on): when
  context enters the pre-threshold band [threshold - lead, threshold)
  with lead = clamp(threshold * 0.125, 8192, 32000), maintenance
  speculatively summarizes in the background off a branch snapshot
  (first configured LLM-backed method: remote, handoff, or soft) using
  a side session id isolated from the live turn. Crossing the threshold
  splices the armed result in instantly instead of blocking on a
  summarization round-trip. Armed results are invalidated by branch
  changes, reset boundaries, model switches that strand provider-native
  replay payloads, and context growth past keepRecentTokens (which
  re-speculates); extensions registering session_before_compact keep
  exact blocking semantics (speculation disabled).
- Reworked handoff to commit in place: /handoff and the auto handoff
  method now write the generated document as a regular compaction entry
  on the current session (summary = document + <files> tag, cut from
  prepareCompaction) instead of starting a new session. SessionHandoff
  shrank to a document generator; session_before_switch/session_switch
  no longer fire with reason "handoff"; mid-turn maintenance no longer
  suppresses the handoff preference; overflow recovery can apply an
  armed handoff result.
- Extracted the shared auto-compaction commit tail
  (#commitAutoCompactionResult / #commitCompactionEntry) used by the
  blocking production path, the armed speculative apply, manual
  compaction, and manual handoff.
- Status line pulses the auto-compact icon while a speculation runs and
  holds it in accent once a result is armed.
- Exported remotePreserveReusable from pi-agent-core/compaction for
  apply-time validation of speculative remote results.
This commit is contained in:
can1357
2026-08-20 03:45:44 +02:00
parent eced7ab08a
commit 2f54e760c1
26 changed files with 1218 additions and 1043 deletions
+1
View File
@@ -11,6 +11,7 @@
- Added `Tokenizer.checkTokenBudget(text, budget)`: a cheap-first budget probe. Byte length is a hard upper bound on token count, so text whose raw bytes already fit answers "fits" without tokenizing at all; only text that busts the bound pays for an exact count (and that count is returned, so a proportional clamp gets the denominator it needs). Since the bound overshoots ~4x on prose, the common "comfortably under budget" answer is free. Compaction's summary-window fit check and OpenAI remote-compaction trimming now route through it.
- Added provider-anchored transcript accounting (`findTranscriptUsageAnchor`, `isTranscriptUsageAnchor`, `estimateTranscriptTokens`). Every settled assistant turn carries `usage` covering the exact prompt it was sent, so transcript sizing charges that report for the prefix and tokenizes only the tail appended after it — counting proportional to one turn instead of the whole history, every turn. The four hand-rolled copies of the anchor trust rules (session stats ×3, shake) now share one predicate, and the deliberately provider-independent compaction floor counts every message locally via `tokenizer.countMessages`.
- Exported `remotePreserveReusable(preserveData, activeModel, settings)` — whether a prior remote compaction's provider-native replay payload is still readable by the active model — so hosts can validate speculatively produced compaction results before committing them.
### Changed
+1 -1
View File
@@ -1268,7 +1268,7 @@ export interface CompactionPreparation {
* let the active model replay it, so keying reuse on "any candidate shares the
* provider" left a provider-switched session permanently context-less (#6343).
*/
function remotePreserveReusable(
export function remotePreserveReusable(
preserveData: Record<string, unknown> | undefined,
activeModel: Model,
settings: CompactionSettings,