feat(coding-agent): added async compaction and in-place handoff
- Added compaction.asyncEnabled (Async Compaction, default on): when context enters the pre-threshold band [threshold - lead, threshold) with lead = clamp(threshold * 0.125, 8192, 32000), maintenance speculatively summarizes in the background off a branch snapshot (first configured LLM-backed method: remote, handoff, or soft) using a side session id isolated from the live turn. Crossing the threshold splices the armed result in instantly instead of blocking on a summarization round-trip. Armed results are invalidated by branch changes, reset boundaries, model switches that strand provider-native replay payloads, and context growth past keepRecentTokens (which re-speculates); extensions registering session_before_compact keep exact blocking semantics (speculation disabled). - Reworked handoff to commit in place: /handoff and the auto handoff method now write the generated document as a regular compaction entry on the current session (summary = document + <files> tag, cut from prepareCompaction) instead of starting a new session. SessionHandoff shrank to a document generator; session_before_switch/session_switch no longer fire with reason "handoff"; mid-turn maintenance no longer suppresses the handoff preference; overflow recovery can apply an armed handoff result. - Extracted the shared auto-compaction commit tail (#commitAutoCompactionResult / #commitCompactionEntry) used by the blocking production path, the armed speculative apply, manual compaction, and manual handoff. - Status line pulses the auto-compact icon while a speculation runs and holds it in accent once a result is armed. - Exported remotePreserveReusable from pi-agent-core/compaction for apply-time validation of speculative remote results.
This commit is contained in:
@@ -11,6 +11,7 @@
|
||||
|
||||
- Added `Tokenizer.checkTokenBudget(text, budget)`: a cheap-first budget probe. Byte length is a hard upper bound on token count, so text whose raw bytes already fit answers "fits" without tokenizing at all; only text that busts the bound pays for an exact count (and that count is returned, so a proportional clamp gets the denominator it needs). Since the bound overshoots ~4x on prose, the common "comfortably under budget" answer is free. Compaction's summary-window fit check and OpenAI remote-compaction trimming now route through it.
|
||||
- Added provider-anchored transcript accounting (`findTranscriptUsageAnchor`, `isTranscriptUsageAnchor`, `estimateTranscriptTokens`). Every settled assistant turn carries `usage` covering the exact prompt it was sent, so transcript sizing charges that report for the prefix and tokenizes only the tail appended after it — counting proportional to one turn instead of the whole history, every turn. The four hand-rolled copies of the anchor trust rules (session stats ×3, shake) now share one predicate, and the deliberately provider-independent compaction floor counts every message locally via `tokenizer.countMessages`.
|
||||
- Exported `remotePreserveReusable(preserveData, activeModel, settings)` — whether a prior remote compaction's provider-native replay payload is still readable by the active model — so hosts can validate speculatively produced compaction results before committing them.
|
||||
|
||||
### Changed
|
||||
|
||||
|
||||
@@ -1268,7 +1268,7 @@ export interface CompactionPreparation {
|
||||
* let the active model replay it, so keying reuse on "any candidate shares the
|
||||
* provider" left a provider-switched session permanently context-less (#6343).
|
||||
*/
|
||||
function remotePreserveReusable(
|
||||
export function remotePreserveReusable(
|
||||
preserveData: Record<string, unknown> | undefined,
|
||||
activeModel: Model,
|
||||
settings: CompactionSettings,
|
||||
|
||||
Reference in New Issue
Block a user