estimateTokens now charges for serialized anthropicServerTool blocks so context maintenance sees the server-tool payload replayed on the wire; excluded from the compaction floor like other encrypted reasoning.
After an OpenAI remote compaction, prepareCompaction decided whether to
keep the provider-native replay boundary or re-expand its originals by
asking whether *any* compaction candidate (every role model plus the
largest-context available model) shared the payload's provider. In a
multi-role setup where a role such as modelRoles.smol stays on OpenAI,
the check passed forever, so a session switched to a non-OpenAI active
model kept a placeholder-only summary and never recovered the compacted
span for the rest of the session.
Judge reusability against the active model — the one that assembles the
request context every turn — instead of the candidate set. When the
active model cannot replay the payload, re-expand the originals into a
portable local summary, matching the self-healing already present for
single-provider migrations.
Fixes#6343
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.
- Added a per-message estimate cache in agent-core keyed by identity, with a
settle gate (assistants cache only with real usage + terminal non-error
stopReason; streaming partials bypass) and dual option-split WeakMaps for the
default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
with an exact-repeat outer-array reuse and slice-on-growth for append-only
turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
the prewalk plan-nudge scrub, via invalidateMessageCache /
registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
convert and repeat estimate are all >10x faster with noise under 20%.
Fixes#5934
Escaped Harmony control-token markers when compaction serializes transcripts into plain summary prompts so Copilot gpt-5.6 models do not reject analysis-channel text.
Added regression coverage for summary-bound Harmony serialization while keeping native transcript rendering unchanged.
Fixes#5184
- Introduced `stringifyJson` helper to preserve bigint precision by serializing them as decimal strings.
- Replaced native `JSON.stringify` across compaction and session management modules to prevent serialization errors when handling bigint values in tool arguments.
- Added regression tests in `agent` and `coding-agent` packages to ensure bigint tool arguments remain intact through compaction and persistence flows.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
Retried handoff generation with toolChoice auto when a provider rejects the cache-preserving toolChoice none request as auto-only.
Kept unrelated provider 400s terminal so bad request failures still surface without masking the cause.
Fixes#4715
- Sent OpenAI-compatible chat messages when compaction.remoteEndpoint targets /chat/completions while preserving the existing custom summarizer payload elsewhere.
- Added regressions for direct wire formatting and end-to-end openai-completions compaction against a configured chat endpoint.
Fixes#4630
calculateContextTokens returned usage.totalTokens which, with the new
Usage.orchestration sidecar, folds provider-side orchestration back into the
context size used by auto-compaction/context promotion thresholds. Subtract
the orchestration sidecar so context sizing stays conversation-only while
cost and totalTokens keep the orchestration spend visible.
Refs #4469
- Made CompactionSettings.reserveTokens optional so field presence carries provenance; the proportional small-window fallback only applies to genuinely defaulted reserves.
- Clamped the fallback reserve to >= 1 and the derived threshold strictly below the context window.
- Changed the coding-agent settings-schema default from 16384 to unset so Settings.get() no longer materializes a default that masks provenance.
Cherry-pick of the reserve-budget clamp only (resolveBudgetReserveTokens + no-op compaction guard): applies compaction.ts + agent-session.ts + compaction/shake/progress-guard tests. Excludes unrelated Julia prelude timeout and ai/test churn from the PR head.
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
- Implement provider-native replay logic to enable reuse of remote compaction data across compatible models.
- Enhance compaction logic to re-expand and locally summarize remote history when provider-native replay is unavailable.
- Update OpenAI request setup to include session and routing identifiers for improved traceability.
- Refine token estimation for image content during truncation to ensure more accurate budget management.
- Update agent-session to resolve compaction model candidates before persistence, ensuring authentication availability.
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
- Fixed stale `preserveData.snapcompact` frames leaking into context-full compaction after switching from `snapcompact` to `context-full` strategy, which inflated context usage and made sessions appear to compact prematurely.
- Added secret redaction for migrated snapcompact archive plaintext (`text`/`textHead`/`textTail`) during the snapcompact->context-full transition, while preserving opaque provider-replay state byte-identical.
- Added `archiveSourceText()` and `stripPreservedArchive()` utilities to snapcompact module for archive extraction and cleanup.
- Consolidated duplicate `stripSnapcompactPreserveData` functions into `snapcompact.stripPreservedArchive`.
- Added unit tests to verify archive removal and empty state collapse behavior.
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
CI caught that flooring by the raw local estimate falsely triggers compaction on
thinking-heavy turns: estimateTokens counts the opaque thinkingSignature /
redactedThinking payloads (providers bill them on replay, #2275), but their local
byte size diverges wildly from what the provider actually charges — so a turn
with a large encrypted-reasoning blob but small provider usage would trip the
floor (broke agent-session-handoff 'provider-anchored usage' test).
estimateTokens now takes { excludeEncryptedReasoning } and the compaction floor
(#estimateStoredContextTokens) uses it: the floor counts only reliably-countable,
on-wire-compressible content (text, tool results, tool calls), while the provider
usage arm of compactionContextTokens still accounts for encrypted reasoning. This
keeps the encrypted-reasoning case provider-anchored while still flooring upward
when a before_provider_request hook compresses tool results.
A before_provider_request extension (a context-compression proxy like Headroom,
an obfuscator, or inline snapcompact) can shrink the outgoing request below the
real stored conversation. The provider then reports deflated prompt tokens, so
the auto-compaction threshold never fires and the stored history grows unbounded
until it overflows the context window and can no longer be compacted at all.
Add compactionContextTokens(provider, storedEstimate) = max(provider, estimate)
and apply it to both the pre-prompt and post-response compaction decisions,
flooring the provider-reported tokens by the agent's own estimate of the stored
conversation. Display and cost accounting still use exact provider usage; only
the compaction trigger takes the floor.
- Extracted native token counting into a new localized `tokenizer.ts` wrapping `@oh-my-pi/pi-natives`.
- Introduced a lightning-fast byte-length estimation logic for token counting when accurate counting is disabled.
- Diverted token calculations to the faster estimator during test environments and when `PI_TOKENIZER_ACCURATE` is falsy.
- Updated agent base and coding-agent sessions to consume the new localized `countTokens` utility.
- Renamed ToolCallSyntax type to Dialect and Grammar interface to DialectDefinition across all packages.
- Moved grammar directory to dialect and updated all import paths in agent, ai, catalog, and coding-agent packages.
- Added renderTranscript and renderThinking methods to DialectDefinition, enabling native dialect-aware conversation serialization.
- Consolidated rendering utilities into new dialect/rendering.ts with shared helpers for ChatML, legacy text, and dialect-specific formatting.
- Updated conversation serialization in agent and coding-agent to use dialect.renderTranscript() for native turn envelope rendering.
- Updated compaction, branch summarization, and session dump formatting to pass preferred model tool syntax into conversation serialization.
- Enhanced shared serializers to render assistant tool calls and tool results through grammar envelopes when syntax is available, with the prior compact format as fallback.
- Aligned prompt, preview, and test fixtures to the new transcript tags: `[Think]`, `[Tool Call]`, and `[Tool Result]`.
- InputController now dispatched Esc to active viewSession operations, aborting compaction, handoff, and retry directly.
- Removed competing onEscape handler swaps across command and event controllers so overlapping auto/manual flow events no longer overwrote cancellation callbacks.
- Compaction now propagated fetch options and rethrew aborted signals so cancellations were not treated as remote failures.
- Renamed all functions, types, and constants in @oh-my-pi/snapcompact to namespace-relative names (`snapcompactCompact` → `compact`, `renderSnapcompactFrames` → `renderMany`, `snapcompactFrameCount` → `frames`, `SnapcompactShape` → `Shape`, `SNAPCOMPACT_SHAPES` → `SHAPES`, …).
- Converted every consumer to `import * as snapcompact` member access: `agent/compaction.ts`, `coding-agent` `agent-session.ts`/`session-manager.ts`/`snapcompact-inline.ts`, and all affected tests.
- Renamed internal `geometry` locals to `geo` in `snapcompact.ts` to avoid TDZ collisions with the new `geometry` export.
- Updated `docs/compaction.md` prose and added a Breaking Changes entry to the snapcompact changelog documenting the full rename map.
The shake-strategy post-shake threshold check was reading
#estimatePendingPromptTokens([]) while #checkCompaction triggered on
calculateContextTokens(assistantMessage.usage). The local estimator
ignored block.thinkingSignature payloads (OpenAI Responses encrypted
reasoning items, Anthropic signed thinking blocks, etc.), so on a
thinking-heavy session the estimate sat ~0.9–2× below provider-reported
usage. Once the two straddled the threshold, the #2119 dead-loop guard
never fired, shake reported 'handled', and #scheduleAutoContinuePrompt
re-injected the auto-continue developer prompt every turn — 53 injections
in a real 25-minute repro session before an external timeout.
Thread the trigger's provider-anchored contextTokens through
#runAutoCompaction → #runAutoShake for the threshold and incomplete
paths, then evaluate residual pressure as triggerContextTokens −
result.tokensFreed with an 80% recovery-band hysteresis. Re-checking
against the raw threshold (even on the corrected metric) would still let
shake reclaim a trickle of the previous turn's elidable blocks and land
just under the line every turn; the band closes that oscillation.
As defense in depth, estimateTokens() now charges thinkingSignature and
redactedThinking.data alongside the visible thinking text so every
other site that uses the estimator (idle compaction, pre-prompt check,
status line) tracks provider usage on replay.
New regression test pins the contract; existing dispatch test bumped
its mocked tokensFreed so its happy-path scenario lands inside the new
recovery band.
Fixes#2275
- Added a new @oh-my-pi/snapcompact package and redirected compaction call sites to it.
- Added provider-aware snapcompact shape resolution for model-specific mixed-frame behavior.
- Added optional image detail support by extending ImageContent and passing hints through OpenAI providers.
- Added native snapcompact render options, including 5x8/8x8 font loading and palette/geometry controls.
Adds snapcompactCompact() in compaction/snapcompact.ts: instead of an LLM-generated summary, discarded history is printed onto dense 2576px PNG frames with the public-domain X.org 5x8 pixel font and re-attached to the compaction summary message as image blocks. Fully local — no model call; ~7x cheaper than raw text at near-parity recall. CompactionSummaryMessage now charges per attached frame in estimateTokens(), frames persist under preserveData.snapcompact with an 8-frame budget that evicts middle-out (session-head frame pinned so head and tail both survive). Rasterization and PNG encoding run in native code via renderSnapcompactPng().
Adds pruneSupersededToolResults() and the opt-in PruneConfig.supersedeKey hook: when a tool call shares a key with a newer one (e.g. a re-read of the same file), the older result is pruned even inside the protectTokens window and replaced with a [Superseded by a newer read of this file] placeholder. Adds readToolSupersedeKey() and the shared splitReadSelector() implementing the read-tool path/selector grammar (including the .. range alias and L-prefix forms) so selector-free reads supersede range reads of the same file and URL-scheme paths are exempt. Strips selectors before tracking in <read-files> compaction lists, so reads dedupe to the base path and match write/edit paths when splitting read-only vs modified lists (selector-polluted lists from earlier compactions self-heal on the next pass). Gated by the new compaction.supersedeReads setting (default on).
Splits the array-form defaultConvertToLlm into a single-message convertMessageToLlm that embedders can delegate every core role to instead of duplicating the conversion. Adds the optional images field on CompactionSummaryMessage so the converter attaches snapcompact frames after the summary text (snapcompact strategy lands next). Renames every convertToLlm call site in compaction.ts and branch-summarization.ts to the canonical defaultConvertToLlm.\n\nNote: the bundled CHANGELOG entries also cover the supersede-reads, snapcompact, and steering-queue fixes that follow in this batch (the entries land in directly adjacent lines and cannot be split by diff).
Converted provider-facing tool parameters to wire JSON Schema before redaction so live Zod instances are not deep-cloned into plain objects.
Fixes#2146
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.
Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.
Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.
BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
Completes the injectable-fetch transport wiring (15.10.8) that the feature
left half-done, fixing the deterministic CI test failures:
- compaction.compact() rebuilt summaryOptions field-by-field but dropped
`fetch`, so the injected transport never reached
requestOpenAiRemoteCompaction / generateSummary's remote path. Thread it.
- Read-tool URL pipeline had no fetch seam: renderHtmlToText gained a
fetchOverride param but renderUrl/ToolSession never carried one, so the
jina/parallel reader backends always used global fetch. Add
ToolSession.fetch -> renderUrl -> renderHtmlToText (defaults to global).
- searchWithParallel mirrored extractWithParallel but missed the fetch
option; add it.
- Repair tests whose deleted hookFetch interceptors were never replaced
with a FetchImpl seam (fetch-kagi-toggle, web-search-parallel,
issue-970 discovery).
- Update issue-1746 POSIX case to the #2154 preserved-scrollback contract:
unknown-viewport streaming deferral is now platform-independent.
- Added optional FetchImpl fields to compaction, proxy, AI, coding-agent, and mnemopi options.
- Threaded injected fetch implementations through OAuth, discovery, and search/LLM request flows.
- Removed exported hookFetch utility and its package entrypoint from utils.
- Replaced global-fetch test monkeypatching with per-test FetchImpl mocks across test suites.
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
- Centralized compaction stop-reason error throws via createSummarizationError().
- Set compaction thrown errors to copy response.errorStatus into Error.status.
- Expanded compaction auth detection to treat HTTP 401/403 as auth failures with regex fallback preserved.
- Added regression tests for 401/403 status propagation and compaction fallback auth behavior.
- Documented both package fixes in Unreleased Fixed changelog entries.