- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
- Implement provider-native replay logic to enable reuse of remote compaction data across compatible models.
- Enhance compaction logic to re-expand and locally summarize remote history when provider-native replay is unavailable.
- Update OpenAI request setup to include session and routing identifiers for improved traceability.
- Refine token estimation for image content during truncation to ensure more accurate budget management.
- Update agent-session to resolve compaction model candidates before persistence, ensuring authentication availability.
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
- Added a fixed-width title slot system to serialize and persist session titles across physical files and backend storage.
- Integrated automated session title updates triggered by todo replan operations using conversation history context.
- Extended storage interfaces across memory, file-system, Redis, and SQL backends to support independent title metadata updates.
- Implemented title metadata parsing within session loaders and list utilities to ensure accurate retrieval and display.
PR #3648 review (codex): the first iteration emitted every envelope
path alongside one combined digest, so a multi-file payload that added
`: any` to a README.md hunk and merely touched src/ok.ts would surface
a *.ts path in the TTSR match context and trip the bundled
tool:edit(*.ts) ts-no-any rule on text that belonged to the Markdown
hunk — aborting valid edits under interruptMode:always.
Add a per-file matcherEntries(args) hook on AgentTool / EditStreaming-
Strategy returning [{ path, digest }] entries, one per touched file
(same-path sections/hunks merged):
- replace / patch: one entry from the top-level path + matcherDigest
- hashline: regex-split by [path#TAG] section, body added-lines per
entry (tolerant of streaming partial payloads)
- apply_patch: expandApplyPatchToPreviewEntries grouped by path
AgentSession.#checkTtsrStream / #checkTtsrAstStream now prefer
matcherEntries and iterate per-file with isolated filePaths + streamKey,
so each file's buffer and repeat-tracking are independent. Tools
without matcherEntries keep the existing combined matcherDigest +
matcherPaths path.
AgentSession's TTSR match context only scanned top-level path/paths
arguments, so hashline and apply_patch edit streams (whose only target
path lives inside the wire payload — section headers or envelope
markers — not as a top-level argument) arrived without any filePaths
and silently skipped path-scoped rules like the bundled ts-no-any
(scope: tool:edit(*.ts)).
Add an optional AgentTool.matcherPaths(args) hook, companion to the
existing matcherDigest(args), so tools whose wire grammar embeds paths
can surface them. Implement on each edit streaming strategy:
- replace / patch: top-level path
- hashline: parse [path#TAG] (and tag-less [path]) section headers
tolerant of streaming partial payloads
- apply_patch: parse *** Add/Update/Delete File: markers, also tolerant
of pre-End-Patch buffers
AgentSession.#getTtsrToolMatchContext consults tool.matcherPaths first,
normalising its output through the existing path-candidate helper, and
falls back to the generic top-level argument scan for tools that don't
implement it.
Fixes#3646
- Update scrubPartialJson to utilize clearStreamingPartialJson for consistent tool-call cleanup.
- Adjust execution order in streamProxy to ensure partial error messages are finalized before scrubbing.
- Remove redundant test expectation comment regarding partialJson leakage.
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
`delete` on object properties degrades V8 hidden class optimization; the new `stripVariant` util sets the property to `undefined` instead, keeping the object shape stable.
`performance.now()` is used in place of `Date.now()` for duration and TTFT measurements to get a monotonic, high-resolution clock that is unaffected by system clock adjustments.
- Removed the pi dialect implementation and associated source files.
- Updated dialect resolution, factory registration, and type definitions to exclude pi.
- Cleaned up settings schema and user options to remove pi-related configurations.
- Deleted corresponding test suites covering pi dialect functionality, in-band tools, and examples.
- Centralized JSON parsing and stream processing logic by moving utilities from `packages/ai` to the shared `@oh-my-pi/pi-utils` package.
- Standardized import paths for JSON parsing and streaming across the agent, ai, and coding-agent packages.
- Refactored SSE stream handling to use consolidated `parseStreamingJson` logic and introduced robust error recovery for malformed container-shaped tail events.
- Cleaned up legacy bundled registry references and updated related module exports and tests to reflect the new utility structure.
Scoped provider-refusal filtering to live replay so compaction and snapcompact summaries retain the refused turn while outbound provider context still drops the refusal.
Fixes#3592
API-level refusals now stay visible as terminal errors without being sent back as assistant dialogue on the next provider request. Added core and coding-agent conversion coverage for Anthropic refusal metadata.
Fixes#3592
- Fixed stale `preserveData.snapcompact` frames leaking into context-full compaction after switching from `snapcompact` to `context-full` strategy, which inflated context usage and made sessions appear to compact prematurely.
- Added secret redaction for migrated snapcompact archive plaintext (`text`/`textHead`/`textTail`) during the snapcompact->context-full transition, while preserving opaque provider-replay state byte-identical.
- Added `archiveSourceText()` and `stripPreservedArchive()` utilities to snapcompact module for archive extraction and cleanup.
- Consolidated duplicate `stripSnapcompactPreserveData` functions into `snapcompact.stripPreservedArchive`.
- Added unit tests to verify archive removal and empty state collapse behavior.
Responses-style providers serialize providerPayload history items instead of the
visible message blocks when replaying native history. Include providerPayload in
the append-only per-message digest so payload-only history rewrites stop the
stable-prefix walk and re-sync the changed message before any later divergent
tail.
Add a regression where an assistant message keeps identical visible content and
id but changes its openaiResponsesHistory providerPayload while a later message
also diverges; syncMessages must preserve the prefix before the assistant and
refresh the assistant payload.
Fixes#3406
Direct callers can clear AppendOnlyContextManager.log without resetting the
private sync cursor. The advisor reset path does this when recycling its helper
agent, leaving lastSyncCount and messageDigests describing the old transcript
while the physical log is empty.
Clamp the stable-prefix reuse count to the current log length before truncating
and appending. A direct log clear now forces the next sync to replay from index 0
instead of starting from a stale private cursor and dropping prefix messages from
the provider context.
Add a regression that clears the public log after syncing two messages, then
resyncs a context with the same first message and a rewritten second; both
messages must be present in the rebuilt append-only log.
Fixes#3406
Track internal tool-result metadata in append-only per-message digests so
metadata-only rewrites of toolCallId, toolName, or isError stop the stable-prefix
walk and re-sync the changed tool result before any later divergent tail.
This prevents stale tool-result pairing or error state from being preserved when
the text content stays unchanged but provider-serialized metadata changes.
Fixes#3406
`AppendOnlyContextManager.syncMessages` hashed a single rolling digest
over the entire synced prefix, so any in-place rewrite of an already-
synced message — per-turn `pruneSupersededToolResults` / `pruneToolOutputs`
collapsing a tool result, image stripping, or a `transformContext` re-render
— triggered `log.clear()` and re-appended the full conversation from
the current (mutated) view. The provider's cached bytes still matched
the prefix, but every position past the divergence had to be re-prefilled.
On llama.cpp / Ollama / LM Studio this re-prefilled tens of thousands
of tokens every few turns (`n_past \u2248 end-of-system-prompt` collapse,
~40k-token full re-prefill, GPU pinned >400W).
Replace the rolling digest with per-message digests in `#messageDigests`,
walk the new sync against them to find the longest byte-stable prefix,
truncate the log down to that prefix via a new `AppendOnlyLog.truncate(count)`,
and only re-append the diverged tail. Genuine compaction (`length <
lastSyncCount`) still clears the log.
- Tail-only rewrite: prefix stays byte-stable; only the trailing message
re-syncs.
- Deep rewrite: prefix up to the divergence stays byte-stable; the
provider re-prefills from the divergent message onward (architectural
minimum).
- True compaction: unchanged, full replay.
Replace the now-misleading `detects in-place rewrite of already-synced
messages` / `detects in-place rewrite via digest mismatch` tests with
`preserves the byte-stable prefix when a deep message is rewritten (#3406)`,
`preserves the prefix when the tail is rewritten (#3406)`, `appended
new messages keep the prefix stable even when the prior tail also
diverged (#3406)`, and `rewriting the first message still re-syncs
from scratch` so each invariant is asserted directly.
Fixes#3406
- Extracted yield logic into a configurable YieldGate class to avoid process-global state.
- Injected time and sleep dependencies to support deterministic testing.
- Handled potential negative time progression by forcing a re-anchor instead of gating indefinitely.
- Maintained existing behavior for the public yieldIfDue export via a shared instance.
Address review feedback: when the SSE stream disconnects after a
toolcall_delta but before toolcall_end/done/error, the catch-block
at lines 183-192 pushes the partial message as the error result
without calling scrubPartialJson. This leaked the internal partialJson
field into the final error message.
Added scrubPartialJson(partial) call in the catch block, before
pushing the error event.
Added test verifying partialJson does not leak when server disconnects
mid-tool-call (toolcall_start + partial toolcall_delta, no terminal event).
Address review feedback: downstream renderers (event-controller.ts:535)
read content.partialJson during toolcall_delta to pace streaming
previews (bash env assignments, write/edit smooth streaming).
Revised approach:
- toolcall_start: initialize partialJson on content via typed
ToolCall & { partialJson: string } intersection (not as any)
- toolcall_delta: accumulate in side-channel Map, write onto content
via typed intersection cast
- toolcall_end: delete partialJson from content + side-channel map
- done/error: scrubPartialJson() cleans any remaining blocks that
never got toolcall_end (the original leak bug, now fixed for all
terminal paths)
Added test verifying partialJson IS present during streaming and
IS absent after completion.
streamProxy stored internal partialJson streaming state directly on typed
ToolCall objects via 4 'as any' casts. If toolcall_end was skipped (stream
error, early done), the field leaked into the final AssistantMessage content,
corrupting downstream serialization.
Replace with a side-channel Map<number, string> keyed by contentIndex:
- toolcall_start initializes the map entry
- toolcall_delta accumulates into it
- toolcall_end cleans it up
The typed ToolCall object never carries non-spec fields. All 4 'as any'
casts are eliminated.
Added 4 contract tests covering argument parsing, partialJson isolation on
normal completion, partialJson isolation when toolcall_end is missing, and
multiple concurrent tool calls with interleaved deltas.
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
- Updated `render`, `renderMany`, and native snapcompact methods to return promises, ensuring scalable async execution.
- Refactored `transformProviderContext` and `buildSideRequestContext` to support asynchronous operations in agent loops.
- Integrated `Promise.all` for improved concurrency when processing frame rendering and rendering batch operations.
- Updated all internal call sites, SDK hooks, and test suites to accommodate the asynchronous API signatures.
CI caught that flooring by the raw local estimate falsely triggers compaction on
thinking-heavy turns: estimateTokens counts the opaque thinkingSignature /
redactedThinking payloads (providers bill them on replay, #2275), but their local
byte size diverges wildly from what the provider actually charges — so a turn
with a large encrypted-reasoning blob but small provider usage would trip the
floor (broke agent-session-handoff 'provider-anchored usage' test).
estimateTokens now takes { excludeEncryptedReasoning } and the compaction floor
(#estimateStoredContextTokens) uses it: the floor counts only reliably-countable,
on-wire-compressible content (text, tool results, tool calls), while the provider
usage arm of compactionContextTokens still accounts for encrypted reasoning. This
keeps the encrypted-reasoning case provider-anchored while still flooring upward
when a before_provider_request hook compresses tool results.
A before_provider_request extension (a context-compression proxy like Headroom,
an obfuscator, or inline snapcompact) can shrink the outgoing request below the
real stored conversation. The provider then reports deflated prompt tokens, so
the auto-compaction threshold never fires and the stored history grows unbounded
until it overflows the context window and can no longer be compacted at all.
Add compactionContextTokens(provider, storedEstimate) = max(provider, estimate)
and apply it to both the pre-prompt and post-response compaction decisions,
flooring the provider-reported tokens by the agent's own estimate of the stored
conversation. Display and cost accounting still use exact provider usage; only
the compaction trigger takes the floor.
- Added `buildSideRequestContext` to the `Agent` class to generate prompt-cache-friendly provider contexts.
- Updated ephemeral side-channel turns to forward the full tool catalog to maintain prompt cache hit rates.
- Injected a `developer` role reminder into ephemeral turns to instruct the model to suppress tool calls.
- Implemented automatic post-processing to strip any tool calls from ephemeral turn responses.
- Exported message and dialect helper functions in `agent-loop.ts` to support context construction.
Azure Responses remote compaction was enabled by shouldUseOpenAiRemoteCompaction
but requestOpenAiRemoteCompaction still built OpenAI-style requests:
Authorization: Bearer plus a URL without api-version. Azure Responses uses an
api-key header and api-version query parameter, so normal Azure configs failed
and fell back to local summarization.
Derive Azure compact URLs from the Azure Responses base URL/resource-name
configuration, append api-version, and send api-key headers while preserving
custom headers. Existing OpenAI and Codex request shapes are unchanged.
Added a regression test that opts into azure-openai-responses compaction and
asserts the compact URL, api-key auth, absence of Authorization, custom header
preservation, and configured compaction model payload.
Fixes#3104
- Added provider/model remoteCompaction metadata and models.yml propagation.\n- Routed configured OpenAI-compatible compaction endpoints for custom providers.\n- Added compactionModel as a summary-only model selector that leaves the active session model unchanged.\n\nFixes #3104
- Added support for "Fast" serving-path variants for select Fireworks models.
- Updated compatibility logic to route `-fast` suffixes to the appropriate router wire format.
- Extended the model generation catalog to include these Fast variants with their respective pricing.
- Updated AI types to allow the `priority` service tier for Fireworks providers.