- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
The idle watchdog aborts the request signal and cursor.ts closes
the Connect stream, so there is no in-flight server exec to race.
Unmarked MCP/todo blocks can continue once every emitted call has
a matching result, same as HTTP/2 RST.
A ThinkingLoop abort is the loop guard asking for a same-model resample
(it injects a thinking-loop-redirect notice that only makes sense on the
model that looped), not a provider failure. #handleRetryableError routed
it through the generic retryable-error branch, so on attempt 1 it called
noteRetryFallbackCooldown + #tryRetryModelFallback and could switch to
another family from retry.fallbackChains while parking the original
selector on a 5-minute cooldown. A healthy Grok 4.6 planning turn got
replaced by whatever the chain listed next.
Carve ThinkingLoop out of the model-fallback branch and out of the
Fireworks Fast->base degrade so the loop guard always re-samples the same
model; the retry budget still bounds a genuinely stuck stream.
Fixes#8760
prepareCompaction walked the branch from the last compaction and ignored reset_boundary markers, so /compact (and auto-compaction) resurrected pre-/clear turns into the summary even though buildSessionContext already starts the model context after the boundary.
Model reset_boundary as a first-class agent-core session entry and start the summarization window after the latest boundary, dropping the superseded pre-reset compaction summary. A boundary before the last compaction stays superseded by it.
Fixes#8718
Mirrored registry status from pre-wire session run-state transitions and required live session corroboration before a peer can sustain bare hub waits.
Fixes#8634
The online difficulty and unexpected-stop classifiers call the tiny/smol model with disableReasoning plus maxTokens=1024. On the openai-completions transport (LiteLLM), disableReasoning on a reasoning model is downgraded to the lowest reasoning effort, so omp still emits reasoning_effort. LiteLLM/Vertex translates that to an Anthropic thinking.budget_tokens of at least 1024, and max_tokens=1024 is not greater than the budget, so every classifier call 400s.
Give the online classifiers 4096 output tokens so the request clears a proxy-injected minimum thinking budget (and leaves room for the keyword). Local reasoning budgets are unchanged.
Fixes#8610
Only the post-commit end/session_compact fan-out is detached.
The start emit still waits so input during that yield lands
in the compaction queue, matching the existing comment.
Rolled partial file appends back to their pre-write size and marked malformed resumed sessions for an atomic rewrite.
Retried transient persistence failures from in-memory state and surfaced the first failure in the interactive TUI.
Fixes#8596
Skipped same-provider cross-model fallback candidates when the latest assistant turn contains signed or redacted Anthropic thinking.
Kept same-model retry available for transient failures and added regression coverage.
Fixes#8558
Prevented deterministic 400 errors for immutable thinking blocks from entering same-model retries or configured model fallback.
Added a signed-thinking session regression covering the terminal error and retry UI lifecycle.
Fixes#8558
A user-invoked /skill:<name> reaches the session as a user-attributed skill custom message (role custom, attribution user) whose expanded SKILL.md body is the task prompt. The auto-thinking gate in #promptWithMessage only accepted role === "user", so these turns skipped classifyDifficulty/applyAutoThinkingLevel and the effort stayed stuck on pending auto.
Broaden the gate to also accept user-invoked skill prompts via the now-exported isUserInvokedSkillPrompt helper; agent-originated and autoload skill injections stay excluded.
Fixes#8554
Address review: #checkpointState is now assigned in the pre-await section alongside the reminder append, so a model that calls rewind immediately after seeing the notice finds an active checkpoint instead of 'No active checkpoint'. The entry id is backfilled post-await once the checkpoint toolResult entry is persisted; #applyRewind runs on a later rewind turn, never before the backfill.
Move the per-request date/cwd line out of the system prompt into a
first-turn system-reminder so open-weight providers keep their tool-schema
prefix cache; the reminder refreshes itself at midnight. Closes#7404.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <noreply@codebuff.com>