- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
Retry: widened agent dequeue-hook deadline budgets from 25ms to 1s — the run loop checks the deadline before invoking dequeue hooks, so a cold or CPU-starved mock roundtrip expired the deadline first and the hooks never ran (deterministic failure in isolation, flaky under CI parallel load).
Retry: deflaked tui IME preedit border test (#5563) — VirtualTerminal.waitForRender's fixed 40ms sleep raced TUI's throttled render timer on starved CI runners, reading the pre-input frame; waitForRender now takes an optional settle predicate polled up to 2s and the test keys on the rendered content.
The startup graphics probe only ran on ConPTY hosts with WT_SESSION, so a
SIXEL-capable terminal that exposes no identifying environment variable
(foot exports TERM=foot and COLORTERM=truecolor only) resolved the
trueColor capability row, kept imageProtocol null, and rendered every
image as the "[Image: …]" text card.
The XTSMGRAPHICS branch also had its status inverted: per xterm ctlseqs a
reply of `CSI ? 2 ; Ps ; Pv S` carries Ps = 0 on success, and a terminal
without SIXEL reports a zero maximum geometry, so a successful reply was
read as unsupported.
Drop the dead DA1 half of the probe with it: ProcessTerminal swallows
every `CSI ? … c` reply for the whole session so a late one cannot leak
into the composer (#8542), which means the attribute list never reached
the probe's input listener on any platform. The bare `CSI c` it wrote was
also unaccounted for in the DA1 sentinel FIFO, so its reply consumed
another probe's sentinel.
PI_FORCE_IMAGE_PROTOCOL, including its off/none kill switch, still wins
over the probe.
Addresses review feedback on the initial commit:
- restore the trailing newline at EOF stripped by the first edit
- add a Fixed entry for this change under ## [Unreleased] in the coding-agent CHANGELOG
StopOnTextCriteria decoded the last STOP_DECODE_WINDOW_TOKENS of the whole
sequence, so prompt tokens were eligible for matching. A prompt that itself
contains the stop string stops generation at the first generated token and
yields an empty title.
Anchor the window to the generation boundary by recording the first
generated index per batch entry. Existing local title models are
unaffected: with the assistant-prefill prompt shape, the example `</title>`
tags sit outside the 32-token window for normal messages, so no shipping
model changes behavior. The bug becomes reachable with any chat-level
few-shot prompt that places the stop string near the generation boundary.
A Codex request to a model the signed-in ChatGPT account is not entitled
to fails with "The '<model>' model is not supported when using Codex with
a ChatGPT account." That was classified as a plain provider error, so the
request failed outright even when a sibling account was signed in and
entitled to the model.
Classify that exact denial as an account-policy error, the same category
`cyber_policy` already uses, so the existing credential-rotation path can
reach an entitled account.
The match is deliberately narrow: it fires only for provider
`openai-codex`, only when the denied model in the message is the model that
was requested, and only for a bounded, non-null model identity. A denial
naming some other model does not trigger rotation, so an unrelated mention
cannot burn sibling credentials.
The memory-extraction prompt concatenated its instructions, few-shot
examples, and the user message into a single user turn, so a small local
model could not distinguish instructions from input and frequently echoed
the Globex/weather examples instead of extracting facts.
Send the instructions as a real system turn and the raw text as the user
turn. The tiny worker protocol gains a systemPrompt field, and Mnemopi
completion input carries task metadata so the backend selects the right
prompt per call.
Drop the code-built MEMORY_EXTRACTION_TEMPLATE rather than porting it:
prompt text belongs in .md files, and resolveMemoryCompletionInput already
overrides that template for every extraction call, so Mnemopi rendered it
only for the result to be discarded.
Measured on ONNX q4 CPU, LFM2.5-1.2B memory extraction improved from 1/8
to 5/8 once the roles were separated.
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.