- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
Follow-up fix round for 71d608aae6 (validate PR review comment
anchors against the diff before submitting). The validation path
broke in several places the original commit did not cover:
- GitHubBackend.submit_review gained no commit_id parameter, so
submitting through GitHubProxyClient raised a production
TypeError on a path the validation code depended on.
- PullRequestInfo lacked head_sha, so the Forgejo commit_id
fallback raised AttributeError; it is now parsed in
_pr_from_payload and carried through the proxy round-trip.
- _pr_file_from dropped the file patch, silently no-oping anchor
validation for anything routed through the proxy; the patch is
now forwarded.
- The hunk parser treated any +++/--- line as a file header,
desyncing line counters when added/removed content began with
those prefixes; file headers are now recognized only before
the first hunk.
- Reworded the comment to "Forgejo only" to match the actual
backend behavior.
Adds tests for forgejo commit_id fetch, fallback double-failure,
empty-patch fail-open, LEFT-side anchoring, proxy commit_id
validation, and file-creation hunk boundaries.
Review follow-up: a safe `-wm` row previously kept backend-parsed capability metadata (no 1M floor, no daybreak pricing) while its synthesized plain listing was enriched — the same model reported two different contexts. Both listings now derive fallback window, the 1M floor, and daybreak cost from the canonical plain slug; unknown `-wm` SKUs and non-worker models keep their verbatim slug-derived metadata.
The authoritative-discovery test now resolves the configured `openai-codex/gpt-5.6-luna` through the real model resolver and asserts the exact bound id (and that an explicit `-wm` config still resolves verbatim) instead of only checking `ids.toContain`.
Codex backend discovery advertises worker-mode SKUs under a `-wm` suffix (gpt-5.6-luna-wm). Authoritative discovery replaced the bundled catalog and kept those slugs verbatim, so a configured `openai-codex/gpt-5.6-luna` vanished from the resolved catalog and the resolver's fuzzy fallback selected `-wm` instead — a route some ChatGPT accounts reject.
Model discovery now recognizes the `-wm` suffix: when the bundled Codex catalog ships the plain SKU, the `-wm` row is also registered under its plain id (re-derived so the 1M-window floor and daybreak pricing keyed on that slug still apply). Unknown `-wm` SKUs keep their authoritative verbatim slug and non-worker models are untouched, so distinct models and other providers are unaffected.
organizeImports sorts the type GeneratedProvider specifier ahead of the named imports, and the formatter collapses the preferred-match arrow chain to one line. Both are biome safe fixes; no behavior change. Resolves the two biome errors that were leaving the 8833 branch's check gate red.
std::env::vars() panics the moment a host env key or value is not valid
Unicode, before any command can run. A corrupt GHOSTTY_BIN_DIR (bytes 9d d9 50)
staged by cmux/Ghostty tripped both sinks:
- pi-shell's session env copy in create_session_for_run (also merged PATH)
- brush-core's get_host_env_vars, which process builtins (sleep, timeout,
pgrep, ...) use to inherit the host env into the shells they build
Both now read via std::env::vars_os() and skip entries that cannot be decoded
as Unicode: a corrupt entry carries no usable meaning. PATH merge behavior is
unchanged. Regression tests inject a non-UTF-8 key and value and assert the
shell still starts and PATH survives.
Reported in issue #8925.
CoreWeave Serverless Inference (W&B Inference) is a reseller with a
rotating model menu, but its catalog entry omitted
dynamicModelsAuthoritative. Runtime /v1/models discovery therefore
merged into the frozen bundled slice instead of replacing it, so stale
ids (e.g. moonshotai/Kimi-K3, Kimi-K2.5) stayed selectable and 404 at
request time. Matches sibling resellers (baseten, gmi-cloud, aiand,
bedrock-mantle). Bundled models remain the offline/failure fallback via
the authoritativeFreshProviders gating in ModelRegistry.
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.
Three independent defects:
1. `generateSummary` serialized the entire span into one prompt with no budget
check. It now plans windows that fit the summarizer's context and folds them
with the update prompt that iterative compaction already uses, so a stranded
boundary is recovered instead of rejected. A provider that rejects a window
the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
beta-gated to 200k on OAuth credentials) halves what was actually sent and
re-plans, because only the rejection knows the real cap.
2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
appends to its own errors classified a deterministic 400 as a transient 503.
Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.
3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
an input that does not fit never fits; both layers now fail fast to the next
candidate.
The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.
Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
Some providers (notably Gemini) serialize array tool arguments with
flattened property paths — questions[0].id, questions[0].options[0].label —
instead of a nested questions array. The schema sees only unrecognized extra
keys and rejects the call (e.g. the ask tool).
Add a pre-validation normalization pass (alongside the existing LLM-quirk
passes) that rebuilds the nested structure. Conservative: fires only when a
key is a well-formed array-index path, preserves non-flattened siblings, and
aborts wholesale on any shape conflict so genuine schema mistakes still
surface as validation errors.
Fixes#8886
The composer accepts three encodings for Shift+Enter (kitty CSI-u, the
legacy \x1b[13;2~ form, and a bare LF from the iTerm2 mapping e.g. Claude
Code's /terminal-setup). The /tree selector only handled the kitty form and
silently routed a bare LF into the plain-Enter branch, so summarize-and-
switch never fired for those terminals.
Mirror the composer: fall through a bare LF to summarize-and-switch while
plain CR (or the decoded Enter key) still does a plain switch.
Fixes#8821
Persistence is lazy: getSessionFile() returns an allocated path from the
start, but the JSONL is only materialized once an assistant message (or an
explicit ensureOnDisk()) crosses the persistence gate. The shutdown banner
printed the resume hint based solely on the path, so any session that ended
before the first assistant message advertised a copy-pasteable command that
always fails with Session not found.
Gate the hint on the new SessionManager#isSessionOnDisk() (file exists in
the active storage backend) and add unit tests.
Fixes#8860
A boolean latch survived /new and unnamed session switches, so the
replacement session skipped titling and could inherit the previous
skill's title. Bind the latch and the apply check to the originating
session id.
maybeStartTitleGeneration used to fire again for every untitled skill
prompt, so a queued /skill: during the first title request could race
and rename the session. Latch until the first request settles.
Session titles ignored /skill:<name> <args> because the invocation is a
custom skill-prompt, not a user turn. Feed the chip or reconstructed
/skill line into first-title generation and replan context, never the
expanded SKILL.md body.