- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
Follow-up fix round for 71d608aae6 (validate PR review comment
anchors against the diff before submitting). The validation path
broke in several places the original commit did not cover:
- GitHubBackend.submit_review gained no commit_id parameter, so
submitting through GitHubProxyClient raised a production
TypeError on a path the validation code depended on.
- PullRequestInfo lacked head_sha, so the Forgejo commit_id
fallback raised AttributeError; it is now parsed in
_pr_from_payload and carried through the proxy round-trip.
- _pr_file_from dropped the file patch, silently no-oping anchor
validation for anything routed through the proxy; the patch is
now forwarded.
- The hunk parser treated any +++/--- line as a file header,
desyncing line counters when added/removed content began with
those prefixes; file headers are now recognized only before
the first hunk.
- Reworded the comment to "Forgejo only" to match the actual
backend behavior.
Adds tests for forgejo commit_id fetch, fallback double-failure,
empty-patch fail-open, LEFT-side anchoring, proxy commit_id
validation, and file-creation hunk boundaries.
Review follow-up: a safe `-wm` row previously kept backend-parsed capability metadata (no 1M floor, no daybreak pricing) while its synthesized plain listing was enriched — the same model reported two different contexts. Both listings now derive fallback window, the 1M floor, and daybreak cost from the canonical plain slug; unknown `-wm` SKUs and non-worker models keep their verbatim slug-derived metadata.
The authoritative-discovery test now resolves the configured `openai-codex/gpt-5.6-luna` through the real model resolver and asserts the exact bound id (and that an explicit `-wm` config still resolves verbatim) instead of only checking `ids.toContain`.
Codex backend discovery advertises worker-mode SKUs under a `-wm` suffix (gpt-5.6-luna-wm). Authoritative discovery replaced the bundled catalog and kept those slugs verbatim, so a configured `openai-codex/gpt-5.6-luna` vanished from the resolved catalog and the resolver's fuzzy fallback selected `-wm` instead — a route some ChatGPT accounts reject.
Model discovery now recognizes the `-wm` suffix: when the bundled Codex catalog ships the plain SKU, the `-wm` row is also registered under its plain id (re-derived so the 1M-window floor and daybreak pricing keyed on that slug still apply). Unknown `-wm` SKUs keep their authoritative verbatim slug and non-worker models are untouched, so distinct models and other providers are unaffected.
organizeImports sorts the type GeneratedProvider specifier ahead of the named imports, and the formatter collapses the preferred-match arrow chain to one line. Both are biome safe fixes; no behavior change. Resolves the two biome errors that were leaving the 8833 branch's check gate red.
std::env::vars() panics the moment a host env key or value is not valid
Unicode, before any command can run. A corrupt GHOSTTY_BIN_DIR (bytes 9d d9 50)
staged by cmux/Ghostty tripped both sinks:
- pi-shell's session env copy in create_session_for_run (also merged PATH)
- brush-core's get_host_env_vars, which process builtins (sleep, timeout,
pgrep, ...) use to inherit the host env into the shells they build
Both now read via std::env::vars_os() and skip entries that cannot be decoded
as Unicode: a corrupt entry carries no usable meaning. PATH merge behavior is
unchanged. Regression tests inject a non-UTF-8 key and value and assert the
shell still starts and PATH survives.
Reported in issue #8925.
CoreWeave Serverless Inference (W&B Inference) is a reseller with a
rotating model menu, but its catalog entry omitted
dynamicModelsAuthoritative. Runtime /v1/models discovery therefore
merged into the frozen bundled slice instead of replacing it, so stale
ids (e.g. moonshotai/Kimi-K3, Kimi-K2.5) stayed selectable and 404 at
request time. Matches sibling resellers (baseten, gmi-cloud, aiand,
bedrock-mantle). Bundled models remain the offline/failure fallback via
the authoritativeFreshProviders gating in ModelRegistry.
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.
Three independent defects:
1. `generateSummary` serialized the entire span into one prompt with no budget
check. It now plans windows that fit the summarizer's context and folds them
with the update prompt that iterative compaction already uses, so a stranded
boundary is recovered instead of rejected. A provider that rejects a window
the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
beta-gated to 200k on OAuth credentials) halves what was actually sent and
re-plans, because only the rejection knows the real cap.
2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
appends to its own errors classified a deterministic 400 as a transient 503.
Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.
3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
an input that does not fit never fits; both layers now fail fast to the next
candidate.
The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.
Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
Some providers (notably Gemini) serialize array tool arguments with
flattened property paths — questions[0].id, questions[0].options[0].label —
instead of a nested questions array. The schema sees only unrecognized extra
keys and rejects the call (e.g. the ask tool).
Add a pre-validation normalization pass (alongside the existing LLM-quirk
passes) that rebuilds the nested structure. Conservative: fires only when a
key is a well-formed array-index path, preserves non-flattened siblings, and
aborts wholesale on any shape conflict so genuine schema mistakes still
surface as validation errors.
Fixes#8886
The composer accepts three encodings for Shift+Enter (kitty CSI-u, the
legacy \x1b[13;2~ form, and a bare LF from the iTerm2 mapping e.g. Claude
Code's /terminal-setup). The /tree selector only handled the kitty form and
silently routed a bare LF into the plain-Enter branch, so summarize-and-
switch never fired for those terminals.
Mirror the composer: fall through a bare LF to summarize-and-switch while
plain CR (or the decoded Enter key) still does a plain switch.
Fixes#8821
Persistence is lazy: getSessionFile() returns an allocated path from the
start, but the JSONL is only materialized once an assistant message (or an
explicit ensureOnDisk()) crosses the persistence gate. The shutdown banner
printed the resume hint based solely on the path, so any session that ended
before the first assistant message advertised a copy-pasteable command that
always fails with Session not found.
Gate the hint on the new SessionManager#isSessionOnDisk() (file exists in
the active storage backend) and add unit tests.
Fixes#8860
A boolean latch survived /new and unnamed session switches, so the
replacement session skipped titling and could inherit the previous
skill's title. Bind the latch and the apply check to the originating
session id.
maybeStartTitleGeneration used to fire again for every untitled skill
prompt, so a queued /skill: during the first title request could race
and rename the session. Latch until the first request settles.
Session titles ignored /skill:<name> <args> because the invocation is a
custom skill-prompt, not a user turn. Feed the chip or reconstructed
/skill line into first-title generation and replan context, never the
expanded SKILL.md body.
The post-interrupt immune-window downgraded end-of-turn blockers to non-interrupting asides, so a blocker that means the agent handed off broken work never woke a new turn. Concerns keep the cooldown; blockers now bypass it and steer a triggered turn, consistent with the #5628 blocker-after-terminal-answer exception.
Apply the same recursive placeholder expansion used by native MCP configs before extension-package servers are validated and surfaced. This prevents stdio credentials and remote headers from reaching servers as literal placeholders.\n\nSolves: Extension-package MCP environment expansion\nTests: bun test packages/coding-agent/test/discovery/omp-plugins.test.ts
Keep host-specific secret injection in the maintained plugin layer instead
of carrying a behavioral divergence in the OMP fork.
Solves: Unwanted fork maintenance for MCP injection
Tests: Reverts only drycode/oh-my-pi PR #1
Claude marketplace plugins may reference runtime secrets in stdio
server environment values. Resolve those placeholders after plugin-root
substitution so child processes receive credentials instead of literal
template strings.
Solves: Stdio MCP credentials remain unexpanded
Tests: Claude plugin discovery tests; coding-agent check; live CH auth
The idle watchdog aborts the request signal and cursor.ts closes
the Connect stream, so there is no in-flight server exec to race.
Unmarked MCP/todo blocks can continue once every emitted call has
a matching result, same as HTTP/2 RST.
Cursor hosted web search / Exa / unnamed field-9 WebFetch send
interaction_query and block the Run RPC until the client writes
interaction_response. Dropping the frame leaves the HTTP/2 stream
alive on heartbeats that are not semantic progress, so the 300s
idle watchdog aborts with "Provider stream stalled while waiting
for the next event".
Approve network permission gates and reject interactive
ask / switch-mode / create-plan. Leave VM setup unanswered
rather than inventing a success result.
A single comment on a line outside a diff hunk 422s the whole review on
GitHub, which the 422/500 fallback (ported here from the Forgejo work,
including payload shaping and commit_id) then degraded into plain issue
comments. Now we fetch /pulls/{n}/files, compute per-hunk anchorable line
sets, drop non-anchorable comments, fold them into the review summary, and
only fall back to issue comments for genuine 422/500s that survive
filtering. Files-fetch failure fails open (submit unfiltered).
Retry: widened agent dequeue-hook deadline budgets from 25ms to 1s — the run loop checks the deadline before invoking dequeue hooks, so a cold or CPU-starved mock roundtrip expired the deadline first and the hooks never ran (deterministic failure in isolation, flaky under CI parallel load).
Workers spawned via resolveWorkerSpawnCmd ran with cwd anchored at the CLI
install directory and shared the agent's foreground process group. Terminal
cwd heuristics such as kitty's new_tab_with_cwd read the newest process in
that group, so new terminal tabs opened in
~/.bun/install/global/node_modules/@oh-my-pi/pi-coding-agent/dist while any
worker was alive. Spawn workers with the absolute host entry and inherit the
agent cwd instead; the bun-test fallback branch keeps its cwd-relative form.
Retry: deflaked tui IME preedit border test (#5563) — VirtualTerminal.waitForRender's fixed 40ms sleep raced TUI's throttled render timer on starved CI runners, reading the pre-input frame; waitForRender now takes an optional settle predicate polled up to 2s and the test keys on the rendered content.
browseHtmlPage wrapped its navigations in untilAborted(signal) but awaited
applyViewport, applyStealthPatches, and the finally-block page.close() raw.
When the shared headless daemon or the page target dies mid-setup, those
puppeteer calls never settle: the search hard timeout fires into a signal
with no listener at those await points, the provider promise hangs past
SEARCH_HARD_TIMEOUT_MS, and the turn never ends (only kill -9 recovers).
- Wrap applyViewport/applyStealthPatches in untilAborted(signal) so the
existing hard timeout can abort a dead-session setup.
- Bound the teardown page.close() with a fresh 5s deadline; .catch() only
covers rejection, not a hang, and the caller signal may already be fired.
Fixes#8865
The startup graphics probe only ran on ConPTY hosts with WT_SESSION, so a
SIXEL-capable terminal that exposes no identifying environment variable
(foot exports TERM=foot and COLORTERM=truecolor only) resolved the
trueColor capability row, kept imageProtocol null, and rendered every
image as the "[Image: …]" text card.
The XTSMGRAPHICS branch also had its status inverted: per xterm ctlseqs a
reply of `CSI ? 2 ; Ps ; Pv S` carries Ps = 0 on success, and a terminal
without SIXEL reports a zero maximum geometry, so a successful reply was
read as unsupported.
Drop the dead DA1 half of the probe with it: ProcessTerminal swallows
every `CSI ? … c` reply for the whole session so a late one cannot leak
into the composer (#8542), which means the attribute list never reached
the probe's input listener on any platform. The bare `CSI c` it wrote was
also unaccounted for in the DA1 sentinel FIFO, so its reply consumed
another probe's sentinel.
PI_FORCE_IMAGE_PROTOCOL, including its off/none kill switch, still wins
over the probe.
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
The exact-echo check used accent-insensitive compare, but the collision
suffix path was a case-sensitive startsWith. AuthLoader-3 vs authloader
therefore leaked through as a real description.
The first HUD commit hid Name: Name. The cause was earlier: task
name was copied into identity.label, which became progress.description
and skipped generateTaskLabel. Keep the handle for id allocation, but
only treat eval label as a real UI description so the tiny-model
summary can run.
The anchored Subagents list printed only `Id: description` and treated a
label that repeated the spawn handle as a real description. Show the same
⟨role⟩ badge as inline task rows and omit descriptions that only echo the id.