Codex Responses Lite only accepts text input, so GPT image prompts with input_image failed before reaching the full multimodal transport.
The request transformer and Codex headers now fall back to full Responses whenever a request body contains input_image, with regression coverage for the lite header.
Fixes#3421
Mapped cacheRetention="long" to cache_control.ttl="1h" on OpenRouter Anthropic Responses requests; guarded against overwriting a caller-supplied cache_control.\n\nRefs #3397
When the inner-terminal env markers (KITTY_WINDOW_ID, GHOSTTY_RESOURCES_DIR,
WEZTERM_PANE, ITERM_SESSION_ID) leak into a tmux session, TERMINAL_ID
resolves to the inner terminal and notifyProtocol becomes OSC 9 / OSC 99.
sendNotification() wrote that raw OSC straight to stdout, but tmux does not
forward bare OSC 9/99 to the outer terminal and the bare sequence does not
flag tmux's own monitor-bell / monitor-activity — so a backgrounded omp pane
under the common tmux + kitty/ghostty/wezterm/iTerm2 stack had no signal at
all for completion or 'ask' blockage.
Under TMUX, OSC-protocol notifications are now wrapped in tmux's
\x1bPtmux;…\x1b\\ DCS passthrough envelope (so users with
allow-passthrough on still get the real desktop toast) and followed by a
\x07 BEL (so monitor-bell flags the window otherwise — the standard
'needs attention' workflow). The OSC 99 capability probe is wrapped the
same way so rich notifications keep working across tmux. Bell-protocol
paths are unchanged.
Fixes#3395
Ollama Cloud rejects any chat request whose options.num_predict exceeds
65536 with HTTP 400, but the existing safety relied on the load-time
omitMaxOutputTokens policy in model-registry.ts. Stale models.db rows
predating that policy (and custom modelOverrides re-enabling output
caps) carried maxTokens: 1048576 forward to the wire layer untouched,
so every request to deepseek-v4-pro / deepseek-v4-flash 400'd with:
max_tokens (1048576) exceeds model's maximum output tokens (65536)
for model deepseek-v4-pro
createChatBody now resolves num_predict through resolveNumPredict,
which clamps every ollama-cloud request at the documented cap before
serialization — independent of the cached spec, on top of the existing
omitMaxOutputTokens path. Self-hosted ollama traffic is unaffected.
Fixes#3392
`buildAnthropicHeaders` filtered `Authorization` and `X-Api-Key` out of
`model.headers` (and `ANTHROPIC_CUSTOM_HEADERS`) via `enforcedHeaderKeys`,
then re-attached the auto-built `Bearer <apiKey>` / `X-Api-Key: <apiKey>`,
so custom-proxy users could never override anthropic-messages auth — unlike
`openai-responses` (`headers.Authorization ??= Bearer apiKey`) and the
GitHub Copilot anthropic branch (`mergeHeaders(..., model.headers, ...)`).
Non-OAuth, non-Cloudflare branches now honor the caller's `Authorization` /
`X-Api-Key`; case-insensitive capture keeps the spread free of duplicate
keys. OAuth still forces `Authorization: Bearer <oauth-token>` (the OAuth
credential itself) and Cloudflare AI Gateway still owns
`cf-aig-authorization` — both drop + log caller overrides.
`buildAnthropicClientOptions` no longer requires the surviving
`Authorization` to start with `Bearer ` before suppressing the
client-level `X-Api-Key`; any custom Authorization on a non-official
endpoint suppresses it so the proxy gets a single credential.
Fixes#3391
Trailing empty assistant 'stop' arriving after a successful 'yield'
revived the already-yielded subagent. AgentSession.agent_end maintenance
compared #assistantEndedWithSuccessfulYield(msg) against the trailing
empty-stop message — not the yield-bearing one — so the empty-stop
recovery path appended a retry reminder and scheduled agent.continue().
Track a sticky #yieldTerminationPending flag set when the yield tool
finishes without error and cleared on the next #promptWithMessage. The
agent_end routing extends the existing successful-yield branch: when the
flag is set, or the current message ended with yield, short-circuit
empty-stop / unexpected-stop / compaction continuations for the rest of
the run, so a successful yield is terminal regardless of trailing stops.
Fixes#3389
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.
- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
forces input=['text'] whenever the resolved baseUrl is not the
canonical personal-Copilot host, so the upstream's vision flag is
honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
OR-upgrading with the bundled reference) when the merged baseUrl
differs from the bundled one, so a bundled spec pinned to the
personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
rasterizer for any github-copilot model whose baseUrl is non-personal,
catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
pi-catalog/wire/github-copilot so catalog and coding-agent share one
canonical check.
Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).
Fixes#3387
With collapseCompactedHistory the live display fell into the LLM compaction
branch, which skips the firstKeptEntryId..compaction turns whenever an OpenAI
remote-compaction replacementHistory payload is present. That payload feeds the
provider only and is not rendered, so a remotely-compacted session showed just
the summary plus post-compaction rows, hiding recent turns that were visible
before. Emit the kept SessionEntry rows in transcript mode regardless. Adds a
regression.
The append fast-path opened the session file for sentinel comparison and
recompute without a guard, so a file unlinked/rotated between #refresh's
statSync and the sentinel read threw out of the 250ms poll timer (no catch),
risking a TUI crash. Treat sentinel read/recompute failures as non-appendable
and fall back to the guarded full reload. Adds a fake-timer regression.
The "skips snapcompact entirely" case asserted .rejects.toThrow() on
session.compact(), but after the maxFrames<1 skip the manual /compact path
falls through to the LLM summarizer whose outcome is provider/network
dependent (resolves when a summary lands, rejects only without
credentials/network). In a sandbox it resolves and blows the 5s default
timeout. Tolerate either outcome and pin only the deterministic skip
contract (snapcompact.compact not invoked + the kept-history notice), with
a generous explicit timeout. Addresses chatgpt-codex P2 review on #3249.
renderUsageReports (command-controller) carried the #3268 dedup contract
with no regression test; the PR's added CLI test asserts the opposite
(per-limit CLI rendering shows the note twice). Export renderUsageReports
and add a real regression through it: two accounts sharing one window group
render a provider-wide UsageReport.note once and an identical per-limit note
once. Verified failing on the pre-fix flatMap form (0 and 2 occurrences) and
passing on head (1 and 1).
The PR's absolute-path tests pass identically on baseline and head on
POSIX CI (they use a /-prefixed outsideDir), so they do not protect the
Windows fix on Ubuntu. This adds a relative @..\\ backslash-prefix test
that fails on the pre-fix code (empty results on POSIX) and passes after
the backslash-normalization fix, giving CI real regression coverage.
Reviewer flagged both as unrelated to the TUI OSC 11 fix. The blank-line addition reverts the PR branch's coding-agent CHANGELOG back to its merge-base blob (the package is otherwise untouched by this PR); the empty file `0` was accidental.