- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
- Deleted the IPv6-wildcard coexistence test: on Linux `::` binds dual-stack
by default, so the flow's 127.0.0.1 bind genuinely conflicts and port
fallback is correct there — the test premise only holds on macOS.
- Stale /launch assertion no longer requires the freed ephemeral port to stay
unbound (parallel test files reclaim it); it now asserts the stale authorize
URL is never served again.
After a prewalk hand-off plus mid-run compaction, the openai-responses input
builder re-encoded replayed assistant turns via convertResponsesAssistantMessage,
demoting their reasoning to <think> output_text and emitting no reasoning item.
DeepSeek (opencode-go) then rejected the thinking-mode continuation with
400 "The reasoning_text in the thinking mode must be passed back to the API"
and OMP looped on the unchanged request.
The encoder now synthesizes a reasoning_text reasoning item for every replayed
assistant turn when the target requires reasoning replay in thinking mode
(requiresReasoningContentForAllAssistantTurns / requiresReasoningContentForToolCalls),
carrying surviving thinking text when present, mirroring the chat-completions
reasoning_content safety net. Gated on reasoning being active for the request,
so non-DeepSeek Responses targets and reasoning-disabled turns are unaffected.
Fixes#8248
- Default reasoning.summary to auto in openai-codex requests to ensure summaries are emitted by the backend.
- Gate stream_options reasoning_summary_delivery behind the PI_CODEX_CONCURRENT_SUMMARIES environment variable opt-in.
- Update tests in openai-codex-responses-lite to verify default summary behavior and opt-in concurrent delivery controls.
Added acceptEmptyResponse to the pi-native gateway option allow-list so a Google advisor on the auth-gateway transport still accepts a silent STOP server-side instead of exhausting the empty-response retries.
Added a pi-native parseRequest regression asserting the option survives the wire.
Gated empty-STOP silence acceptance on the last Cloud Code Assist endpoint so an earlier endpoint returning only empty streams still fails over instead of being recorded as a valid silent review.
Added an Antigravity auto-mode regression covering failover exhaustion before silence.
Tracked whether Cloud Code Assist removed a planning-leak payload so advisor silence acceptance cannot consume it as a valid empty review.
Updated the planning-leak regression to exercise acceptEmptyResponse and require the recovered function call.
Allowed advisor streams to treat Google STOP responses without visible content as successful silence while preserving the default retry behavior for interactive agents.
Added provider and advisor-path regressions covering retry counts, system instructions, and the advise declaration.
Fixes#8223
Mirror resolver readiness checks in the Bedrock availability probe so role profiles using Environment, EcsContainer, or Ec2InstanceMetadata are only advertised when the named source can produce credentials.
The AWS credential resolver only understood static keys, SSO, and
credential_process profiles, plus env-based web identity. The standard
EKS/IRSA and multi-account shape (role_arn + source_profile, chaining to
web_identity_token_file) was ignored, so hasConfiguredAwsProfile returned
false and amazon-bedrock was dropped before any request -- surfacing as
"No models available" while Claude Code and OpenCode worked on the same box.
- Resolve role_arn profiles via STS AssumeRole/AssumeRoleWithWebIdentity,
with source_profile recursion (cycle-guarded), credential_source
(Environment/Ec2InstanceMetadata/EcsContainer), and web_identity_token_file,
honoring role_session_name/duration_seconds/external_id.
- Mirror the same capability set in hasConfiguredAwsProfile so detection no
longer diverges from resolution.
- Broaden the EC2-host probe to recognize Nitro/EKS DMI markers so IMDS-backed
Bedrock is detected there.
Fixes#8209
Kimi's usage provider built window ids from the raw duration/timeUnit ("300time_unit_minute") and left the aggregate quota on "default", so the status-line usage segment, which matches canonical "5h"/"7d" windows, never rendered for kimi-code sessions.
Derive window ids from the reported span instead (300 minutes -> "5h", 7 days -> "7d"), mirroring the minimax-code convention, and tag the aggregate quota as the weekly window. The payload carries only resetTime for the aggregate, but the observed reset horizon matches the plan's weekly cycle. The segment aggregator also falls back to the reported span (duration within a minute of 5h or 7d) when the id isn't canonical, so cache rows written by older builds still light up the 5h meter in mixed-version fleets.
Cursor bash/grep frames wrote optional kwargs as present-undefined, and
tools.format gemini projectors dropped kCursorExecResolved so settled
calls ran twice.
Co-authored-by: Cursor <cursoragent@cursor.com>
The fetch layer recomputes Bedrock's content-length from the serialized body,
so a caller value would be covered by the SigV4 signature but never sent, and
AWS rejects the mismatch. Cursor streams its Connect body after the headers
(initial frame, heartbeats, tool responses), so no caller-supplied length can
describe it and an HTTP/2 peer resets the stream once the body diverges.
Reserved in both, and both regressions now send a Content-Length to prove it
never reaches the wire.
Three review findings.
host is transport-owned even though this request never sets it: node's http2
client suppresses the :authority it derives from the URL when a plain host
header is present, so a caller value silently retargeted the request at another
virtual host. Added to the reserved set.
The tests exercised only the exported sanitizer, so if streamCursor stopped
merging callerHeaders every one of them would still have passed while the wire
lost the headers. They now drive streamCursor against a local HTTP/2 server and
assert what the server actually received: an ordinary header arrives, names are
lower-cased, connection-specific and pseudo-headers neither arrive nor kill the
request, the request's own headers are not overridable, and the authority is
not retargetable. All five fail if the merge is removed. That also lets
sanitizeCursorCallerHeaders stop being exported.
The changelog entry pointed at #8046, the parent work, rather than the PR being
released here, and lacked the contributor credit the external-contribution
convention requires.
Both transports built their request header maps from scratch and never read
options.headers, so caller-supplied tracing or attribution headers were
silently dropped while working on every other provider. Bedrock now merges
caller headers before SigV4 signing so the signature covers them; Cursor
merges them into its HTTP/2 request.
Merging exposed three rules each transport enforces that a caller cannot be
allowed to violate, all covered by regressions that fail against the previous
code:
- SigV4 generates host/x-amz-date/x-amz-content-sha256/x-amz-security-token
for itself. signRequest signs a caller value but returns its own, so a
caller supplying any of them would sign one set of bytes and send another.
- HTTP/2 forbids the HTTP/1 connection-specific headers, and node's
http2.request() throws on them rather than ignoring them.
- Header names are compared case-insensitively on the wire, so a caller
Content-Type beside a fixed content-type is a duplicate rather than an
override. Caller names are lower-cased and copies of headers each request
sets itself are dropped.
The callback flow advertised `http://localhost:<port>/callback` but bound only
`127.0.0.1`. `localhost` resolves to both loopback families and clients try
`::1` first, so any process holding the IPv6 loopback on that port received the
authorization code instead. Nothing detected it either: a specific-address bind
coexists with another process's wildcard bind, so `Bun.serve` reported the port
as free and the random-port fallback never ran. The login then waited out the
full 5-minute timeout while the browser showed the other process's response.
A dev server is the common case. `next dev` binds `*:3000`, which is exactly
the MCP OAuth flow's default callback port.
Bind both loopback literals instead of probing. The kernel routes a connection
to the most specific matching bind, so the `::1` listener reclaims `localhost`
traffic from a wildcard-bound process, and both listeners answer the same
routes. A genuine collision on one exact loopback address still raises
EADDRINUSE and reaches the existing in-use policy; a host that cannot bind
`::1` at all keeps serving on IPv4 alone.
- Added account-scoped policy error detection to correctly identify Codex cyber-policy rejections.
- Updated credential storage and retry logic to route denied accounts through sibling rotation instead of bypassing it.
- Ensured coding-agent sessions exhaust all sibling accounts before falling back on cyber denials.
- Added comprehensive test coverage for credential rotation and retry behavior on policy errors.
Gate status-line monthly rendering to Cursor, prefer personal dashboard
rails over legacy /auth/usage request fractions, fall back from unusable
overall buckets to plan, ignore disabled plan percent fields, and keep
on-demand when the included bucket is empty. Also add changelog PR/author
attribution for the external contribution.
Cursor Pro/Pro+ usage-summary still uses individualUsage.plan, but the
web dashboard percent is autoPercentUsed / apiPercentUsed — not
plan.used/limit cents. Prefer those rails for Cursor Models / Other
Models, keep overall+cents fallbacks, floor monthly status-line %, and
repaint after async usage fetch so mo N% does not stay blank.
Cursor's /api/usage-summary now returns individualUsage.plan (and optional
onDemand) for Pro/Pro+/Ultra instead of the overall bucket that #7613
parsed. Fall back to plan when overall is absent, keep overall preferred
when both exist, and render monthly Cursor quotas in the status-line usage
segment.
Verified locally against a live Pro+ account and with focused unit tests.
- Replace the compat cast with an annotated CompatOf<Api> local narrowed
via the in operator: assignment up-cast, compiler-checked field type,
no as assertion.
- Cover the 0 sentinel end-to-end through the lazy wrapper with fake
timers (mirrors the direct-Anthropic 0-disable test): advance 400s
past the generic budget, assert no watchdog abort, then cancel
cleanly.
- Add Unreleased changelog entries for pi-ai and pi-catalog with
external attribution for #7892.
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).
- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
and lazy terminal errors carry the structural errorId classification
so session auto-retry classifies stalls without text matching.
#purgeSupersededDisabledRows matched disabled rows against active ones by
exact identity-key string equality, but the active-replacement path it
mirrors (matchesReplacementCredential) claims pre-org legacy rows
(`<b>` vs `<b>|org:<o>`). So a later org-scoped login of the same account
never purged the pre-org tombstone, which then rendered forever as a red
row in `omp usage` with no CLI/TUI escape. This is the OAuth half of the
class of bug #2943 fixed for api_key rows in the same function.
Reuse matchesReplacementCredential in the purge so an org-scoped login
claims and hard-deletes its pre-org tombstone, inheriting the one-way
upgrade and shared-workspace guards unchanged.
Fixes#7876
resolveAnthropicBaseUrl() resolved the chat base URL from github-copilot,
FOUNDRY_BASE_URL, then model.baseUrl -> hardcoded api.anthropic.com, and
never read $env.ANTHROPIC_BASE_URL. The stock anthropic descriptor pins
model.baseUrl to api.anthropic.com, so the env fallback was unreachable:
gateway-scoped keys were sent to api.anthropic.com (401, credential leak)
regardless of ANTHROPIC_BASE_URL, contradicting docs and the web-search
fix in #1693.
- Chat resolver now returns ANTHROPIC_BASE_URL (after Foundry, ahead of the
official default); an explicit non-official model.baseUrl still wins.
- resolveAnthropicCustomHeaders keys off the resolved base URL so
ANTHROPIC_CUSTOM_HEADERS reach env-configured non-official gateways.
- stream.ts leaked-thinking heal exemption mirror updated to the same
effective-endpoint precedence.
Fixes#7874