Commit Graph

2129 Commits

Author SHA1 Message Date
can1357 10fd42289c feat: introduced external thinking support and private scratchpad think tool
- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
2026-08-11 20:39:57 +02:00
can1357 d3b22a0db6 test(ai): removed env-dependent dual-stack test, made launch-route port-reuse safe
- Deleted the IPv6-wildcard coexistence test: on Linux `::` binds dual-stack
  by default, so the flow's 127.0.0.1 bind genuinely conflicts and port
  fallback is correct there — the test premise only holds on macOS.
- Stale /launch assertion no longer requires the freed ephemeral port to stay
  unbound (parallel test files reclaim it); it now asserts the stale authorize
  URL is never served again.
2026-08-11 16:31:23 +02:00
roboomp 54ce9e47fc fix(ai): synthesized reasoning_text on DeepSeek Responses replay
After a prewalk hand-off plus mid-run compaction, the openai-responses input
builder re-encoded replayed assistant turns via convertResponsesAssistantMessage,
demoting their reasoning to <think> output_text and emitting no reasoning item.
DeepSeek (opencode-go) then rejected the thinking-mode continuation with
400 "The reasoning_text in the thinking mode must be passed back to the API"
and OMP looped on the unchanged request.

The encoder now synthesizes a reasoning_text reasoning item for every replayed
assistant turn when the target requires reasoning replay in thinking mode
(requiresReasoningContentForAllAssistantTurns / requiresReasoningContentForToolCalls),
carrying surviving thinking text when present, mirroring the chat-completions
reasoning_content safety net. Gated on reasoning being active for the request,
so non-DeepSeek Responses targets and reasoning-disabled turns are unaffected.

Fixes #8248
2026-08-11 15:43:41 +02:00
can1357 f22cb606bf feat(ai): enabled automatic reasoning summaries in codex requests
- Default reasoning.summary to auto in openai-codex requests to ensure summaries are emitted by the backend.
- Gate stream_options reasoning_summary_delivery behind the PI_CODEX_CONCURRENT_SUMMARIES environment variable opt-in.
- Update tests in openai-codex-responses-lite to verify default summary behavior and opt-in concurrent delivery controls.
2026-08-11 15:29:26 +02:00
can1357 733c3271bb Merge PR #8212: fix(ai): resolve AWS role_arn/source_profile role chains for Bedrock (@roboomp) 2026-08-11 15:15:45 +02:00
can1357 3b66177abd Merge PR #8226: fix(advisor): accept silent Gemini reviews (@roboomp) 2026-08-11 15:15:44 +02:00
can1357 64baa7c1bd chore(format): applied biome formatting and removed dead code from merged prs 2026-08-11 15:14:15 +02:00
can1357 f95ba550b5 Merge PR #8228: fix(ai): support Daybreak Responses Lite (@pickpocket) 2026-08-11 15:06:17 +02:00
can1357 24c348862f Merge PR #8197: fix(status-line): recognized kimi-code subscription windows in the usage segment (@tehfiend) 2026-08-11 15:06:16 +02:00
can1357 270fe8d454 Merge PR #8183: fix(cursor): omit undefined exec args and keep exec-resolved under owned dialects (@jairuspace) 2026-08-11 15:06:16 +02:00
can1357 814f31bb6b Merge PR #8130: fix(ai): fail over stalled Antigravity streams (@usr-bin-roygbiv) 2026-08-11 15:06:14 +02:00
can1357 42357889b0 test(ai): isolated Bedrock bearer credentials 2026-08-11 15:06:14 +02:00
can1357 cd7a86423f Merge PR #8107: fix(ai): honor StreamOptions.headers in Bedrock and Cursor (@svperfecta) 2026-08-11 15:06:14 +02:00
can1357 4644bf4f99 Merge PR #8081: fix(ai): bind the IPv6 loopback so dev servers stop stealing OAuth callbacks (@jwaldrip) 2026-08-11 15:06:14 +02:00
can1357 dedddb86f5 Merge PR #7998: fix(ai): map Cursor plan rails to dashboard percents and show mo N% (@dnth) 2026-08-11 15:06:12 +02:00
roboomp 433ba7df1f fix(ai): forwarded advisor silence flag over pi-native
Added acceptEmptyResponse to the pi-native gateway option allow-list so a Google advisor on the auth-gateway transport still accepts a silent STOP server-side instead of exhausting the empty-response retries.

Added a pi-native parseRequest regression asserting the option survives the wire.
2026-08-11 08:46:35 +00:00
pickpocket b6818ab603 fix(ai): support Daybreak Responses Lite 2026-08-11 04:43:45 -04:00
roboomp 0a64b6d965 fix(ai): kept antigravity failover before advisor silence
Gated empty-STOP silence acceptance on the last Cloud Code Assist endpoint so an earlier endpoint returning only empty streams still fails over instead of being recorded as a valid silent review.

Added an Antigravity auto-mode regression covering failover exhaustion before silence.
2026-08-11 08:39:08 +00:00
roboomp c396e38726 fix(ai): preserved advisor planning leak retries
Tracked whether Cloud Code Assist removed a planning-leak payload so advisor silence acceptance cannot consume it as a valid empty review.

Updated the planning-leak regression to exercise acceptEmptyResponse and require the recovered function call.
2026-08-11 08:32:09 +00:00
roboomp 86b8f510c8 fix(advisor): accepted silent Gemini reviews
Allowed advisor streams to treat Google STOP responses without visible content as successful silence while preserving the default retry behavior for interactive agents.

Added provider and advisor-path regressions covering retry counts, system instructions, and the advise declaration.

Fixes #8223
2026-08-11 08:22:54 +00:00
roboomp 3e2b431680 fix(ai): gated AWS credential_source availability
Mirror resolver readiness checks in the Bedrock availability probe so role profiles using Environment, EcsContainer, or Ec2InstanceMetadata are only advertised when the named source can produce credentials.
2026-08-11 05:29:55 +00:00
roboomp ba8bb17fa3 fix(ai): resolve AWS role_arn/source_profile role chains for Bedrock
The AWS credential resolver only understood static keys, SSO, and
credential_process profiles, plus env-based web identity. The standard
EKS/IRSA and multi-account shape (role_arn + source_profile, chaining to
web_identity_token_file) was ignored, so hasConfiguredAwsProfile returned
false and amazon-bedrock was dropped before any request -- surfacing as
"No models available" while Claude Code and OpenCode worked on the same box.

- Resolve role_arn profiles via STS AssumeRole/AssumeRoleWithWebIdentity,
  with source_profile recursion (cycle-guarded), credential_source
  (Environment/Ec2InstanceMetadata/EcsContainer), and web_identity_token_file,
  honoring role_session_name/duration_seconds/external_id.
- Mirror the same capability set in hasConfiguredAwsProfile so detection no
  longer diverges from resolution.
- Broaden the EC2-host probe to recognize Nitro/EKS DMI markers so IMDS-backed
  Bedrock is detected there.

Fixes #8209
2026-08-11 05:21:03 +00:00
Bejay Cole e7f487381b fix(status-line): recognized kimi-code subscription windows in the usage segment
Kimi's usage provider built window ids from the raw duration/timeUnit ("300time_unit_minute") and left the aggregate quota on "default", so the status-line usage segment, which matches canonical "5h"/"7d" windows, never rendered for kimi-code sessions.

Derive window ids from the reported span instead (300 minutes -> "5h", 7 days -> "7d"), mirroring the minimax-code convention, and tag the aggregate quota as the weekly window. The payload carries only resetTime for the aggregate, but the observed reset horizon matches the plan's weekly cycle. The segment aggregator also falls back to the reported span (duration within a minute of 5h or 7d) when the id isn't canonical, so cache rows written by older builds still light up the 5h meter in mixed-version fleets.
2026-08-10 17:01:41 -06:00
jairuspace 5cbd482f8e fix(cursor): omit undefined exec args and keep exec-resolved under owned dialects
Cursor bash/grep frames wrote optional kwargs as present-undefined, and
tools.format gemini projectors dropped kCursorExecResolved so settled
calls ran twice.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-10 13:39:40 -06:00
usr-bin-roygbiv c5d1a0aa9c fix(ai): fail over stalled Antigravity streams 2026-08-10 01:58:11 +00:00
Brian Corrigan 3bfce1d7b4 Treat content-length as transport-owned in both providers
The fetch layer recomputes Bedrock's content-length from the serialized body,
so a caller value would be covered by the SigV4 signature but never sent, and
AWS rejects the mismatch. Cursor streams its Connect body after the headers
(initial frame, heartbeats, tool responses), so no caller-supplied length can
describe it and an HTTP/2 peer resets the stream once the body diverges.

Reserved in both, and both regressions now send a Content-Length to prove it
never reaches the wire.
2026-08-09 15:32:48 -06:00
Brian Corrigan 9dc3594187 Test the Cursor transport contract, and treat host as transport-owned
Three review findings.

host is transport-owned even though this request never sets it: node's http2
client suppresses the :authority it derives from the URL when a plain host
header is present, so a caller value silently retargeted the request at another
virtual host. Added to the reserved set.

The tests exercised only the exported sanitizer, so if streamCursor stopped
merging callerHeaders every one of them would still have passed while the wire
lost the headers. They now drive streamCursor against a local HTTP/2 server and
assert what the server actually received: an ordinary header arrives, names are
lower-cased, connection-specific and pseudo-headers neither arrive nor kill the
request, the request's own headers are not overridable, and the authority is
not retargetable. All five fail if the merge is removed. That also lets
sanitizeCursorCallerHeaders stop being exported.

The changelog entry pointed at #8046, the parent work, rather than the PR being
released here, and lacked the contributor credit the external-contribution
convention requires.
2026-08-09 15:15:47 -06:00
Brian Corrigan 7527fd53a2 fix(ai): honor StreamOptions.headers in Bedrock and Cursor
Both transports built their request header maps from scratch and never read
options.headers, so caller-supplied tracing or attribution headers were
silently dropped while working on every other provider. Bedrock now merges
caller headers before SigV4 signing so the signature covers them; Cursor
merges them into its HTTP/2 request.

Merging exposed three rules each transport enforces that a caller cannot be
allowed to violate, all covered by regressions that fail against the previous
code:

- SigV4 generates host/x-amz-date/x-amz-content-sha256/x-amz-security-token
  for itself. signRequest signs a caller value but returns its own, so a
  caller supplying any of them would sign one set of bytes and send another.
- HTTP/2 forbids the HTTP/1 connection-specific headers, and node's
  http2.request() throws on them rather than ignoring them.
- Header names are compared case-insensitively on the wire, so a caller
  Content-Type beside a fixed content-type is a duplicate rather than an
  override. Caller names are lower-cased and copies of headers each request
  sets itself are dropped.
2026-08-09 14:10:12 -06:00
Jason Waldrip b11f3e6311 fix(ai): bound the IPv6 loopback so dev servers stop stealing OAuth callbacks
The callback flow advertised `http://localhost:<port>/callback` but bound only
`127.0.0.1`. `localhost` resolves to both loopback families and clients try
`::1` first, so any process holding the IPv6 loopback on that port received the
authorization code instead. Nothing detected it either: a specific-address bind
coexists with another process's wildcard bind, so `Bun.serve` reported the port
as free and the random-port fallback never ran. The login then waited out the
full 5-minute timeout while the browser showed the other process's response.

A dev server is the common case. `next dev` binds `*:3000`, which is exactly
the MCP OAuth flow's default callback port.

Bind both loopback literals instead of probing. The kernel routes a connection
to the most specific matching bind, so the `::1` listener reclaims `localhost`
traffic from a wildcard-bound process, and both listeners answer the same
routes. A genuine collision on one exact loopback address still raises
EADDRINUSE and reaches the existing in-use policy; a host that cannot bind
`::1` at all keeps serving on IPv4 alone.
2026-08-09 00:34:45 -06:00
can1357 ac55ea7697 fix(catalog): toggle qwen3.8 max thinking on wire 2026-08-08 19:38:31 +02:00
can1357 b28c2da29a Merge PR #8021: fix(catalog): correct qwen3.8 max discovery metadata (@roboomp) 2026-08-08 19:38:31 +02:00
can1357 7ca140f66e fix(ai): routed policy-rejected accounts through sibling rotation
- Added account-scoped policy error detection to correctly identify Codex cyber-policy rejections.
- Updated credential storage and retry logic to route denied accounts through sibling rotation instead of bypassing it.
- Ensured coding-agent sessions exhaust all sibling accounts before falling back on cyber denials.
- Added comprehensive test coverage for credential rotation and retry behavior on policy errors.
2026-08-08 19:29:49 +02:00
roboomp c0eda613c8 fix(catalog): kept qwen3.8 max preview on enable_thinking
Restored the preview to its main compat so the reasoning_effort dialect is scoped to qwen3.8-max, leaving the preview ladder unchanged.
2026-08-08 14:54:29 +00:00
roboomp 155fdaedba fix(catalog): corrected qwen3.8 max discovery metadata
Curated reasoning, multimodal input, context limits, and the provider-specific effort ladder for the discovered Alibaba Token Plan model.

Fixes #8019
2026-08-08 14:37:24 +00:00
dnth aa92c7e97c fix(ai): address Cursor usage review feedback
Gate status-line monthly rendering to Cursor, prefer personal dashboard
rails over legacy /auth/usage request fractions, fall back from unusable
overall buckets to plan, ignore disabled plan percent fields, and keep
on-demand when the included bucket is empty. Also add changelog PR/author
attribution for the external contribution.
2026-08-08 17:36:19 +08:00
dnth 69f790fa6e fix(ai): map Cursor plan rails to dashboard percents
Cursor Pro/Pro+ usage-summary still uses individualUsage.plan, but the
web dashboard percent is autoPercentUsed / apiPercentUsed — not
plan.used/limit cents. Prefer those rails for Cursor Models / Other
Models, keep overall+cents fallbacks, floor monthly status-line %, and
repaint after async usage fetch so mo N% does not stay blank.
2026-08-08 17:26:40 +08:00
dnth cc3835342e fix(ai): parse Cursor plan/onDemand usage and show monthly status-line
Cursor's /api/usage-summary now returns individualUsage.plan (and optional
onDemand) for Pro/Pro+/Ultra instead of the overall bucket that #7613
parsed. Fall back to plan when overall is absent, keep overall preferred
when both exist, and render monthly Cursor quotas in the status-line usage
segment.

Verified locally against a live Pro+ account and with focused unit tests.
2026-08-08 16:46:01 +08:00
can1357 85b5d0ac1f fix(ai): keep Cursor session tokens on default origin 2026-08-07 13:37:55 +02:00
can1357 f1ad890d3f Merge PR #7613: feat(ai): add Cursor personal usage reporting (@chessl) 2026-08-07 13:37:55 +02:00
can1357 10863ef6b4 Merge PR #7892: fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat (@voonfoo) 2026-08-07 13:37:54 +02:00
can1357 f087a9050d fix(ai): keep plan per-minute limits transient 2026-08-07 13:37:54 +02:00
can1357 68dece477f Merge PR #7772: fix(session): handle subscription-cap retry exhaustion (@roboomp) 2026-08-07 13:37:54 +02:00
can1357 3d12722dba fix(ai): keep Chinese transient caps out of quota rotation 2026-08-07 13:37:54 +02:00
can1357 50cc615e45 Merge PR #7828: fix(ai): classify Simplified Chinese quota exhaustion as credential-rotatable (@IceCodeNew) 2026-08-07 13:37:54 +02:00
can1357 6be1b41d3b Merge PR #7821: fix(ai): preserve Copilot integrator entitlement errors (@roboomp) 2026-08-07 13:37:53 +02:00
can1357 3c35180de7 Merge PR #7878: fix(auth): purge pre-org OAuth tombstone on org-scoped re-login (@roboomp) 2026-08-07 13:37:53 +02:00
Voon Foo 61ff318e6d review: typed compat access, 0-disable lazy watchdog test, changelogs
- Replace the compat cast with an annotated CompatOf<Api> local narrowed
  via the in operator: assignment up-cast, compiler-checked field type,
  no as assertion.
- Cover the 0 sentinel end-to-end through the lazy wrapper with fake
  timers (mirrors the direct-Anthropic 0-disable test): advance 400s
  past the generic budget, assert no watchdog abort, then cancel
  cleanly.
- Add Unreleased changelog entries for pi-ai and pi-catalog with
  external attribution for #7892.
2026-08-07 15:39:04 +08:00
Voon Foo 2f24d4457e fix(ai,catalog): widen Bedrock stream-stall watchdog via model compat
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).

- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
  a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
  Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
  tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
  and lazy terminal errors carry the structural errorId classification
  so session auto-retry classifies stalls without text matching.
2026-08-07 14:18:46 +08:00
roboomp 648d858bb1 fix(auth): purged pre-org oauth tombstone on org-scoped re-login
#purgeSupersededDisabledRows matched disabled rows against active ones by
exact identity-key string equality, but the active-replacement path it
mirrors (matchesReplacementCredential) claims pre-org legacy rows
(`<b>` vs `<b>|org:<o>`). So a later org-scoped login of the same account
never purged the pre-org tombstone, which then rendered forever as a red
row in `omp usage` with no CLI/TUI escape. This is the OAuth half of the
class of bug #2943 fixed for api_key rows in the same function.

Reuse matchesReplacementCredential in the purge so an org-scoped login
claims and hard-deletes its pre-org tombstone, inheriting the one-way
upgrade and shared-workspace guards unchanged.

Fixes #7876
2026-08-07 04:17:10 +00:00
roboomp 65f0d5869b fix(anthropic): honor ANTHROPIC_BASE_URL for chat requests
resolveAnthropicBaseUrl() resolved the chat base URL from github-copilot,
FOUNDRY_BASE_URL, then model.baseUrl -> hardcoded api.anthropic.com, and
never read $env.ANTHROPIC_BASE_URL. The stock anthropic descriptor pins
model.baseUrl to api.anthropic.com, so the env fallback was unreachable:
gateway-scoped keys were sent to api.anthropic.com (401, credential leak)
regardless of ANTHROPIC_BASE_URL, contradicting docs and the web-search
fix in #1693.

- Chat resolver now returns ANTHROPIC_BASE_URL (after Foundry, ahead of the
  official default); an explicit non-official model.baseUrl still wins.
- resolveAnthropicCustomHeaders keys off the resolved base URL so
  ANTHROPIC_CUSTOM_HEADERS reach env-configured non-official gateways.
- stream.ts leaked-thinking heal exemption mirror updated to the same
  effective-endpoint precedence.

Fixes #7874
2026-08-07 03:54:45 +00:00