Commit Graph

3322 Commits

Author SHA1 Message Date
usr_bin_roygbiv 21cb04bd3e fix(ai): retry safe truncated responses streams 2026-07-17 16:48:02 -05:00
can1357 2c225a0d00 chore: bump version to 17.0.3 2026-07-17 21:38:00 +02:00
can1357 0e05691cdb style: formatted kimi-code k3 reasoning test 2026-07-17 21:24:10 +02:00
can1357 69148a89ca chore: normalized changelogs after farm PR sweep 2026-07-17 21:23:27 +02:00
can1357 8577371187 Merge PR #5896: fix(kimi-code): preserve K3 native effort contract (@roboomp) 2026-07-17 21:22:08 +02:00
can1357 cd97af0e19 Merge PR #5853: fix(ai): order parallel tool outputs before result images (@roboomp) 2026-07-17 21:22:08 +02:00
can1357 0422357bff Merge PR #5880: fix(extensions): restore legacy Kimi provider compatibility (@roboomp) 2026-07-17 21:22:05 +02:00
can1357 d6666684da Merge PR #5829: fix(cursor): surface actionable error when h2 ALPN negotiation fails (@roboomp) 2026-07-17 21:22:04 +02:00
can1357 3f3e731d2f Merge PR #5815: fix(ai): preserve Anthropic account identity in usage reports (@roboomp) 2026-07-17 21:22:03 +02:00
can1357 242b3866e8 fix(kimi-code): enforce mandatory K3 reasoning on the Kimi dispatch
The Kimi branch in streamSimple forwarded raw options to streamKimi,
bypassing normalizeMandatoryReasoningOptions. With supports_thinking_type
'only' now surfaced as thinking.requiresEffort, disabled/omitted requests
(e.g. title generation) serialized thinking:{type:disabled}, which the
mandatory K3 endpoint rejects. Clamp to the lowest supported effort in the
Kimi path, mirroring the mapOptionsForApi contract every other provider
uses, and add a regression test.
2026-07-17 21:09:40 +02:00
roboomp 913ec0baae fix(kimi-code): preserved k3 native effort contract
- Parsed live named efforts, mandatory-thinking state, and model protocol metadata.

- Sent native Kimi named efforts and adaptive Anthropic override efforts without generic token budgets.

Fixes #5893
2026-07-17 18:19:00 +00:00
roboomp 7e8c4fc01f fix(extensions): restored legacy provider compatibility
Restored the historical assistant-message stream factory and synchronous auth-storage facade needed by pi-provider-kimi-code.

Fixes #5879
2026-07-17 16:54:47 +00:00
roboomp c77518851f fix(ai): ordered parallel tool outputs before images
Tracked synthetic tool-image messages by identity and inserted consecutive tool outputs before their trailing image block.

Covered generic and Codex Responses serialization while preserving the following genuine user-message boundary.

Fixes #5850
2026-07-17 14:37:13 +00:00
roboomp 148a48a215 fix(cursor): surfaced actionable error when h2 ALPN negotiation fails
The Cursor run RPC is HTTP/2-only (api2 rejects HTTP/1.1 with 464), and bun
only opens an HTTP/2 session when TLS-ALPN negotiates h2. Behind an
ALPN-stripping TLS-intercepting proxy (e.g. Zscaler) the handshake yields no
h2 and bun throws ERR_HTTP2_ERROR: "h2 is not supported", which streamCursor
passed through verbatim. Model discovery masks the same failure by falling
back to bundled models.json, so listing works while runs fail opaquely.

Map the h2-negotiation failure on both the session and request error paths to
a ProviderResponseError that names the ALPN-stripping proxy as the cause and
points at the providers.cursor.baseUrl HTTP/2 bridge workaround.

Fixes #5828
2026-07-17 10:27:57 +00:00
roboomp 5a54a70e39 fix(ai): separated anthropic account and organization identity
- Kept the stored OAuth account identifier authoritative when Anthropic only returns an organization header.

- Prevented org-only usage metadata from failing active-account matching in the status line.

Fixes #5698
2026-07-17 07:55:41 +00:00
can1357 d527259c26 chore: bump version to 17.0.2 2026-07-17 07:40:54 +02:00
can1357 0f26558d03 chore: reformat 2026-07-17 07:40:21 +02:00
can1357 eac51b6a04 Merge: darkphilosophy/feat/advisor-per-agent-toggle
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:

- #failing latch: waitForCatchup resolves immediately while an advisor is
  mid-failure; parked waiters wake the moment a turn fails, before any
  async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
  dedup snapshot and never propagates into the primary's turn-end callback
  (per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
  runtime latches quotaExhausted, requeues the batch, and notifies —
  cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
  before any assistant turn is recorded.
2026-07-17 07:37:29 +02:00
can1357 38b73a5cbc style(tests): removed unused fixture variable and formatted plan test
The redaction-contract rewrite left credentialTokens unused in the
GitLab Duo provider test, failing biome's lint gate.
2026-07-17 07:24:54 +02:00
can1357 4b4bad430a test(ai): aligned kimi-code thinking tests with the zai compat policy
The regenerated catalog stamps kimi-for-coding with the zai thinking
format, under which reasoning yields to a forced tool choice (#5758
review) instead of downgrading the choice: chat-completions carries an
explicit thinking {type: disabled} and the Anthropic wire keeps the
forced choice with no thinking block.
2026-07-17 06:38:27 +02:00
can1357 731a2cb5b2 test(natives): gave timeout drain repro a spawn-proof deadline
The 50ms budget raced external-process spawn on cold CI runners: cancel
could fire before yes produced output, so the builtin tail flushed an
empty ring buffer (0 lines instead of 5, Linux x64 modern). 750ms keeps
the post-cancel drain scenario while outlasting spawn latency.
2026-07-17 06:09:51 +02:00
can1357 5f46138c4d fix(ai): redacted the GitLab Duo goal transcript
The Duo goal bypasses transformMessages; apply the outbound credential
scrub (#5655) to the rendered ChatML transcript and latest-prompt goal,
and updated the provider test to the redaction contract.
2026-07-17 05:40:13 +02:00
can1357 32fecd0c8d fix(tests): repaired type errors from owner-approved merges
Dropped a nonexistent AnthropicCompat field from the redaction test and
made the handoff agent_end capture handler return void.
2026-07-17 05:36:09 +02:00
can1357 6f0d61f7ca fix(ai): dropped stale signatures after credential redaction 2026-07-17 05:26:18 +02:00
usr_bin_roygbiv fab31a2bae fix(ai): redact sensitive credentials in system prompt instructions 2026-07-17 05:22:35 +02:00
usr_bin_roygbiv c55196eb30 fix(ai): retry full transcript on blocked stateful OpenAI Responses prompts 2026-07-17 05:22:35 +02:00
usr_bin_roygbiv 99ecfd6b6c fix(ai): avoid JSON.stringify BigInt serialization crash in redactSensitiveInObject 2026-07-17 05:22:35 +02:00
usr_bin_roygbiv b59ee9a99c fix(ai): redact sensitive credentials from outbound messages 2026-07-17 05:22:34 +02:00
can1357 f36394ef55 style: fixed formatting in merged test and session files 2026-07-17 05:02:31 +02:00
can1357 f4bbf2dec6 chore: normalized changelogs after merge sweep
Ran scripts/fix-changelogs.ts --since 1424cae06 to merge duplicate
Unreleased headings and promote entries the union merge driver misfiled.
2026-07-17 05:01:09 +02:00
DarkPhilosophy 1b4c292f8f Merge remote-tracking branch 'can1357/main' into feat/advisor-per-agent-toggle 2026-07-17 05:47:09 +03:00
can1357 7b21704a72 merge PR #5772 via eval/pr-5772: fix(codex): force tool_choice auto on responses-lite requests 2026-07-17 04:42:11 +02:00
can1357 34cd2df096 merge PR #5758 via eval/pr-5758: fix(ai): honored kimi-k3 131k output limit 2026-07-17 04:40:05 +02:00
can1357 f59dea091c merge PR #5724 via eval/pr-5724: fix(ai): honored custom kimi anthropic base urls 2026-07-17 04:40:03 +02:00
can1357 93c6f219a5 merge PR #5723 via eval/pr-5723: fix(ai): classify 402/balance-exhausted quota errors as usage limits 2026-07-17 04:40:03 +02:00
can1357 b587da5059 merge PR #5673 via eval/pr-5673: fix(ai): recognize indented markdown fences as fenced code 2026-07-17 04:37:11 +02:00
can1357 bcc7d34a9a merge PR #5651 via eval/pr-5651: fix(cursor): gated mounted device execution through approval 2026-07-17 04:37:09 +02:00
can1357 731448c53c merge PR #5636 via eval/pr-5636: fix(ai): deferred Cursor completion until protocol end 2026-07-17 04:37:09 +02:00
can1357 2d41785b3a fix(cursor): prevent duplicate MCP execution 2026-07-17 04:18:20 +02:00
can1357 b0d04e5173 fix(ai): fixed snapshot validation for login-sourced API keys
- Added optional `source?` field with value `'login'` to `apiKeyCredentialSchema` so snapshots accept login-sourced API keys.
- Updated CHANGELOG with a fixed entry describing the correction.
- Added a test verifying that a snapshot containing `source: "login"` passes client wire validation.
2026-07-17 04:15:58 +02:00
can1357 9e926c6cb5 docs(changelog): describe Lite choice handling 2026-07-17 03:58:01 +02:00
can1357 1f0d8d9750 fix(codex): preserve Lite tool-use constraints 2026-07-17 03:57:41 +02:00
can1357 e99fb80741 fix(ai): document and cover 402 usage limits 2026-07-17 03:57:02 +02:00
roboomp f961d824bf fix(codex): force tool_choice auto on responses-lite requests
The Responses-Lite rewrite moves tools into an `additional_tools` developer
input and deletes top-level `tools`, but preserved a forced top-level
`tool_choice` (e.g. `{ type: "web_search" }`). With no top-level tools to
validate against, the ChatGPT Codex endpoint rejected the request with
`HTTP 400 Tool choice '…' not found in 'tools' parameter`, and web search
silently fell back to Gemini.

`applyCodexResponsesLiteShape` now sets `tool_choice: "auto"`, matching
codex-rs `build_responses_request`. Classic (non-Lite) Responses requests
keep their forced choice since top-level `tools` remains present.

Fixes #5771
2026-07-16 23:42:30 +00:00
roboomp aa4386d8ca fix(ai): honored kimi-k3 131k output limit
The generic OpenAI-compatible output policy capped native K3 requests at 64,000 tokens even though its catalog metadata advertises Moonshot’s 131,072-token output limit.

- Generalized the Chat Completions provider clamp resolver.
- Allowed native moonshot/kimi-k3 to clamp against model.maxTokens.
- Preserved the existing 64k default for other OpenAI-compatible models and the existing raised GLM-5.2 reasoning clamp.
- Covered default and explicit 131,072-token K3 requests on the wire.

Fixes #5756
2026-07-16 21:50:57 +00:00
roboomp 3eda7154f9 fix(ai): honored custom kimi anthropic base urls
Derived the Anthropic Messages root from the configured OpenAI-compatible model base URL and covered custom gateway routing with a regression test.

Fixes #5722
2026-07-16 15:46:11 +00:00
metaphorics c0ec090a9c fix(ai): classify 402/balance-exhausted quota errors as usage limits
Grok Build reports exhausted account balances as HTTP 402, which bypassed credential rotation. Treat that status and wording as persistent account-local quota exhaustion while preserving informative non-quota bodies in the backoff lane.
2026-07-17 00:38:55 +09:00
roboomp e014577138 fix(ai): recognize indented markdown fences as fenced code
The fence-vs-span decision only treated a backtick run as a fenced block
when it sat immediately after a newline. CommonMark allows a fence indented
by up to three spaces, so an indented ```md block was read as an inline
span, closed at an inner triple-backtick string literal, and healed a later
literal reasoning tag as thinking.

Track the current line's leading-space count (reset on newline, invalidated
by the first non-space char) and open a fenced block when the run is >= 3
backticks at an indent of 0-3 spaces, matching the sibling FencedThinking
scanner's fence-line rule.

Fixes #5665
2026-07-16 08:14:38 +00:00
roboomp 398ab3f7b6 fix(ai): close fenced code blocks only on a fence line
Code mode treated inline spans and fenced blocks identically, closing at
the first matching backtick run anywhere in the buffer. An inline triple-
backtick literal inside a fenced block (e.g. `const fence = '```';`) exited
code mode early, so a later literal reasoning tag in the same block was
healed as thinking and dropped from the rendered code.

Track whether the opener was a fence (a backtick run >= 3 at line start) or
an inline span. A fenced block now closes only on a fence line — a line of
backticks at least as long as the opener — streaming committed lines while
holding the last partial line; inline spans keep closing on the matching
backtick run.

Fixes #5665
2026-07-16 08:03:52 +00:00
roboomp 45c7d7f173 fix(ai): keep literal think tags inside markdown code visible
ThinkingInbandScanner scanned the visible-text channel for leaked reasoning
open tags with a plain indexOf, ignoring Markdown code spans. A literal
`<think>` inside inline code or a fenced block was read as a reasoning
boundary, so the unmatched tag split the text into text + thinking and
corrupted the rendered Markdown.

The scanner now tracks code-span state: a backtick run enters code mode and
suppresses reasoning-tag detection until the matching closing run, streaming
the content through as verbatim text. Reasoning tags still win at any position
so the gemini ```thinking fence keeps healing.

Fixes #5665
2026-07-16 07:26:35 +00:00