Commit Graph

1551 Commits

Author SHA1 Message Date
can1357 2f54e760c1 feat(coding-agent): added async compaction and in-place handoff
- Added compaction.asyncEnabled (Async Compaction, default on): when
  context enters the pre-threshold band [threshold - lead, threshold)
  with lead = clamp(threshold * 0.125, 8192, 32000), maintenance
  speculatively summarizes in the background off a branch snapshot
  (first configured LLM-backed method: remote, handoff, or soft) using
  a side session id isolated from the live turn. Crossing the threshold
  splices the armed result in instantly instead of blocking on a
  summarization round-trip. Armed results are invalidated by branch
  changes, reset boundaries, model switches that strand provider-native
  replay payloads, and context growth past keepRecentTokens (which
  re-speculates); extensions registering session_before_compact keep
  exact blocking semantics (speculation disabled).
- Reworked handoff to commit in place: /handoff and the auto handoff
  method now write the generated document as a regular compaction entry
  on the current session (summary = document + <files> tag, cut from
  prepareCompaction) instead of starting a new session. SessionHandoff
  shrank to a document generator; session_before_switch/session_switch
  no longer fire with reason "handoff"; mid-turn maintenance no longer
  suppresses the handoff preference; overflow recovery can apply an
  armed handoff result.
- Extracted the shared auto-compaction commit tail
  (#commitAutoCompactionResult / #commitCompactionEntry) used by the
  blocking production path, the armed speculative apply, manual
  compaction, and manual handoff.
- Status line pulses the auto-compact icon while a speculation runs and
  holds it in accent once a result is armed.
- Exported remotePreserveReusable from pi-agent-core/compaction for
  apply-time validation of speculative remote results.
2026-08-20 03:46:34 +02:00
can1357 eced7ab08a feat: implemented live status board utility and agent progress tracking
- Implement a live status board utility for transient multi-line CLI status displays with TTY fallback support.
- Add progress callback support and forward subagent progress events to runner hooks.
- Update cleanse execution flow to track checker runs, agent progress, and status rendering.
- Add comprehensive unit tests for live board repainting and cleanse progress assertions.
2026-08-20 03:22:54 +02:00
can1357 0cdd37fc15 feat(pi-natives/tools): implemented utok tokenizer for multiple models
- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
2026-08-20 01:45:04 +02:00
can1357 12238f55ca feat: implemented native ctok tokenization engine with model scopes
- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
2026-08-19 23:27:29 +02:00
can1357 22b40b47e1 chore: bump version to 17.3.8 2026-08-19 12:35:37 +02:00
roboomp 4b07f409f6 fix(ai): hoist assistant message interleaved in responses tool batch
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.

buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.

Fixes #8789
2026-08-19 08:46:28 +00:00
can1357 3566bd9b41 fix(ai): unified Cursor interaction-query handling after merging #8889 and #8830
- Kept the shared cursor/interaction-query module as the single handler and deleted the duplicate local implementation in cursor.ts
- Added the named webFetchRequestQuery approval case (field 9 is named under the regenerated proto)
- Preserved the deliberate no-fake-VM-success semantics for setupVmEnvironmentArgs (review of #8047)
- Updated the field-9 regression test to assert the named decode of the raw same-field reply, which also pins the LEN-prefix wire framing
2026-08-19 01:47:20 +02:00
can1357 20bd4ab97b chore(changelog): normalized [Unreleased] sections after merges and added missing entries
- Repaired union-merge artifacts in packages/coding-agent/CHANGELOG.md (duplicated 17.3.6/17.3.7 blocks; promoted the new entries back to [Unreleased])
- Added missing [Unreleased] entries for PRs #8833, #8866, #8879, #8903, #8905, #8915, #8916, #8917, #8920, #8923, #8928, #8929, #8937
2026-08-19 01:42:50 +02:00
can1357 75179e1dc0 fix(compaction): scale summary window floor to the model's context
The absolute 16,384-token floor plus the carried summary and output
reserves exceeds windows below ~58k outright, and overflow recovery then
bailed at the very floor that caused the rejection, leaving compaction
unusable on small-context models. Scale the floor to window/8 (min 1k)
and use the same floor in overflow recovery.
2026-08-19 01:39:17 +02:00
can1357 17e47b3eb7 Merge PR #8920: fix(compaction): bound summarization input and stop retrying overflow (@PaleRoses)
# Conflicts:
#	packages/agent/src/compaction/compaction.ts
2026-08-19 01:39:07 +02:00
can1357 79de0acd44 docs(agent): add changelog entry for compaction prompt-injection hardening (#8727) 2026-08-19 01:36:04 +02:00
can1357 996f562247 Merge PR #8727: fix(agent): harden compaction summaries against prompt injection (@koopmannleon19977-cmyk) 2026-08-19 01:36:04 +02:00
can1357 e88fb70afe Merge PR #8720: fix(compaction): honor /clear reset boundary in prepareCompaction (@roboomp) 2026-08-19 01:36:03 +02:00
PaleRoses 753c86ea72 fix(compaction): bound summarization input and stop retrying overflow
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.

Three independent defects:

1. `generateSummary` serialized the entire span into one prompt with no budget
   check. It now plans windows that fit the summarizer's context and folds them
   with the update prompt that iterative compaction already uses, so a stranded
   boundary is recovered instead of rejected. A provider that rejects a window
   the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
   beta-gated to 200k on OAuth credentials) halves what was actually sent and
   re-plans, because only the rejection knows the real cap.

2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
   the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
   appends to its own errors classified a deterministic 400 as a transient 503.
   Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.

3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
   identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
   an input that does not fit never fits; both layers now fail fast to the next
   candidate.

The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.

Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
2026-08-18 12:25:13 -07:00
can1357 8500092296 chore: bump version to 17.3.7
Retry: widened agent dequeue-hook deadline budgets from 25ms to 1s — the run loop checks the deadline before invoking dequeue hooks, so a cold or CPU-starved mock roundtrip expired the deadline first and the hooks never ran (deterministic failure in isolation, flaky under CI parallel load).
2026-08-18 11:34:10 +03:00
can1357 0a912cc467 chore: bump version to 17.3.7 2026-08-17 22:29:25 +03:00
can1357 54e1a8c900 chore: bump version to 17.3.6 2026-08-17 17:16:40 +03:00
koopmannleon19977-cmyk 39c908bc90 fix(agent): harden compaction summarizer against prompt injection 2026-08-16 16:22:51 +02:00
roboomp 722c4aa0ae fix(compaction): honor /clear reset boundary in prepareCompaction
prepareCompaction walked the branch from the last compaction and ignored reset_boundary markers, so /compact (and auto-compaction) resurrected pre-/clear turns into the summary even though buildSessionContext already starts the model context after the boundary.

Model reset_boundary as a first-class agent-core session entry and start the summarization window after the latest boundary, dropping the superseded pre-reset compaction summary. A boundary before the last compaction stays superseded by it.

Fixes #8718
2026-08-16 11:37:59 +00:00
can1357 37eee71978 chore: bump version to 17.3.5 2026-08-16 10:21:05 +03:00
Can Bölük ca1f184823 chore: rewritten changelogs 2026-08-16 09:28:34 +03:00
can1357 544c414296 Merge PR #8561: fix(ai): preserve Anthropic tool-search replay blocks (@roboomp) 2026-08-16 02:43:33 +02:00
can1357 3e64a24714 test(agent): typed abort-signal listener without DOM lib 2026-08-16 02:19:32 +02:00
can1357 55e5da3d17 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:15:11 +02:00
can1357 df6d1e1ac5 fix(ai): completed oneshot transient retry handling 2026-08-16 02:14:12 +02:00
can1357 b445134b4d Merge PR #8370: fix: retry transient Anthropic failures at oneshot LLM call sites (@wonjun3991)
# Conflicts:
#	packages/coding-agent/src/utils/title-generator.ts
2026-08-16 02:14:12 +02:00
roboomp bee1d44bf9 fix(ai): preserved anthropic tool-search replay blocks
- Retained tool-search server calls and opaque results in signed assistant history across direct streams, gateways, and custom-endpoint projection.

- Added replay regressions for interleaved thinking and client tool continuations.

Fixes #8559
2026-08-14 14:34:27 +00:00
can1357 ffd53ff92a chore: bump version to 17.3.4 2026-08-14 14:38:16 +02:00
roboomp 2996f16a61 fix(agent): added codex v2 compaction feature header
- Negotiated remote_compaction_v2 on Codex compatibility fetches.

- Covered explicit endpoints and non-Codex request scoping.

Fixes #8524
2026-08-14 06:56:57 +00:00
can1357 039728ad80 chore: bump version to 17.3.3 2026-08-14 05:44:05 +02:00
can1357 ae2d3d6ea1 chore: bump version to 17.3.2 2026-08-14 00:28:43 +02:00
can1357 0bc2c342f4 chore: bump version to 17.3.1 2026-08-13 19:39:21 +02:00
can1357 b279db1790 test: refactored test suites to eliminate time-based sleeps and polling loops
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
2026-08-13 19:32:22 +02:00
can1357 8b0f400d3c chore: bump version to 17.3.0 2026-08-13 08:28:43 +02:00
can1357 6b4823181b test: cleaned test suites and documented filtering guidelines
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
2026-08-13 08:28:42 +02:00
can1357 fc1fd664b9 chore(changelog): rewritten 2026-08-13 05:51:01 +02:00
wonjun3991 720ac2168a fix(agent): retry handoff, branch summary and manual /compact on a blip
Each is a single, side-effect-free completion whose result is parsed after
it resolves, so one transient provider failure previously aborted the
whole operation - for /compact that left the user's context full.

Adds SummaryOptions.oneshotRetry, because both compaction paths call the
same generateSummary and the policy therefore cannot be a constant inside
it. Manual /compact has no outer loop and gets retry by default;
auto-compaction passes false because session-maintenance already retries
the whole attempt, and nesting would multiply the budget (10 outer x 3
inner) while stacking each outer wait on an inner backoff.
2026-08-13 10:54:42 +09:00
wonjun3991 0523c7112f feat(agent): add opt-in transient retry to instrumentedCompleteSimple
instrumentedCompleteSimple is the single funnel for every oneshot LLM
call in the agent, so retry belongs here rather than in a try/catch above
each caller - the failure arrives as a resolved AssistantMessage.

Opt-in rather than default-on: oneshotKind is free-form and callers may
pass arbitrary ctx.tools, so the funnel cannot itself prove a request is
replay-safe.

Response headers are captured per attempt and cleared between attempts,
so a stale retry-after can never be reused for a later failure.
2026-08-13 10:54:42 +09:00
can1357 4d73392621 chore: fix changelog 2026-08-13 02:02:23 +02:00
can1357 c25b008f4a Merge PR #8067: fix(compaction): manual /shake keeps a recent tail of tool results (@zhang17-24) 2026-08-13 02:00:49 +02:00
can1357 5481d8b9b0 chore: bump version to 17.2.15 2026-08-12 03:26:12 +02:00
can1357 c101452bb5 refactor: restructured and condensed agent prompts and system instructions
- Refactored and condensed numerous system prompts, agent instructions, and tool documentation files across packages.
- Streamlined workflow rules, formatting constraints, and execution guidelines for improved clarity and brevity.
- Updated discovery rules, recommendation criteria, and syntax standards in prompt templates.
2026-08-12 01:48:14 +02:00
can1357 e5ebb2aee0 chore: bump version to 17.2.14 2026-08-11 20:43:02 +02:00
left-to-right 835a7db15a fix(compaction): keep rescue shake able to elide the newest result
RESCUE_SHAKE_CONFIG spreads AGGRESSIVE_SHAKE_CONFIG, so the new 4k manual
tail leaked into dead-end recovery and could block eliding the very
result that caused the dead end. Override protectTokens back to 0 in the
rescue preset and add the coding-agent changelog entry for the manual
/shake behavior change.

Addresses review on #8067.
2026-08-11 22:33:40 +08:00
can1357 2157becbe9 chore: bump version to 17.2.13 2026-08-11 16:03:05 +02:00
can1357 64baa7c1bd chore(format): applied biome formatting and removed dead code from merged prs 2026-08-11 15:14:15 +02:00
can1357 270fe8d454 Merge PR #8183: fix(cursor): omit undefined exec args and keep exec-resolved under owned dialects (@jairuspace) 2026-08-11 15:06:16 +02:00
can1357 00d90809fc test: removed redundant policy-key suite 2026-08-11 15:06:11 +02:00
can1357 8504e4865f Merge PR #7945: fix(coding-agent): resolve xd:// device dispatches against device user policy first (@re2zero) 2026-08-11 15:06:11 +02:00
jairuspace 5cbd482f8e fix(cursor): omit undefined exec args and keep exec-resolved under owned dialects
Cursor bash/grep frames wrote optional kwargs as present-undefined, and
tools.format gemini projectors dropped kCursorExecResolved so settled
calls ran twice.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-10 13:39:40 -06:00