Commit Graph

896 Commits

Author SHA1 Message Date
can1357 4e6f04220f chore: reformat 2026-05-31 13:04:25 +02:00
can1357 6cee7a666d test(tests): added deterministic yield tests and viewport mutation regression coverage
- Replaced real-time yield assertions with a mocked clock and scheduler wait spy to validate gate timing deterministically.
- Updated fact consolidation conflict tests to capture console warnings and verify duplicate-resolution messaging.
- Added TUI offscreen-expansion regressions for unknown-viewports, covering deferred rebuild and user-driven rebuild behavior.
2026-05-31 13:02:32 +02:00
can1357 a043826723 chore: bump version to 15.7.3 2026-05-31 12:30:20 +02:00
can1357 417a1a1d32 feat(agent): added shake compaction strategy primitives
Introduce the shake compaction strategy in the core agent: exports, summary prompt, implementation, and unit coverage.
2026-05-31 07:39:38 +02:00
can1357 ff9a6826dd fix(compaction): made read tool-results prunable except for skill:// paths
- Replaced flat `protectedTools: string[]` with `ProtectedToolMatcher[]` supporting predicate functions.
- Regular file/URL `read` calls are now eligible for pruning and shake compaction.
- `read` calls whose `path` starts with `skill://` remain protected like native `skill` results.
- Added `collectToolCallsById` to correlate tool results with their originating call arguments.
2026-05-31 07:12:27 +02:00
Can Bölük 650c4e0b61 Merge pull request #1507 from oldschoola/audit/fix-top5-plus-honorable-mentions
audit: fix top 5 perf/correctness findings + honorable mentions
2026-05-31 06:46:44 +03:00
can1357 570b439b31 chore: enabled parallel Bun test runs across package scripts
- Updated each package's test script to run `bun test` with the `--parallel` flag.
- Updated the tui package test command to preserve its `test/*.test.ts` file filter while enabling parallel execution.
2026-05-31 05:24:51 +02:00
can1357 5db3bcabad fix(agent): snapshot initial mutable state
Addresses review feedback on #1507.
2026-05-31 04:49:17 +02:00
oldschoola 21d152c70c feat(coding-agent): async ConfigFile + ModelRegistry.create() factory; fix MemorySessionStorage O(N^2) appends
F6: defer the JSON -> YAML migration out of ConfigFile's constructor and
add an async path so the boot sequence stops blocking the event loop on
sync I/O. New ConfigFile.tryLoadAsync/loadAsync/loadOrDefaultAsync/
getMtimeMsAsync and static ConfigFile.warmup. The migration is now
idempotent (per-process cache) so relocate() does not re-run it.
ModelRegistry.create(authStorage, modelsPath?) is a new async factory
that runs the warmup before the sync constructor's bundled-model load.
Production call sites (main.ts, sdk.ts, task/executor.ts, commit
pipelines, SDK example) all switched. Sync new ModelRegistry(...)
constructor is still supported for tests.

F2: rewrite MemorySessionStorage's mirror as { chunks: string[]; byteLen;
mtimeMs } so writeLineSync appends a single chunk in O(1) instead of
read-modify-writing the entire file (which was O(N) per append, O(N^2)
per session). statSync now reports true UTF-8 byte length instead of
character count. readTextPrefix walks chunks until the byte budget is
exhausted instead of materialising the full mirror.

Also rolls in per-package CHANGELOG entries for F1-F8.
2026-05-31 04:46:06 +02:00
oldschoola 3a733c480b perf(agent,ai,coding-agent): in-place state mutation + per-delta json-parse throttle + drop hot-path structuredClone
F3 (agent): mutate state.messages and state.pendingToolCalls in place on
appendMessage/popMessage/clearMessages/reset/tool_execution_start/_end
instead of allocating a fresh array/Set on every transition. Subscribers
that capture state.messages by reference now observe updates directly.
Public type signature unchanged.

F5 (ai): add parseStreamingJsonThrottled to utils/json-parse — a per-delta
wrapper around parseStreamingJson that skips the re-parse until the
buffer has grown by minGrowthBytes (default 256). Wired into every
provider's tool-call argument accumulator (anthropic, amazon-bedrock,
openai-completions, openai-codex-responses, openai-responses-shared) so
per-delta cost becomes O(N) in total buffer length instead of O(N²).
Every provider's toolcall_end still runs a final unthrottled parse, so
the published block.arguments is unchanged.

F8 (coding-agent): drop the per-delta structuredClone of streaming tool
arguments in ToolExecutionComponent.updateArgs. event-controller.ts and
ui-helpers.ts already spread their input into a fresh object on each
delta, so cloning here was dead work on the rendering hot path. Added a
reference-equality short-circuit so repeat calls with the same args
object skip the preview-diff and display refresh.
2026-05-31 04:46:06 +02:00
can1357 f9f213e53e chore: bump version to 15.7.2 2026-05-31 03:46:17 +02:00
can1357 73de3bf926 chore: bump version to 15.7.1 2026-05-31 02:50:54 +02:00
can1357 e72fe367e5 chore: bump version to 15.7.0 2026-05-31 02:17:33 +02:00
can1357 474eb92047 chore: bump version to 15.6.0 2026-05-30 19:01:54 +02:00
can1357 e831c2c758 chore: reformat 2026-05-30 18:08:51 +02:00
can1357 5f3856ee7c chore: bump version to 15.5.15 2026-05-30 04:49:01 +02:00
can1357 ae905fb3cf feat(agent): added agent tool-call cap enforcement to stream loop
- Added `maxToolCallsPerTurn` support to `AgentOptions` and `AgentLoopConfig`, with Agent getter/setter and serialized state wiring.
- Implemented stream-loop cap handling by normalizing bad values and halting after `toolcall_end` reaches the limit.
- Added `ANTHROPIC_TOOL_CALL_BATCH_CAP`=8 and wired session cap sync on init, model changes, and restore.
- Added tests that truncated a 10-call stream to 8 tool calls, and verified non-Claude models resolve no cap.
2026-05-30 04:47:58 +02:00
can1357 e004585f39 chore: bump version to 15.5.14 2026-05-30 01:25:28 +02:00
can1357 445fe99fbc fix(agent): executed tool calls emitted under end_turn stop reason
- Fixed agent loop abandoning tool_use blocks on `stop`/`end_turn` turns; only `length` (truncation) now skips execution.
- Verified against live Anthropic API: stop_reason is never replayed on wire and doesn't gate continuation validity.
- Added tests pinning the wire-safety contract: thinking blocks with stale/missing signatures are downgraded to text before replay.
2026-05-30 00:27:03 +02:00
can1357 1422afd54d chore: bump version to 15.5.13 2026-05-29 18:21:43 +02:00
can1357 14c366b049 chore: bump version to 15.5.12 2026-05-29 14:06:38 +02:00
can1357 1fe843de2e fix(agent): handled skipped tool calls for non-toolUse assistant turns
- Updated the agent loop to execute tool calls only when the assistant stop reason was `toolUse`.
- Added skipped placeholder `tool_result` messages for leftover `toolCall` blocks when a turn ended without `toolUse`.
- Promoted OpenAI/Ollama `stop` tool-call turns to `toolUse` and stripped thinking signatures on abandoned tool-use turns during message transforms.
2026-05-29 13:31:37 +02:00
can1357 e707b90652 chore: bump version to 15.5.11 2026-05-29 07:19:54 +02:00
can1357 3f073b82de chore: bump version to 15.5.10 2026-05-28 14:50:02 +02:00
can1357 c4f93eca20 fix(agent): patched compaction 401/403 fallback to copy error status
- Centralized compaction stop-reason error throws via createSummarizationError().
- Set compaction thrown errors to copy response.errorStatus into Error.status.
- Expanded compaction auth detection to treat HTTP 401/403 as auth failures with regex fallback preserved.
- Added regression tests for 401/403 status propagation and compaction fallback auth behavior.
- Documented both package fixes in Unreleased Fixed changelog entries.
2026-05-28 12:50:52 +02:00
can1357 eca3a08d25 chore: bump version to 15.5.9 2026-05-28 12:09:01 +02:00
can1357 6cc91fcb37 chore: bump version to 15.5.8 2026-05-28 10:37:49 +02:00
can1357 c1fa0e9f50 refactor(agent): replaced keepalive utility with disposable EventLoopKeepalive
- Replaced the `keepaliveWhile` Promise wrapper with a new `EventLoopKeepalive` class that registers and disposes an interval timer through `Symbol.dispose`.
- Updated `Agent` to instantiate `EventLoopKeepalive` via `using` during prompt execution instead of manually managing an interval.
- Wrapped interactive mode's await path with the new helper and removed redundant `keepaliveWhile` usage from the CLI entrypoint.
2026-05-28 09:51:57 +02:00
hezhiyang2000 f71e1db0c2 fix: restore #emit listener isolation (isPromise + try/catch)
The previous commit accidentally dropped the try/catch in #emit
because the local agent.ts was based on an older main that lacked
the listener-isolation code. Restore it to match main.
2026-05-28 14:09:01 +08:00
hezhiyang2000 af2011f5a1 fix: use inline setInterval instead of EventLoopKeepalive import
EventLoopKeepalive is not exported from yield.ts on main.
Use setInterval + unref() directly, matching the pattern in
keepaliveWhile(). Biome import order also fixed (type imports
before value imports within the same group).
2026-05-28 11:51:24 +08:00
hezhiyang2000 38d77f819d style: fix Biome import order (type imports before value imports)
Biome organizeImports sorts type imports before value imports within
the same relative-path group. Move EventLoopKeepalive import after
all type imports.
2026-05-28 11:50:15 +08:00
hezhiyang2000 6fb1983fbc fix: add EventLoopKeepalive to Agent.prompt()
Bun 1.3.x event loop busy-waits when the only pending work is an
unresolved Promise. Agent.prompt() sets #runningPromise via
Promise.withResolvers() which stays unresolved during the entire
agent loop execution (LLM calls + tool iterations), causing ~100%
CPU even when the process is idle.

PR #1419 added keepaliveWhile() to getUserInput() in main.ts, but
session.prompt() callers (interactive mode, resume, etc.) still
await the unresolved #runningPromise, bypassing the keepalive.

Install EventLoopKeepalive directly in Agent.prompt() so all
callers are covered. Dispose in the finally block after the agent
loop completes.
2026-05-28 10:07:22 +08:00
can1357 b87cd8f93f chore: bump version to 15.5.7 2026-05-27 19:01:33 +02:00
cognitive 0455c164f9 fix(agent): compact() propagates thinkingLevel to all three summarizers
The thinking-level fix (e07b47ee4) added SummaryOptions.thinkingLevel
and threaded it from agent-session.ts into compact(), but the
field-by-field rebuild of summaryOptions inside compact() (and a
second inline rebuild for generateShortSummary) silently dropped it.

Effect on every call site that fans through compact():
  generateSummary, generateTurnPrefixSummary, generateShortSummary all
  see options?.thinkingLevel === undefined => resolveCompactionEffort
  falls back to Effort.High => user's /model :off selection is
  silently overridden, and xai-oauth/grok-build still trips on the
  unsupported-effort path even though fix #2 strips it at the wire
  layer of the openai-responses mapper.

Add `thinkingLevel: options?.thinkingLevel` at both rebuild sites in
compact(): the summaryOptions literal feeding generateSummary and
generateTurnPrefixSummary, and the inline options literal feeding
generateShortSummary.

Extend compaction-thinking-level.test.ts with four compact()-level
cases driving isSplitTurn:true so all three summarizers fire:
  - Off                  -> every fan-out call gets reasoning=undefined
  - Low                  -> every fan-out call gets reasoning="low"
  - <unset>              -> every fan-out call gets reasoning="high" (default)
  - grok-build + High    -> every fan-out call gets reasoning=undefined (clamp)

TDD red-green verified: stashing the source fix flips the Off and Low
cases to fail with received="high" (exactly the reviewer's prediction);
restoring the fix returns all four to green. Suite: 131 pass / 0 fail
(baseline 127 + 4 new). biome + tsgo --noEmit clean.

Op: correct
Restores: spec:compaction-honors-session-thinking-level
(cherry picked from commit 9b501e369b820cb992345ceb3f2a207ccc9ee233)
2026-05-27 15:01:24 +00:00
cognitive 25c6794cd5 fix(agent): compaction honors session thinking level and silent-clamps unsupported-effort models
Triple-stacked failure on the same axis (thinking effort) produced the
user-visible

    Error: Compaction failed: Thinking effort high is not supported by
           xai-oauth/grok-build.
    Supported efforts:

(empty list after the colon) whenever the active model was a curated
xAI catalog entry with compat.supportsReasoningEffort: false.

Three defects lined up. (1) Behavior: compaction at four call sites
in packages/agent/src/compaction/compaction.ts hardcoded
reasoning: Effort.High and never threaded session.thinkingLevel —
the user's /model :off selection (and any explicit low/medium) was
silently overridden. On every other model this was invisible.
(2) Validation: requireSupportedEffort threw at the openai-flavored
mapper layer before the wire-side omitReasoningEffort gate in
providers/xai-responses.ts ever ran; two contradictory guards on the
same wire param. (3) Message: when getSupportedEfforts returned [],
the rendered error tail was 'Supported efforts: ' with nothing after
the colon — disappears as a side-effect of fix #2.

Fix #1 — thread ThinkingLevel | undefined end-to-end. Add
SummaryOptions.thinkingLevel and HandoffOptions.thinkingLevel.
Convert via a single exhaustive switch (effortFromThinkingLevel) in
the new resolveCompactionEffort helper:
  - Off            → undefined  (omit reasoning entirely)
  - undefined/Inherit → Effort.High → clamp per model (preserves the
                                       historical default for users
                                       who never touched the dial)
  - explicit Effort → respect user → clamp per model

resolveCompactionEffort lives in compaction.ts; all four call sites
(generateSummary, generateHandoff, generateShortSummary,
generateTurnPrefixSummary) route through it. agent-session.ts threads
this.thinkingLevel into all three production compaction entry points
(manual /compact at L6201, auto-compaction at L6458 — the most-fired
path, originally missed in plan review — and direct generateHandoff
at L5465). The audit-gate test
(test/agent-session-compaction-thinking-threading.test.ts) scans the
file with a brace-balanced extractor and refuses any unthreaded site.

Fix #2 — silent-clamp at the openai-flavored mapper layer. Extract
exported modelOmitsReasoningEffort(model) in model-thinking.ts as the
single source of truth for compat.supportsReasoningEffort: false on
openai-responses* APIs. getSupportedEfforts now calls it instead of
inlining the check (pure refactor — observable behavior preserved).
resolveOpenAiReasoningEffort in stream.ts early-returns undefined
when the predicate is true, so the wire-side omitReasoningEffort
gate (providers/xai-responses.ts:78) becomes the single source of
truth for the actual strip — no redundant throw.

Three regression tests pin the contract:
  - packages/ai/test/xai-oauth-effort-strip.test.ts (5 tests):
    modelOmitsReasoningEffort returns true for grok-build and
    grok-4.20-0309-reasoning, false for grok-4.3 / Anthropic /
    openai-completions.
  - packages/agent/test/compaction-thinking-level.test.ts (5 tests):
    every ThinkingLevel outcome through generateHandoff — Off stays
    undefined (not coerced to High), Low stays Low, Inherit / undefined
    default to High, grok-build clamps to undefined regardless of
    requested level. Covers the Codex-caught Off-vs-not-provided
    distinction.
  - packages/coding-agent/test/agent-session-compaction-thinking-threading.test.ts
    (2 tests): brace-balanced source scan asserts every direct
    compact() / generateHandoff() in agent-session.ts threads
    'thinkingLevel: this.thinkingLevel'; floor of 3 threaded sites.

TDD red-green verified for fix #1: temporarily reverted the handoff
call-site back to hardcoded Effort.High → compaction-thinking-level
went 2 pass / 3 fail (Off coerced, Low overridden, grok-build throws);
restored → 5 pass / 0 fail.

Verified:
  - packages/agent:  127 pass / 0 fail
  - packages/ai:     1061 pass / 337 skip / 0 fail
  - packages/coding-agent (focused): 179 pass / 5 skip / 0 fail
  - biome + tsgo --noEmit clean across all three packages

Out of scope (follow-ups):
  - branch-summarization.ts:307 already passes no reasoning — no edit.
  - The empty-list error message at model-thinking.ts:296 is now
    structurally unreachable from the openai-responses path.
  - modelOmitsReasoningEffort and grokSupportsReasoningEffort
    (xai-responses.ts:22) overlap; collapse into a single predicate
    in a future commit.

Op: correct
Restores: spec:compaction-honors-session-thinking-level
Restores: spec:xai-oauth-grok-build-compaction-no-throw
(cherry picked from commit e07b47ee46769053c658819437e2478389a4cee0)
2026-05-27 15:01:20 +00:00
can1357 bf92a3dd4d chore: bump version to 15.5.6 2026-05-27 15:04:44 +02:00
can1357 ca86239bda Revert "wip: gentle"
This reverts commit 99bae2ce6c.
2026-05-27 15:01:59 +02:00
can1357 99bae2ce6c wip: gentle 2026-05-27 14:45:02 +02:00
can1357 f0c7df60d5 chore: bump version to 15.5.5 2026-05-27 14:44:30 +02:00
can1357 6643178d6a refactor: replaced instanceof Promise checks with isPromise utility
- Imported isPromise from node:util/types in four modules.
- Replaced four instanceof Promise checks with isPromise calls for more reliable promise detection.
2026-05-27 13:02:52 +02:00
can1357 0e9137a703 refactor(agent/utils): simplified keepaliveWhile implementation using unrefed interval
- Removed the EventLoopKeepalive class and replaced it with a direct setInterval call.
- Added an unref call to the interval timer to prevent blocking the process exit.
2026-05-27 13:01:49 +02:00
Can Bölük ad8bf18189 Merge pull request #1426 from oldschoola/fix/session-emit-listener-isolation
fix(session): isolate event listener failures in agent fan-out
2026-05-27 13:45:01 +03:00
Can Bölük 0b52b8c0d0 Merge pull request #1419 from hezhiyang2000/fix/bash-busy-wait-yield
fix: eliminate Bun event-loop busy-wait with setTimeout keepalive (v2)
2026-05-27 13:39:01 +03:00
can1357 69ab1656ff chore: bump version to 15.5.4 2026-05-27 10:27:32 +02:00
oldschoola 48d63d4ddf fix(session): satisfy tsgo strict checks on void-return listener
Two TS errors in CI:
1. `agent-session.ts:1317` — `error TS1345: An expression of type 'void'
   cannot be tested for truthiness`. The listener type is
   `(event: AgentSessionEvent) => void`, so the returned value can't be
   directly tested. Same shape in `agent.ts:1079`.
2. `test/session/emit-listener-isolation.test.ts:19` — the test fixture
   for `AgentEvent.tool_execution_start` was missing the required `args`
   field.

Cast the return to `unknown` and check `instanceof Promise` instead of
duck-typing `.then` — type-safe and matches what async functions actually
return. Add `args: {}` to the test fixture.
2026-05-27 01:07:09 -07:00
hezhiyang2000 1126934300 fix: replace ReturnType<typeof setInterval> with NodeJS.Timeout
Project convention (AGENTS.md) prohibits ReturnType<> — use the
concrete type name instead. NodeJS.Timeout matches the existing
pattern used throughout the codebase (e.g. interactive-mode.ts).
2026-05-27 14:59:08 +08:00
hezhiyang2000 2159f56386 fix: add EventLoopKeepalive to eliminate idle busy-wait
Root cause: Bun 1.3.x (JavaScriptCore) busy-waits when the only
pending work is an unresolved Promise. A setInterval keepalive
keeps the event loop in epoll_wait instead of userspace spinning.

- EventLoopKeepalive: setInterval-based keepalive (re-arms after each
  firing, addressing the bot review concern about setTimeout expiry)
- keepaliveWhile(): wrapper to await a Promise with keepalive active
- Applied to getUserInput() in main.ts
- Retains yieldIfDue() and ExponentialYield from #1396

Idle CPU drops from ~100% to ~0% (wchan=do_epoll_wait).
2026-05-27 14:57:46 +08:00
oldschoola 08bccfe0ec fix(session): isolate event listener failures in agent fan-out
Both `AgentSession.#emit` (session/agent-session.ts) and `Agent.#emit`
(packages/agent/src/agent.ts) iterated listeners with no error isolation.
A synchronous throw in any subscriber aborted the for-loop, so later
subscribers (TUI rendering, ACP bridge, task executor progress,
hindsight) silently missed events. Many listeners — see
`modes/controllers/event-controller.ts:141` and
`modes/controllers/input-controller.ts:576` — are registered as
`async (event) => { await this.handleEvent(event); }`; the returned
Promise was dropped, so any rejection became an unhandled rejection.

Wrap each listener invocation in try/catch and attach a `.catch` to any
returned thenable. Errors are logged via `logger.warn` (already imported
in agent-session.ts) and `console.error` (agent.ts has no logger
dependency, keep it that way).

Test: new `test/session/emit-listener-isolation.test.ts` registers two
listeners on both classes; first listener throws (or returns a rejecting
Promise); asserts the second listener still receives the event AND no
`unhandledRejection` fires. 4 cases (sync+async × Agent+AgentSession).
All fail on current main; all pass with the fix.
2026-05-26 23:26:15 -07:00
can1357 280b638cc8 chore: bump version to 15.5.3 2026-05-27 03:20:33 +02:00
can1357 2ae57b85dc chore: bump version to 15.5.2 2026-05-27 00:46:37 +02:00