Commit Graph

7623 Commits

Author SHA1 Message Date
can1357 af33d4bfe4 ci: added macOS release signing and Homebrew automation to CI
- Added macOS CI signing and notarization steps when APPLE_* secrets are configured.
- Added strict darwin verification checks to reject ad-hoc signatures and run smoke tests.
- Added Homebrew formula publishing from release assets with SHA-256 checksums.
- Added helper scripts for signing secret upload, entitlements, and release workflows.
2026-06-08 11:49:28 +02:00
can1357 fd40148dcb fix(coding-agent): widened tiny-title worker smoke timeout for slow runners
The release_binary smoke probe spawns the tiny-title worker as a cold subprocess
of the compiled binary and pings it. On the contended macos-15-intel runner the
cold start (decompress + module-graph load, with a cold bun cache) blew past the
5s bound and failed the 15.10.3 release, while arm64/linux/win passed. The probe
only needs to prove the worker spawns and ponges at all, so the timeout is raised
to 30s — a dead worker still never ponges, so the check is unchanged in substance.
2026-06-08 07:39:40 +02:00
can1357 dad939c465 test(ai): de-flaked parallel oauth refresh test
The codex oauth ranking test asserted both `maxConcurrent === 3` and a wall-clock
bound (`elapsed < 2x refreshDelay`). The concurrency counter already proves parallel
execution deterministically — serial refreshes can never raise peak in-flight above
1 — while the wall-clock bound flaked on loaded CI runners (192ms vs 150ms) and
broke the 15.10.3 release. Dropped the redundant timing assertion.
2026-06-08 07:26:28 +02:00
can1357 3b7a44cb49 test(coding-agent): realigned stale tests with intentional behavior changes
These three suites encoded pre-refactor behavior and broke the 15.10.3 release CI.

- event-controller read grouping: the group header count now reflects aggregated
  display rows (distinct files), and a mid-turn visible-reasoning break finalizes
  the prior group but keeps it live until pending reads settle (97ae0fd3, b5eff5b3).
- gallery harness: lsp no longer attaches an instance renderer (custom rendering
  removed in d9d06134), so the custom-branch regression guard is retargeted to
  task, which still attaches its renderer and merges call+result.
- eager-todo enforcement: the prelude reminder converts to a developer-role
  message now that auxiliary messages map to developer for compaction (e13f2de5).
2026-06-08 07:17:08 +02:00
can1357 246d7dd0ab chore: bump version to 15.10.3 2026-06-08 06:52:49 +02:00
can1357 74cc940195 feat(coding-agent/edit): added streaming diff builder to stabilize in-flight preview cursor
- Added an `insertCursorLine` helper to translate streaming insert cursors into preview line numbers.
- Implemented `buildStreamingSectionDiff` to group resolved edits by operation and emit deletions before insertions, stabilizing streamed cursor progression.
- Updated `computeHashlineSectionDiff` to use the streaming builder when `options.streaming` is enabled, while preserving existing Myers diff behavior for non-streaming paths.
2026-06-08 06:52:34 +02:00
can1357 28dade85c3 fix(agent): resolved deferred asides at injection to drop stale messages
- Introduced `AsideMessage` as a message-or-thunk union so aside providers can defer injection decisions.
- Updated agent loop handling to resolve aside thunks at injection time and skip entries that returned `null`, then switched the session yield queue to `drainLazy` for deferred message building.
- Added tests validating lazy aside evaluation and staleness-aware dropping when everything becomes stale after dequeueing.
2026-06-08 06:46:41 +02:00
can1357 b075fb79b6 fix(coding-agent/modes): fixed grouped read rows freezing on pending terminal read previews
- Replaced `readGroup.finalize()` with `readGroup.seal()` when ending assistant content, after tool runs, and at trailing read-group flushes to keep grouped reads in the live region until late results settle.
- Documented the fix in `packages/coding-agent/CHANGELOG.md` for grouped read rows that could freeze on terminals when results arrived after a sibling tool closed the run.
2026-06-08 06:31:21 +02:00
can1357 587993eb1e fix(coding-agent): shared file mutation versions across edit and write tools
- Added a session-global file mutation counter and accessor methods to tool sessions.
- Updated the write tool to bump a file's mutation version after each write.
- Updated the edit tool to use session-wide mutation versions when checking stale deferred diagnostics.
2026-06-08 06:29:02 +02:00
can1357 b5eff5b3d3 fix(coding-agent/modes): sealed read groups to prevent lingering pending previews at turn end
- Added a sealed state to `ReadToolGroupComponent` to force-close a read group when a turn ends without a read result.
- Updated transcript finalization to treat pending read entries as active unless the group is sealed, preventing premature finalization when outputs are still in flight.
- Expanded turn-end pending-tool cleanup to seal `ReadToolGroupComponent` instances in addition to `ToolExecutionComponent`.
2026-06-08 06:28:56 +02:00
can1357 573fb8a9a6 chore: reformat 2026-06-08 06:20:52 +02:00
can1357 49495a072e Merge remote-tracking branch 'origin/farm/6bb79ebd/fix-resume-missing-session-friendly-error' 2026-06-08 06:20:30 +02:00
can1357 a37e6c26ca Merge remote-tracking branch 'origin/farm/b7f4bc6d/fix-tmux-resize-viewport-flash' 2026-06-08 06:20:09 +02:00
can1357 9a2db86375 feat(hashline): added resolved span reporting for block edit operations
- Added a BlockResolution type and `onResolved` callback so `resolveBlockEdits` reports each anchor's resolved span for `replace block`/`delete block` edits.
- Updated patch application to return those spans on hash-match applies and omit them during drift-recovery paths where line numbers no longer align.
- Propagated the surfaced spans through the edit tool output and added/updated tests covering replace/delete block echo and drift behavior.
2026-06-08 06:20:00 +02:00
can1357 7e72d364ce fix(coding-agent/tools): fixed multi-path find target processing and duplicate result emission
- Processed explicit multi-path `find` targets independently and merged scoped results.
- Ignored missing or invalid extra targets in multi-path `find` queries instead of failing the whole run.
- Deduplicated overlapping matches when combining per-target `find` results.
- Tracked streamed match paths to prevent repeated update rows during long-running scans.
2026-06-08 06:18:54 +02:00
can1357 d9d061348c feat(coding-agent): delivered job and LSP notifications via aside channel
- Routed background-job completions and late LSP diagnostics through the new non-interrupting aside channel so the model sees them mid-run between requests.
- Removed inline custom rendering from the LSP tool now that diagnostics surface through the shared transcript renderer.
2026-06-08 06:16:37 +02:00
can1357 dbd09cf4ee test(agent): added tests for aside timing and stale-yield queue draining
- Added an agent loop test proving aside messages are delivered after tool results, before the next model request, without interrupting tool execution.
- Added a yield-queue test confirming stale entries are excluded, the queue is cleared, and re-draining yields no messages.
2026-06-08 06:11:04 +02:00
can1357 4a2f33c905 ux(coding-agent/tools): disabled background fill for todo tool rendering
- Added `applyBg: false` to the todo tool renderer so todo output is rendered without a background.
2026-06-08 06:08:54 +02:00
can1357 de03d7f3d1 feat(agent): added non-interrupting aside message support
- Added a new aside-message source on `Agent` and exposed it in `AgentLoopConfig` as `getAsideMessages`.
- Updated the agent loop to poll aside messages after tool batches and before yielding, merging them with follow-up messages before continuing.
- Changed coding-agent session and yield queue handling to pull queued background messages via `drainMessages` at step boundaries instead of streaming-only injection.
2026-06-08 06:08:43 +02:00
can1357 c7c0b67b8c feat(coding-agent/modes): reworked late diagnostics rendering in chat transcript
- Added a LateDiagnosticsMessageComponent to render delayed LSP diagnostics as grouped tree nodes in transcript messages.
- Updated message handling to use that component and honor tool-output expansion while reusing shared diagnostics formatting.
- Added tests for shared rendering output, collapsed/expanded limits, and empty diagnostics behavior.
2026-06-08 06:04:46 +02:00
roboomp 9ec17f51a9 fix(tui): routed forced renders through the multiplexer resize debounce
The earlier guard only suppressed throttled `requestRender(false)`; forced repaints (e.g. `#finishSixelProbe`, image-budget eviction, `resetDisplay`) bypassed it through `requestRender(true)` or the direct `#doRender` path and could still paint into a still-reflowing pane inside the 50 ms settle window — recreating the flash race.

Extracted `#armMultiplexerResizeTimer(clearScrollback)` as the single arm-or-extend point used by the SIGWINCH callback, the deferral branch in `requestRender(true)`, and `resetDisplay()`. Each call cancels any queued throttled render, ORs the callers `clearScrollback` intent into `#deferredForcedClearScrollback`, and re-arms the timer. When the timer fires it consumes the flag and re-enters `requestRender(true, { clearScrollback })`; clearing `#multiplexerResizeTimer` before re-entry lets the deferred call pass straight to `#prepareForcedRender`. `stop()` clears the flag.

Added regression tests for forced `requestRender(true)` and `resetDisplay()` landing inside the debounce window.
2026-06-08 04:04:34 +00:00
roboomp a31f088ffc fix(tui): superseded queued renders during multiplexer resize debounce
A streamed-token tick that landed `requestRender(false)` in the same 30fps frame as the SIGWINCH left `#renderTimer` armed, so the throttled render could fire inside the 50 ms settle window and paint at the new geometry while tmux was still reflowing — the same race the debounce is meant to avoid. The immediate-force path used to clear that timer through `#prepareForcedRender`; the debounce path now does the same when arming, and `#scheduleRender` short-circuits while the multiplexer timer is active so any follow-on `requestRender(false)` is held off until the eventual forced render. Added a regression test covering a queued render + SIGWINCH in the same frame.
2026-06-08 03:57:22 +00:00
can1357 4fe2572629 feat(scripts/session-stats): added since-window filtering and hashline edit classification
- Added a new `--since` CLI window option and applied a timestamp cutoff to tools, edits, and followups queries, including session-count updates.
- Updated edit analytics to treat `is_error` as authoritative, added hashline op parsing from edit input, and expanded failure classification patterns for hashline error cases.
2026-06-08 05:56:03 +02:00
can1357 bd079eef6a feat(coding-agent/modes): enabled mouse-hover highlighting for plan review options
- Handled SGR mouse motion reports in `PlanReviewOverlay` to set hovered options from hit-test rows and clear highlights outside option rows.
- Updated option rendering to paint a background band on hovered, non-disabled options while preserving keyboard selection behavior.
- Enabled any-motion mouse tracking for overlays and expanded regression coverage to verify hover tracking setup and teardown.
2026-06-08 05:53:33 +02:00
roboomp 227aaef984 fix(tui): coalesced multiplexer resize events into one settled render
Tmux/screen/zellij send SIGWINCH while the pane is still mid-reflow and fire several events during drag-resize or pane-close animations. Forcing an immediate render on each event raced those mid-reflow paints — the multiplexer overwrote the TUI output and the user saw the viewport flash blank before the next throttled frame.

The SIGWINCH callback now debounces inside multiplexer sessions through a 50 ms render-scheduler timer; subsequent events cancel and re-arm it, so a single forced render fires at the final geometry once the pane is quiet. `#resizeEventPending` is still set on every event so the eventual render classifies as a resize, and stop() cancels the pending timer.

Fixes #2088
2026-06-08 03:50:50 +00:00
can1357 53917fc382 fix(agent): fixed OpenAI compaction history builder call-id tracking
- Updated `buildOpenAiNativeHistory` to maintain known and custom tool-call ID sets incrementally while appending provider payload history.
- Rebuilt call-ID state when a full-snapshot payload replaced history so stale outputs were no longer emitted after reset.
- Added compaction regression tests for codex provider payload call-ID registration and stale-result dropping.
- Scope.
- Summary.
- Type.
2026-06-08 05:46:57 +02:00
can1357 02dc03d03f fix(coding-agent): updated late diagnostics to batch messages and drop stale results
- Added per-path edit versioning in `EditTool` to drop stale late diagnostics after later edits.
- Added deferred diagnostics queueing through `queueDeferredDiagnostics` and late-diagnostic yield batching.
- Updated `UiHelpers` to render late diagnostic file path and summary lines in the chat transcript.
2026-06-08 05:46:43 +02:00
can1357 f11cbed185 feat: enabled clickable read links and made deterministic reads non-abortable
- Enabled clickable read-path output for result rows, summaries, and previews.
- Resolved read links from result paths, source metadata, internal URLs, and absolutes.
- Preserved selector suffixes while rendering line-anchor hyperlinks.
- Ignored aborted signals for plain-file and directory reads while keeping conflicts-cancel behavior.
- Added tests for non-abortable read behavior and link-label rendering regression coverage.
2026-06-08 05:40:04 +02:00
can1357 bf518d8a5e fix(coding-agent/edit): fixed edit tool headers to truncate long paths and compact diff stats
- Trimmed first-changed-line suffixes from edit path display and passed the line as a hyperlink target.
- Added middle-eliding path truncation for edit and rename file labels.
- Simplified diff-stat output to show only added/removed counts in a compact bracketed form.
2026-06-08 05:37:18 +02:00
can1357 4f6ea5ca24 test(coding-agent): added writethrough deferred diagnostics benchmark and test coverage
- Added a new `edit-lsp-writethrough.bench.ts` benchmark to measure writethrough latency with and without deferred diagnostics.
- Added a slow-server case to `lsp-diagnostics-freshness.test.ts` that confirmed deferred diagnostics arrive after a prompt inline return.
2026-06-08 05:28:06 +02:00
can1357 8307f7107a fix(coding-agent): suppressed redundant user-interrupt assistant transcript lines
- Added `isUserInterruptAbort` and `shouldRenderAbortReason` helpers so interrupt handling can distinguish Esc-based aborts from other abort reasons.
- Updated assistant transcript rendering to suppress the `Interrupted by user` line while continuing to show generic or other abort labels.
- Updated plan-review and transcript container tests to assert interrupted assistant messages no longer render the redundant interrupt line.
2026-06-08 05:27:25 +02:00
can1357 31309d6212 fix(agent): removed tool-level abort-signal bypasses for read/write/edit tool execution
- Updated `executeToolCalls` to always pass the active `toolSignal` into `tool.execute` rather than bypassing it for non-abortable tools.
- Removed the `nonAbortable` option from `AgentTool` and from the read, write, and edit tools so they can no longer opt out of abort handling.
- Documented the cancellation behavior change in the affected tool docs and package changelogs.
2026-06-08 05:27:00 +02:00
can1357 0a196a43f3 docs: update changelogs 2026-06-08 05:22:32 +02:00
can1357 b1563e4e80 fix(lsp): adjusted LSP diagnostics polling to settle unversioned publishes
- Updated `waitForDiagnostics` to accept exact document-version matches immediately and otherwise wait for a quiescence window before using the latest publish.
- Removed the old unversioned-acceptance option and applied the settle-based wait logic through inline and deferred diagnostics fetch paths.
- Added an LSP writethrough regression test ensuring stale unversioned diagnostics are ignored in favor of later fresh publishes.
2026-06-08 05:20:30 +02:00
can1357 e13f2de58a fix(coding-agent): mapped auxiliary messages to developer role for compaction
- Updated `convertToLlm` logic to emit `developer` role for custom, hook, and file-mention inputs.
- Simplified OpenAI compact output filtering to retain only `user` and `assistant` messages, removing legacy `system-reminder` pattern checks.
- Adjusted compaction and session tests to match the new developer-role mapping and expected compacted content.
2026-06-08 05:19:57 +02:00
can1357 dfeb9af61a feat(coding-agent/lsp): deferred slow LSP diagnostics during writethrough
- Added a configurable diagnostics timeout to `getDiagnosticsForFile` and wired callers to pass custom budgets.
- Refactored writethrough diagnostics retrieval to `fetchDiagnosticsWithDeferral`, waiting inline briefly and handing off slow results to the deferred callback.
- Extended deferred fetches to use a longer timeout so late diagnostics are still delivered for slow servers.
2026-06-08 05:17:37 +02:00
can1357 4442b3bfcc fix(coding-agent): marked agent reminder prompt as synthetic
- Flagged the injected reminder turn with `synthetic: true` to keep it hidden from the transcript.
2026-06-08 05:16:03 +02:00
roboomp 70e4cfcbf2 fix(coding-agent): convert createSessionManager throws into friendly CLI errors
`omp --resume <id>` and `omp --fork <id>` previously crashed with
`[Uncaught Exception] Error: Session "..." not found.` followed by a
stack trace whenever the id did not match an existing session. The
throws in `createSessionManager` were never caught by `runRootCommand`,
so they fell through to the global `unhandledRejection` handler in
`postmortem.ts` and printed the raw stack instead of a clean message.

Add `SessionResolutionError` (a dedicated subclass of `Error` with an
optional usage `hint`) and use it for every user-facing resolution
failure: unknown `--resume`/`--fork` id, `--fork` combined with
`--no-session`, and the non-interactive cross-project / moved-cwd
prompts. `runRootCommand` catches it around the `createSessionManager`
call, writes `Error: <msg>` (and the hint when present) to stderr, and
exits with code 1. Other (unexpected) errors still propagate so they
remain visible to the postmortem handler.

Fixes #2084
2026-06-08 03:15:50 +00:00
can1357 e1e75526d9 fix(coding-agent): removed system-reminder wrapper from file context
- Sent file contents as plain text instead of wrapping in `` tags.
- Joined file entries with single newlines instead of double.
2026-06-08 05:15:18 +02:00
can1357 0e78cac716 fix(coding-agent): removed abort marker and updated continue shortcut prompts
- Removed `<turn-aborted>` guidance injection from `transformMessages`, deleted `turn-aborted-guidance.md`, and dropped the synthetic abort note path for aborted/error turns.
- Updated `c`/`.` continue shortcuts to submit `manual-continue.md` as a hidden synthetic `developer` message via `session.prompt(..., { synthetic: true })` instead of sending an empty user turn.
- Updated tests and changelog notes in AI and coding-agent to reflect the revised abort-context and continue behavior.
2026-06-08 05:02:49 +02:00
can1357 cdafa6f371 fix(coding-agent): fixed live scrollback commit behavior for streaming transcript tool rows
- Computed live commit state from render diffs and drove native scrollback via safeLength.
- Fixed volatile live tool rows committing only after stream finalization in native scrollback.
- Stored append-only and volatile flags in FrozenRender snapshots for live-row tracking.
- Removed append-only streaming predicates from assistant and tool rendering paths.
2026-06-08 04:57:36 +02:00
can1357 97ae0fd34d fix(coding-agent): fixed grouped read output by splitting and merging selector rows
- Split top-level and delimited read selectors into separate rows before grouping.
- Merged duplicate same-file read selectors into one summarized row with ellipsis truncation.
- Computed grouped-read status and totals from aggregated rows for accurate summaries.
- Updated changelog entries and fixtures to document/read expectations.
2026-06-08 04:51:25 +02:00
can1357 54404878bf feat(coding-agent/cli): added custom gallery renderer for grouped read fixtures
- Gallery fixtures gained a `renderState` hook to render non-standard component states.
- Added a `read_group` filesystem fixture that uses `ReadToolGroupComponent` to render delimited, full-file, and range read states.
- Added gallery CLI coverage to verify grouped read fixture output and expected ranges.
2026-06-08 04:47:57 +02:00
can1357 da0502008e docs(coding-agent): rewrote plan-mode prompt as execution spec
- Reframed the plan artifact as a self-contained execution spec a fresh agent runs after context is cleared.
- Folded high-consensus requirements (sequencing, contracts, cutover, verification) into existing sections as inline rules.
- Banned decision-free sections and conversation back-references; kept the decision-complete self-check and completeness-wins tiebreak.
2026-06-08 04:44:19 +02:00
can1357 46578e57ba fix(tui): fixed handling of split DEC 2048 in-band resize reports
- Added a reassembly buffer for DEC 2048 in-band resize reports split across stdin reads.
- Dropped invalid or stale partial resize sequences to prevent leaked trailing bytes from entering terminal input.
- Added tests that verify split resize reports are handled and split kitty key sequences are forwarded correctly.
2026-06-08 04:42:26 +02:00
can1357 fe5d27c4a3 feat(coding-agent): removed animated shimmer border on exec blocks
- Removed sweeping bottom-edge segment from pending bash, eval, and ssh blocks; pending now shows a static accent border.
- Dropped border shimmer geometry, animate options, and related tests.
- Repurposed pending-tool 30fps redraws solely for the running task row shimmer.
2026-06-08 04:31:58 +02:00
can1357 f741431b76 fix(coding-agent/modes): fixed resume ranking to prefer literal recency and history matches
- Fixed blank-query behavior so rankSessionSearchMatches returns all sessions unchanged.
- Deduplicated prompt-history results by session path in mergeSessionRanking to avoid duplicate suggestions.
- Adjusted resume ranking to favor literal-token matches by recency, then fallback to fuzzy score.
2026-06-08 04:30:56 +02:00
can1357 ea14cee2bc test(coding-agent): refined log_experiment flagging test to use storage-run setup
- Reworked the log_experiment flagging test to create sessions and runs directly through storage APIs.
- Logged a baseline run, completed a second run, and then invoked log.execute using the baseline run ID in flag_runs.
- Verified the baseline run was marked flagged with the expected reason via storage.listLoggedRuns output.
2026-06-08 03:48:42 +02:00
can1357 65bddab719 feat(coding-agent/eval): enforced JS eval helper options as trailing object literals
- Added strict option parsing for JavaScript helpers so read/sort/uniq/counter/tree/llm/agent now reject positional options and non-object option values.
- Introduced `optionsArg` and `isPlainObject` to enforce a single trailing options object and throw clear TypeError messages on invalid calls.
- Updated the eval tool prompt to document that JavaScript helpers require one trailing options object and no extra positional arguments.
2026-06-08 03:33:45 +02:00
can1357 04a931337f fix(openai-responses): keep open function calls when starting text/thinking
- Stopped prematurely closing function call parts on output_text and thinking_start.
- Only close the open part when it is not a function_call.
2026-06-08 03:15:27 +02:00