- Added a new `--since` CLI window option and applied a timestamp cutoff to tools, edits, and followups queries, including session-count updates.
- Updated edit analytics to treat `is_error` as authoritative, added hashline op parsing from edit input, and expanded failure classification patterns for hashline error cases.
- Handled SGR mouse motion reports in `PlanReviewOverlay` to set hovered options from hit-test rows and clear highlights outside option rows.
- Updated option rendering to paint a background band on hovered, non-disabled options while preserving keyboard selection behavior.
- Enabled any-motion mouse tracking for overlays and expanded regression coverage to verify hover tracking setup and teardown.
- Updated `buildOpenAiNativeHistory` to maintain known and custom tool-call ID sets incrementally while appending provider payload history.
- Rebuilt call-ID state when a full-snapshot payload replaced history so stale outputs were no longer emitted after reset.
- Added compaction regression tests for codex provider payload call-ID registration and stale-result dropping.
- Scope.
- Summary.
- Type.
- Added per-path edit versioning in `EditTool` to drop stale late diagnostics after later edits.
- Added deferred diagnostics queueing through `queueDeferredDiagnostics` and late-diagnostic yield batching.
- Updated `UiHelpers` to render late diagnostic file path and summary lines in the chat transcript.
- Enabled clickable read-path output for result rows, summaries, and previews.
- Resolved read links from result paths, source metadata, internal URLs, and absolutes.
- Preserved selector suffixes while rendering line-anchor hyperlinks.
- Ignored aborted signals for plain-file and directory reads while keeping conflicts-cancel behavior.
- Added tests for non-abortable read behavior and link-label rendering regression coverage.
- Trimmed first-changed-line suffixes from edit path display and passed the line as a hyperlink target.
- Added middle-eliding path truncation for edit and rename file labels.
- Simplified diff-stat output to show only added/removed counts in a compact bracketed form.
- Added a new `edit-lsp-writethrough.bench.ts` benchmark to measure writethrough latency with and without deferred diagnostics.
- Added a slow-server case to `lsp-diagnostics-freshness.test.ts` that confirmed deferred diagnostics arrive after a prompt inline return.
- Added `isUserInterruptAbort` and `shouldRenderAbortReason` helpers so interrupt handling can distinguish Esc-based aborts from other abort reasons.
- Updated assistant transcript rendering to suppress the `Interrupted by user` line while continuing to show generic or other abort labels.
- Updated plan-review and transcript container tests to assert interrupted assistant messages no longer render the redundant interrupt line.
- Updated `executeToolCalls` to always pass the active `toolSignal` into `tool.execute` rather than bypassing it for non-abortable tools.
- Removed the `nonAbortable` option from `AgentTool` and from the read, write, and edit tools so they can no longer opt out of abort handling.
- Documented the cancellation behavior change in the affected tool docs and package changelogs.
- Updated `waitForDiagnostics` to accept exact document-version matches immediately and otherwise wait for a quiescence window before using the latest publish.
- Removed the old unversioned-acceptance option and applied the settle-based wait logic through inline and deferred diagnostics fetch paths.
- Added an LSP writethrough regression test ensuring stale unversioned diagnostics are ignored in favor of later fresh publishes.
- Updated `convertToLlm` logic to emit `developer` role for custom, hook, and file-mention inputs.
- Simplified OpenAI compact output filtering to retain only `user` and `assistant` messages, removing legacy `system-reminder` pattern checks.
- Adjusted compaction and session tests to match the new developer-role mapping and expected compacted content.
- Added a configurable diagnostics timeout to `getDiagnosticsForFile` and wired callers to pass custom budgets.
- Refactored writethrough diagnostics retrieval to `fetchDiagnosticsWithDeferral`, waiting inline briefly and handing off slow results to the deferred callback.
- Extended deferred fetches to use a longer timeout so late diagnostics are still delivered for slow servers.
- Removed `<turn-aborted>` guidance injection from `transformMessages`, deleted `turn-aborted-guidance.md`, and dropped the synthetic abort note path for aborted/error turns.
- Updated `c`/`.` continue shortcuts to submit `manual-continue.md` as a hidden synthetic `developer` message via `session.prompt(..., { synthetic: true })` instead of sending an empty user turn.
- Updated tests and changelog notes in AI and coding-agent to reflect the revised abort-context and continue behavior.
- Computed live commit state from render diffs and drove native scrollback via safeLength.
- Fixed volatile live tool rows committing only after stream finalization in native scrollback.
- Stored append-only and volatile flags in FrozenRender snapshots for live-row tracking.
- Removed append-only streaming predicates from assistant and tool rendering paths.
- Split top-level and delimited read selectors into separate rows before grouping.
- Merged duplicate same-file read selectors into one summarized row with ellipsis truncation.
- Computed grouped-read status and totals from aggregated rows for accurate summaries.
- Updated changelog entries and fixtures to document/read expectations.
- Gallery fixtures gained a `renderState` hook to render non-standard component states.
- Added a `read_group` filesystem fixture that uses `ReadToolGroupComponent` to render delimited, full-file, and range read states.
- Added gallery CLI coverage to verify grouped read fixture output and expected ranges.
- Reframed the plan artifact as a self-contained execution spec a fresh agent runs after context is cleared.
- Folded high-consensus requirements (sequencing, contracts, cutover, verification) into existing sections as inline rules.
- Banned decision-free sections and conversation back-references; kept the decision-complete self-check and completeness-wins tiebreak.
- Added a reassembly buffer for DEC 2048 in-band resize reports split across stdin reads.
- Dropped invalid or stale partial resize sequences to prevent leaked trailing bytes from entering terminal input.
- Added tests that verify split resize reports are handled and split kitty key sequences are forwarded correctly.
- Fixed blank-query behavior so rankSessionSearchMatches returns all sessions unchanged.
- Deduplicated prompt-history results by session path in mergeSessionRanking to avoid duplicate suggestions.
- Adjusted resume ranking to favor literal-token matches by recency, then fallback to fuzzy score.
- Reworked the log_experiment flagging test to create sessions and runs directly through storage APIs.
- Logged a baseline run, completed a second run, and then invoked log.execute using the baseline run ID in flag_runs.
- Verified the baseline run was marked flagged with the expected reason via storage.listLoggedRuns output.
- Added strict option parsing for JavaScript helpers so read/sort/uniq/counter/tree/llm/agent now reject positional options and non-object option values.
- Introduced `optionsArg` and `isPlainObject` to enforce a single trailing options object and throw clear TypeError messages on invalid calls.
- Updated the eval tool prompt to document that JavaScript helpers require one trailing options object and no extra positional arguments.
Codex review on #2082 found that the Responses compatibility encoder still
kept a singleton open function-call item. When MiniMax object-argument chunks
are flushed late by contentIndex, a later parallel toolcall_start could close
the first item and make the late delta append to the second item instead.
Keep OpenFunctionCall state in a map keyed by contentIndex, allocate Responses
output indexes when items open, and close each function item from its matching
toolcall_end. Late deltas now use event.contentIndex, preserving deferred
object-argument flushes for the original tool call even after later parallel
starts.
Added an encodeStream regression that starts two parallel calls, emits the
first call's arguments only after the second start, and verifies both argument
delta/done events and output_item.done payloads stay attached to their own
output indexes.
Fixes#2080
- Updated workflow detection in `workflow.ts` to recognize only the lowercase `workflowz` token for detection and highlighting.
- Changed workflow instructions and tips to reference the new `workflowz` trigger for eval fan-out behavior.
- Updated workflow tests to assert `workflowz` matches correctly and previous `workflow`/`workflows` forms no longer trigger.
- Stubbed `isSettingsInitialized` to false so crest timing stays deterministic.
- Replaced magic timestamp with named `CLASSIC_CREST_VISIBLE_MS` constant.
- Anthropic streaming now yielded explicit `ping` events and propagated them through the stream event types.
- Ping keepalive markers reset idle liveness handling so long-running streams no longer stalled as idle.
- Usage and quota limit errors were treated as non-retryable, and tests were added to keep transient rate-limit retries unchanged.
- Tracked focus transitions in TUI and flagged renders after focus changes to allow unknown viewport mutations.
- Used the new focus-change flag when deciding explicit viewport mutation so subsequent frames repaint even when the terminal lacks a viewport-oracle position.
- Added a regression test covering menu teardown focus changes on unknown-viewport terminals with eager scrollback risk.
- Replaced manual settings callback sets with a shared `SettingSignal` that snapshots listeners and skips over individual callback failures.
- Updated `provider.appendOnlyContext`, `statusLine.sessionAccent`, and `hindsight` hook dispatch to use the new signal and added a Settings test confirming a throwing append-only listener does not stop other listeners.
- Adjusted the duplicate-tool-results regression test to handle `tool_calls` being absent before mapping IDs.
Codex review on #2082 caught that the merged-args fix still emitted
`delta = JSON.stringify(rawArgs)` per object chunk. Every downstream
consumer that follows the OpenAI `toolcall_delta` contract by
concatenating deltas — `packages/agent/src/proxy.ts` reconstructing
`partialJson` (lines 286-290), `openai-chat-server` forwarding
`tool_calls[].function.arguments`, `openai-responses-server` appending
to `cur.argsText`, `anthropic-messages-server` emitting
`input_json_delta` — would have seen an invalid sequence like
`{"input":"a"}{"input":"b"}` once a host fragmented the args across
deltas, and `parseStreamingJson`'s repair-then-partial-parse fallback
collapses that to `{}` or just the first object. The source-side
`.result()` was correct because `block.arguments` already held the
merged result, but every concat-based reader downstream lost the args.
Suppress object-chunk wire deltas during streaming (the merge still
runs into `block.partialArgs`/`block.arguments`) and flush the full
merged JSON as a single concat-safe delta in `finishToolCallBlock`
right before `toolcall_end`. Concat consumers now reconstruct the args
unconditionally — single-chunk case is still `"" + full_json`,
multi-chunk case is `"" + "" + … + full_json`, both parse to the same
merged object.
Two new regression tests in `issue-2080-repro.test.ts` accumulate
`event.delta` the way `proxy.ts` does and assert `JSON.parse(accum)`
matches the source-side merged args, covering both the multi-chunk
fragmented-string shape and the single-chunk shape (no #1776
regression). The three pre-existing tests for the merge behaviour
itself still pass unchanged.
- Added reconstructed SSE event emission for OpenAI, Azure, and Anthropic streams.
- Added raw SSE text to debug report bundles, including raw-sse.txt output.
- Included dropped-record metadata in raw SSE text when events were trimmed.
- Updated raw SSE and sse-debug tests for observer-based capture and safety checks.
- Cached setting path segments and memoized `Settings.get()` results, clearing caches on updates.
- Triggered session-name and session-accent callbacks only when effective values changed, with error-safe dispatch.
- Memoized status-line and interactive accent resolution with cache invalidation on settings/theme/session changes.
- Added regression tests for keybinding precedence, status-line, settings, and accent cache behavior.
- Built canonical and alias key maps for action/custom handlers and rebuilt them on changes.
- Preserved the first handler when custom key alias collisions occurred during rebuilds.
- Exported canonicalKeyId and addKeyAliases from keybindings for external use.
MiniMax-compatible OpenAI-completions hosts stream `function.arguments` as
an object instead of the OpenAI JSON-string contract. The accumulator
added in #1776 wrote `block.partialArgs = rawArgs` per chunk, which
worked for the assumed single-chunk shape but silently dropped every
chunk but the last when the host fragmented the args across deltas. For
an `edit` call that meant a tail-slice of the patch text was applied —
e.g. a `replace 91..91:` body whose surrounding rows were lost, surfacing
as the hashline applier "widening" the delete across the original line's
neighbours (the symptom the reporter saw with MiniMax-M3).
Shallow-merge chunks into the accumulated args object. For shared string
keys, distinguish cumulative restatements from per-chunk-delta fragments
with `startsWith` so cumulative payloads are not doubled and delta
fragments are not lost. The single-chunk shape covered by the existing
#1776 regression test remains a no-op because there is no prior value to
merge with. New `issue-2080-repro.test.ts` covers the three multi-chunk
shapes: delta fragments, cumulative restatements, and chunks that
introduce new keys.
Fixes#2080
- Updated classic shimmer to compute band position from elapsed time at 30 cells per second instead of a fixed sweep duration.
- Updated KITT shimmer to use fixed-speed ping-pong motion over the same 2xrange cycle so round-trip timing scales with message length.
- Added hardware cursor state caching to reuse last known position and visibility.
- Updated cursor flow to preserve target row/column/visibility across rerenders.
- Reworked cursor output emission to skip unchanged cursor moves and hide/show writes.
- Adjusted `LspTool` to retry only zero/decl-only references with two 250ms waits for project-aware servers.
- Removed raw-mode GitHub repo README API fallback in `read`, so `:raw` now renders the page directly.
- Replaced regex heuristics in `isImagePlaceholderAnswer` with a fixed normalized placeholder set.
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.