Only a successful snapshot settled a native todo block. A `read_todos`
narrowed by a filter and a server `UpdateTodosError` both went
unanswered: no `tool_execution_end`, so the card animated forever, and
no `toolResult`, so `buildSessionContext` stripped the block on rebuild.
Every completed native todo call now settles. The refusal path carries
no `details.phases` -- `event-controller` feeds that straight into
`setTodos`, so echoing the current list back would let a call that
changed nothing overwrite live panel state. A server error is carried
through as a failed result instead of collapsing into the benign no-op.
Each regression is covered by a test verified to fail without its fix.
The previous commit paired every server-resolved todo block with a
result, but built that result in the provider from the flat snapshot.
`todoToolRenderer.renderResult` reconstructs the list exclusively from
`details.phases`, so the block survived the dangling-strip only to replay
as `Todo 0 tasks`.
Only the host computes that grouping -- the provider sees a flat list --
so `todoSync` now returns the result it already assembled and the
provider persists it verbatim. A refused snapshot never reaches the host,
so the provider's summary-only fallback still covers that path, and
exactly one result is emitted either way.
Separately, `Agent`'s Cursor buffering wrapper pushed its entry only
after awaiting the optional `cursorOnToolResult` transformer. The
provider dispatches decoded messages with `void handleServerMessage(...)`,
so a `message_end` from the same chunk could drain the buffer while a
transformer was still pending, dropping the result. The entry is now
reserved synchronously and patched in place when the transformer
resolves, keeping buffer order and still applying the customization.
Production is unaffected -- `sdk.ts` sets no transformer -- but the
option is supported and its contract returns a Promise.
Tests: a delayed-transformer case that loses the result without the
buffering change, and a replay case driving the persisted result through
`buildSessionContext` and asserting `details.phases` rebuilds a non-empty
list -- the id-pair assertion alone did not catch the empty render.
Review follow-up on two defects in the native todo bridge.
`tool_execution_end` was emitted under a freshly generated UUID, but the
interactive transcript files the visible block under the streamed
`callId` and only clears it when the ids match. The card therefore stayed
pending and animating for the rest of the session. The settled call id is
now passed to `todoSync`, making the parameter required so no caller can
silently reintroduce a mismatch.
Nothing produced a `toolResult` for these blocks either: `todoSync` only
appended a custom entry and emitted a transient event. Since
`buildSessionContext` strips any `toolCall` with no matching result, the
interaction vanished from every rebuilt transcript -- reload, branch
switch, or Ctrl+L -- leaving a "tool call elided" placeholder. A paired
result now travels the same `onToolResult` channel the other
server-resolved Cursor calls already use, including when the snapshot is
refused: the call happened, it just changed no local state.
Cursor resolves its native `update_todos`/`read_todos` tools server-side,
so the todo list never followed the model's intent locally.
Two defects, both silent:
- `agent.v1.ToolCall` is a protobuf oneof. A decoded message exposes the
selected variant as `tool: { case, value }` and has no flattened
`updateTodosToolCall` property, so the bridge recognized no native todo
call at all on the wire path.
- The synthesized `todo` block was emitted as locally runnable carrying a
`{todos}` payload the local tool's schema rejects, turning every update
into a validation error and driving a spurious continuation turn.
Todo calls are now read through the oneof, both native blocks are stamped
resolved, and local state is mirrored only from the server's confirmed
success snapshot. Partial `read_todos` responses -- narrowed by
`status_filter`/`id_filter`, or short of the server's own `total_count` --
are subsets, not the list, and are refused rather than deleting the tasks
they omit. `TODO_STATUS_CANCELLED` maps to `abandoned` instead of
reverting the task to `pending`.
The exec bridge mirrors each snapshot into session state, refreshes the
interactive panel via a synthetic `tool_execution_end`, and persists to
the session branch so the list survives reloads, rewinds, compaction, and
session switches. Existing phase grouping is preserved.
Regression tests drive the bridge with wire-encoded protobuf, which is
the only shape production ever sees; all six fail without this change.
- Added `listDisabledCredentials` and `refreshSnapshot` methods to credential stores along with API endpoints and wire schemas.
- Added `authorizedAt` timestamps and Anthropic OAuth grant TTL constants to track credential lifecycles.
- Updated the usage CLI to render auto-disabled credential tombstones and grant expiration warnings.
- Added comprehensive unit and broker integration tests covering the new credential management features.
Passed session-scoped settings through Edit and Write generated-file checks and fell back to schema defaults when no global singleton exists.
Guarded inline image sizing against an uninitialized global settings proxy and added isolated-session regression coverage.
Fixes#6549
- Remove uniform language inference requirement, allowing mixed-language paths to rewrite each file in its own language.
- Update `ast_edit_blocking` in `crates/pi-natives/src/ast.rs` to compile rewrite rules per language and skip unsupported languages gracefully.
- Update `ast-edit.md` prompt documentation to reflect mixed-language path support.
- Add test coverage verifying mixed-language tree rewrites.
- Treated transient non-array retain items as absent during TUI streaming.
- Added regression coverage for malformed partial renderer arguments.
- Documented the fix in the coding-agent changelog.
Fixes#6528
- Added language-specific code formatters for JavaScript, Julia, Python, and Ruby to improve display rendering.
- Integrated display formatting into browser run and eval render tools while preserving verbatim execution.
- Added comprehensive test suites verifying formatting stability, lexical safety, and streaming behavior.
- Implemented a structured web-search query parsing module supporting directives, tokenization, date parsing, and syntax serialization.
- Updated search providers to map query directives and date bounds to native provider parameters and filters.
- Added lenient result constraint post-filtering and configuration settings for enhanced engine routing.
- Added comprehensive unit and integration tests covering query parsing, constraint filtering, and provider-specific request mapping.
The rebuild dedup dropped any preserved pendingTools component whose toolCallId
had a persisted toolResult. A background task's initial async.state=="running"
result is persisted while EventController#handleToolExecutionEnd deliberately
keeps its component in pendingTools so a later tool_execution_update/_end can
settle it. Dropping that still-live handle stranded those updates on the running
snapshot. Only terminal results are now owned by the replay; running async
handles stay preserved. Added a regression covering the running-task case.
rebuildChatFromMessages preserves the live pendingTools components across its
clear+replay so streaming keeps routing into them. That preservation assumed
every pending-tool component was still dangling (its result outside
state.messages). Once a tool's result had landed in the session entries while
its component still lingered in pendingTools (a rebuild racing the
tool-completion event, or a background/displaceable snapshot), the replay
reconstructed the completed block from the persisted toolResult AND the
preserved live component was re-appended, so the same tool block rendered twice.
Resolved tool calls are now dropped from the preserved live set and owned by the
replay; only genuinely in-flight (dangling, replay-stripped) calls are preserved.
Fixes#6516
Replaced direct assertions on internal ordering and recovery helpers with an isolated broker integration test. The test seeds recovered terminal metadata, starts an active daemon, sends the authenticated list RPC, and asserts active-first ordering, true exit-time recency, the terminal cap, and protocol serialization/parsing.
Made the ordering and recovery helpers internal again because production wiring now supplies their coverage.
Fixes#6517
Recovery unconditionally restamped every non-detached record's exitedAt with the restart timestamp, including already-exited/failed ones, so after an idle-broker restart the list history cap could keep arbitrary directory-order entries and drop the genuinely most-recently-exited process.
Reap only records that were still alive at recovery; terminal records keep their real exitedAt/exitReason. Extract reapRecoveredSnapshot() and cover it with tests.
Fixes#6517
- Added guidelines stating that concurrent edits to the same files are safe.
- Specified prerequisites for safe overlap including skipping validation and defining contracts up front.
The broker `list`/ps op sorted every daemon record by createdAt ascending and never pruned, so in a long-lived project the active (newest) daemon rendered behind all exited ones and the response grew without bound.
List non-terminal daemons first (oldest to newest) and cap exited/failed history at the 10 most recently exited; truncated records stay addressable by name via describe/logs/restart.
Fixes#6517
- Add usage history, reporting, and client summary endpoints to the auth broker.
- Implement SQLite persistence, batching, and periodic flushing for observed client usage.
- Integrate usage reporting into the coding agent session for completed assistant messages.
- Update stats aggregator and dashboard to retrieve usage history snapshots from the broker.
- Added tool schema and description compaction for oversized data payloads in `packages/coding-agent/src/debug/raw-sse-buffer.ts`.
- Implemented head-tail trimming to preserve leading and trailing event fields when events exceed character budgets.
- Added test coverage verifying SSE debug event truncation, elision markers, and JSON-safe tool compaction behavior.
- Added available shell builtins list to the bash tool prompt template.
- Added helper to check if shell builtins are disabled via settings or environment.
- Removes `model` field from task item/schema, TaskParams, and TaskItem types.
- Removes model selector validation, formatting, and approval display logic.
- Updates task tool priority docs to reflect that model is no longer per-call overridable.
- Updates eval agent() helper docs and prompt templates to remove model parameter.
- Updates tests to reflect removal of model override capability.
Required CLI model registries to expose getAvailable() and used that authenticated
set whenever callers omit availableModels. Deferred SDK and bench/dry-balance
resolution now lets configured roles beat unauthenticated catalog id collisions.
Updated resolver test registries and made the #6508 regression omit the explicit
availableModels option, covering the deferred-caller path from the review.
Fixes#6508
`resolveCliModel` ran findExactCliModel's unauthenticated catalog fallback
before configured-role resolution, so the bundled `cursor/default` model (bare
id `default`) shadowed a configured `modelRoles.default`. On machines without
Cursor credentials `--model default` failed with `No API key found for cursor`
instead of resolving the configured, authenticated default role.
Defer the catalog fallback: explicit provider/id references and authenticated
bare ids still win over roles, a configured role now beats an unauthenticated
catalog-only id, and the catalog id is still recovered via the trailing fuzzy
fallback when no role matches.
Fixes#6508
- Large-paste menu now saves pastes as local://paste-N.md instead of
local://attachment-N, giving the artifact a markdown extension and a
clearer name.
- Add `resolveFallbackTool` callback to `AgentOptions` and `AgentLoopConfig` that resolves tool calls not found in the advertised set.
- Use the callback as a third lookup step after `name` and `customWireName` match, enabling side transports like `xd://` device mounts.
- Add test coverage verifying the fallback resolves known devices and preserves "not found" errors for unknown names.
- Wire the coding agent's device registry as `resolveFallbackTool` in both `createAgentSession` and `streamAgentSession` paths.
- Prepended a model-visible PREVIEW_PENDING_NOTICE to ast_edit preview
results; the TUI-only proposed badge never reached the model, so
preview diffs read as already-applied edits.
- Rendered the resolve reminder with the source tool name via
Handlebars instead of a generic 'This is a preview'.
- Documented the xd://resolve / xd://reject two-phase flow in the
ast_edit tool prompt.
- Kept op required in the todo schema; lenientArgValidation now routes raw args to execute(), where resolveTodoParams re-validates and repairs an omitted op (list -> init, phase+items -> append, bare items on empty list -> init).
- Ambiguous op-less calls surface the schema error as a retryable tool error instead of a hard validation failure.
Redacted assistant anthropicServerTool blocks alongside redactedThinking and providerPayload so /share never uploads raw server_tool_use input or web_search_tool_result encrypted_content.
Persisted native server-tool calls and web-search results in assistant content, replayed them only to the issuing Anthropic provider, and retained Umans gateway filtering.
Fixes#6495
- Add `CodexAttestationProvider` hook and `getCodexAttestationHeader` resolver that gates on ChatGPT-OAuth credentials.
- Integrate attestation into WebSocket transport, SSE event stream, and OpenAI compaction requests.
- Wire `setCodexAttestationProvider(generateCodexAttestation)` in model-registry init for OAuth sessions.
- Added `LIVE_DELEGATION_MESSAGE_TYPE` constant and delegation message handling for voice sessions.
- Implemented turn-based transcript coalescing with user and assistant turn counters.
- Added transcript display row with normalized rendering in the live visualizer.
- Refactored controller to send delegation messages via `sendCustomMessage` with configurable frame styling.
- Removed microphone permission error reporting from silence detection logic.
- Switched from `miniaudio` to `maudio` Rust crate and added `AudioCapture` and `AudioPlayback` native classes.
- Removed browser-side audio infrastructure including Web Audio API, audio worklet processor, and WebRTC runtime.
- Migrated STT recorder and transcriber modules to use native `AudioCapture` with callback-based streaming.
- Replaced streaming audio player with native `AudioPlayback` that writes PCM directly without TypeScript intermediaries.
- Removed ffmpeg, wav, and platform-specific playback commands from the audio toolchain.
- Replaced puppeteer-based WebRTC with native LiveWebRtcPeer for cross-platform live audio delivery.
- Added cross-platform microphone capture via miniaudio and Opus codec integration for live encoding/decoding.
- Added Apple DeviceCheck attestation token generation via raw Objective-C FFI for macOS.
- Updated live session model to "gpt-live-1-codex" and default voice to "sol" across protocol and controller.
- Added LiveWebRtcPeer and deviceCheckGenerateToken to the public native bindings API.
- Extracted #formatToolExecution into a standalone formatDefaultToolExecution module.
- Updated xdev renderXdevCall and renderXdevResult to use the default card when no mounted renderer exists.
- Added fallback rendering that shows tool label, args, and output with appropriate theming.
- Added integration test verifying generic card renders for mounted tools without bespoke renderers.
- Catch execution errors in XdevRegistry and render them using the mounted tool's error handling.
- Preserve xdev dispatch context when device execution fails.
- Add `pinSessionOAuthAccount` storage method and active account flag to the auth storage API.
- Introduce session account selector component, controller logic, and interactive mode delegation.
- Implement the `/session pin` builtin slash command with text listing and account pinning capabilities.
- Add unit tests covering session account selection, component navigation, and command handling.
requestRender(true) only queues the paint via setImmediate; calling
prewarm() synchronously afterwards put the worker spawn syscall ahead
of the render callback in the same loop turn. Schedule the prewarm
with its own setImmediate (FIFO after the render callback) so the
first startup frame paints before the subprocess spawns.