- Extracted decodeStreamedToolArgs into tool-args-reveal.ts and used it from both the live event path and transcript rebuilds, so mid-write theme/settings/focus replays no longer show stale streamed write/edit/eval content.
- Fixed the smoothing-off live path returning stale provider-parsed args.
- Documented the mandatory shared decode in the AGENTS.md streaming-preview hazard note; added changelog entries for this batch.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
Mid-turn renderSessionContext (settings overlay close, focus attach during streaming) now hands the rebuilt todo snapshot back to the EventController via the new inheritDisplaceableTodo method instead of sealing it. Idle rebuilds keep the historic seal path.
Added a regression test that asserts the trailing todo snapshot is published to the controller and stays displaceable while session.isStreaming is true.
Fixes#3516
Dropped eager todo snapshot displacement from tool_execution_start, streaming message_update, and the rebuild assistant-iteration step. Displacement now runs only when the next todo's successful result lands, so a failed follow-up leaves the last-good todo panel on screen.
Added regression coverage for the failed follow-up case and updated the streamed-second-todo test to drive displacement from the success result.
Fixes#3516
Resolved any tracked todo snapshot before storing a fresh one in the rebuild paths so an assistant message replaying multiple todo tool calls collapses to the final snapshot.
Added a renderSessionContext regression test for two todo tool calls in one rebuilt assistant message.
Fixes#3516
Kept successful todo result blocks live until a later todo update replaces them or the turn ends.
Added regression coverage for same-turn todo snapshot replacement after intervening tool output.
Fixes#3516
Some providers (MiniMax, GLM, DeepSeek) return thinking blocks in
their responses even when reasoning is disabled — the model generates
thinking content regardless of the reasoning_effort parameter.
When the user sets thinking level to "off", they expect no thinking
content to be visible. Previously, thinking blocks would still appear
because the hideThinkingBlock setting was independent of the thinking
level and defaulted to false (show).
Fix: add effectiveHideThinkingBlock computed property that returns
true when hideThinkingBlock is true OR the session thinking level is
"off". All render paths (streaming, transcript rebuild, component
construction) now read the effective value instead of the raw setting.
The toggle (Ctrl+T) is guarded: when thinking is off, it shows a
status message ("Thinking is off — enable thinking to show blocks")
instead of silently no-op'ing or corrupting the persisted setting.
Fixes#626
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
Tail appended transcript JSONL instead of rebuilding rendered history on every poll, collapse compacted history for live chat rendering, and replace synchronous session rewrites so tailers detect historical changes.
Fixes#3258
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
- Centralized draft state and image management by migrating fields from context to the CustomEditor component.
- Standardized transcript row construction by introducing shared helpers for background jobs, IRC traffic, and file mentions.
- Refactored redundant UI logic and helper functions into reusable utility modules to streamline message submission and component rendering.
- Standardized event handler types by consolidating lifecycle definitions into a shared module while maintaining public API stability.
- Added a `proseOnlyThinking` configuration setting to suppress raw code blocks in AI thinking traces.
- Implemented `formatThinkingForDisplay` utility to replace code blocks with ellipses in the UI.
- Integrated runtime toggling and live refreshing of message components via streaming reveal controllers.
- Added a live tokens-per-second indicator to the assistant thinking pulse.
- Verified logic with new unit and integration tests for thinking block presentation.
- Added tracking for expected cache invalidations during model changes, compactions, and plan-mode transitions.
- Included `cacheMissExplainedAt` metadata in session context to prevent displaying misleading cache miss warnings in the transcript.
- Updated controller logic to reset assistant usage markers when mode-switching or performing actions that invalidate the prompt cache.
- Added `display.cacheMissMarker` setting to enable visual indicators for prompt cache invalidation in assistant messages.
- Included updated theme symbols to represent cache misses across supported icon sets.
- Configured the selector controller to rebuild the chat display when the marker visibility setting is toggled.
- Implemented `detectCacheInvalidation` to identify when model requests lose their prompt cache.
- Added `CacheInvalidationMarkerComponent` to display a slim notice above affected assistant turns.
- Updated `ChatTranscriptBuilder` and `EventController` to track session usage and inject markers dynamically.
- Included comprehensive test coverage for invalidation detection logic and UI rendering.
- Refreshed all injected message frames with a consistent rounded-outline card design.
- Implemented icon-tagged headers for hooks, local skill invocations, and generic custom messages.
- Updated skill invocation cards with home-shortened paths, dynamic line count units, and compact invocation argument display.
- Unified branch summary styling with compaction banners to provide a consistent visual language for history collapse points.
- Removed leaked absolute home directory paths from skill metadata displays.
- Swapped legacy `zodToWireSchema` for `toolWireSchema` to normalize tool schemas.
- Updated `getSchemaPropertyKeys` in `tool-index.ts` to process schemas via `toolWireSchema`.
- Refactored tool token estimation in `context-usage.ts` to utilize the new schema helper.
- Fixed ArkType assertion checks in test helpers to correctly verify instances against `arkType.errors`.
- Aligned test suites and mock specifications with raw schema-based parameters instead of manually stringified JSON structures.
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
- Created AdvisorRuntime and AdviseTool to drive a read-only advisor agent that delivers severity-tagged advice (nit, concern, blocker) with interruption policy and transcript delta rendering.
- Added /advisor slash command with on/off/status/dump subcommands to control advisor lifecycle and inspect advisor metrics (model, messages, tokens, cost).
- Added advisor.enabled and advisor.subagents settings to enable passive advisor review on main agent and spawned task/eval subagents.
- Implemented advisor message rendering with severity-color badges (blocker=error, concern=warning, nit=muted) in chat log and status line indicator (++ badge).
- Extended yield-queue and session-history-format to support advisor batching and optional thinking block inclusion.
- Canonicalized assistant and thinking messages by trimming and collapsing dot text.
- Skipped rendering assistant and thinking blocks when canonicalized content was empty.
- Filtered ACP thinking notifications and session outputs to ignore placeholder content.
- Added canonicalizeMessage tests for undefined, blank, whitespace, and dot-only inputs.
- Removed the streaming guard that previously rejected /tan while the parent response was still generating.
- Passed "deliverAs: \"nextTurn\"" when sending the background dispatch breadcrumb and kept "triggerTurn: false" so an in-flight turn is not steered.
- Skipped rebuilding chat messages during streaming sessions and updated tests to cover the non-blocking dispatch path.
- Added a space-hold gesture state machine in CustomEditor, tracking repeated spaces, detecting holds beyond SPACE_HOLD_THRESHOLD, and firing start/end callbacks via a release timer.
- Hooked editor space-hold callbacks in InputController so STT toggles on hold start and again on release when STT is enabled.
- Added tests for space-hold start/stop behavior and updated keybinding docs to describe the hold-to-record STT workflow.
- Routed `customType: "handoff"` messages to the compact divider path in Agent Hub and UI helpers.
- Added handoff summary expansion that extracts context text and strips `<handoff-context>` wrappers.
- Refactored shared divider rendering into `SummaryDividerComponent` used by compaction and handoff messages.
- Added session-domain modules and exports for session-entries, context, listing, loader, and migrations.
- Changed persistence to async append writes plus writeTextAtomic, removing sync line APIs.
- Added compaction-aware session context rebuild with dangling tool-call cleanup.
- Added resumable session resolution with status inference, id/stem/suffix matching, and backup recovery.
With display.showTokenUsage on, the usage row was rendered inside the
assistant block above the turn's tool blocks. Finalizing the assistant
block was therefore deferred, and the late append recommitted the
already-committed tool rows, duplicating them in scrollback (worst with
parallel tool calls).
The assistant block now always finalizes as soon as a tool-call appears,
and the usage row is emitted as a standalone finalized block below the
turn's tool blocks across all three render paths (live event-controller,
transcript rebuild, agent-hub). The now-dead setUsageInfo/#usageInfo path
on AssistantMessageComponent is removed.
- Filtered dot-only or blank thinking blocks so they no longer render as assistant thought.
- Adjusted assistant-message and streaming-reveal logic to use visible-thinking helpers for consistency.
- Added a SessionFocusController to switch transcript and input context between main and subagent sessions.
- Added agent-hub Enter activation and double-left return behavior for focused local agents.
- Added view-session-based event and render logic to avoid stale focus-session state.
- Added status-line focused agent display with ghost icon and focused-mode border dimming.
- Added snapcompact.shape setting with auto and variant options in agent config.
- Implemented resolveShape support for forced variants, auto provider winners, and repricing.
- Threaded resolved shape into compaction and inline-image flows for pricing and rendering.
- Added optional `hasSteeringMessages` config hook and limited steering checks to boundaries.
- Fixed interrupted tool-batch steering by keeping queued messages until boundary handling.
- Added idle text and image submissions to steer queueing when no input waiter exists.
- Auto-continued resumable sessions after queued steering and preserved submit metadata.
- Fixed `/dump` output (`session-dump-format.ts`) and the RPC `get_state` `dumpTools` payload (`rpc-mode.ts`) serializing Zod schema instances — leaking `def`/`shape` internals and stringified methods — by converting parameters through `zodToWireSchema` like providers receive.
- Fixed context-usage token estimation stringifying the Zod `def` tree, which overcounted tool schema tokens in the status line.
- Fixed tool-discovery indexing (`tool-index.ts`) returning empty `schemaKeys` for Zod tools, weakening BM25 ranking; keys are now recovered via wire conversion.
- Fixed the extension inspector panel rendering "(no arguments)" for Zod tools by reading properties off the converted wire schema.
- Added coverage in `tool-index.test.ts` plus new `context-usage.test.ts` and `session-dump-format.test.ts`, and a changelog entry.