- Implemented model-specific frame-size billing for Anthropic, OpenAI, and Google.
- Changed compaction shape resolution to bind model variants to ideal frame sizes.
- Updated tests and docs to reflect new frame-size and budget behavior.
- Updated snapcompact selectors and previews to pass `ShapeTarget` into `resolveShape` so `auto` is model-tuned.
- Applied `providerFrameBudget` in session compaction so generated archives stay within provider image caps.
- Replaced inline hard-coded image limits with `providerImageBudget` and skipped rasterization at cap.
- Added OpenRouter inline-transformer tests for cap exhaustion and existing-image budget exhaustion behavior.
Sent Hindsight retain timestamps with local timezone offsets and supplied timestamps for automatic and queued retains. Added regression coverage for client serialization and backend timestamp propagation.\n\nFixes #2363
- Repointed `SHAPES.anthropic` and the `anthropic-messages`/unknown-API fallback in `resolveShape` at the `6x12-dim` variant: production mono eval on claude-fable scored f1 .840 vs .877 for the repeated grid (within noise at n=25) at 37% lower cost, with no refusals.
- Reworded the `snapcompact.shape` descriptions in `settings-schema.ts` to drop per-provider eval-winner claims from the variant help text.
- Updated `snapcompact.test.ts` (new `6x12-dim` default render/compact assertions, `8x8r-bw` exercised via `resolveShape`), `snapcompact-inline.test.ts` frame math for the new geometry, and the `docs/compaction.md` shape sentence.
- Amended the snapcompact `[Unreleased]` entry that said the Anthropic default stayed `8x8r-bw` and added the default-switch entry.
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
- Added `thinking.requiresEffort` to `ThinkingConfig` and baked it in `deriveThinking`/`fillThinkingWireDefaults` via `impliesMandatoryReasoning` for reasoning-only upstreams: Gemini 3.x, Gemini 2.5 Pro, the OpenAI o-series, MiniMax M2, and thinking-only `-reasoner`/`-reasoning` SKUs.
- Added `minimumSupportedEffort()` to `model-thinking.ts` as the clamp target for thinking-off requests on flagged models.
- Moved `stripThinkingVariantToken`/`findThinkingVariantToken` into `identity/family.ts`, taught them the `-reasoning`/`-reasoner` spellings, and re-pointed the `variant-collapse` and coding-agent `model-resolver` imports.
- Dropped `requiresEffort` (with `effortRouting`/`suppressWhenOff`) from collapsed-pair thinking surfaces in `derivePairThinkingSurface`, since the collapsed pair routes off to the bare backing id.
- Regenerated `models.json` and covered derivation, backfill, and reasoning-token pairing in `model-thinking.test.ts` and `variant-collapse.test.ts`.
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.
- Resolved OpenAI shape resolution to `openai` and default to `8on16-bw`.
- Fixed catalog generation to collapse effort tiers before provider grouping.
- Updated help text and schemas to describe `8on16-bw` as the OpenAI auto default.
- Added production `render_pages` and `mono_prod` scripts for end-to-end QA output.
- Added the `6x12-dim`, `8x13-bw`, `8on16-bw`, `doc-8on16-bw`, `doc-8on16-sent`, and `doc-8on16-sent-dim` options to the `snapcompact.shape` enum in `settings-schema.ts` with per-variant descriptions naming each eval winner.
- Updated the `snapcompact.shape` paragraph in `docs/compaction.md` to list the new grid and two-column doc variants.
- Updated the coding-agent changelog: rewrote the `snapcompact.shape` Added entry for the new variants and recorded the Changed/Fixed entries for the already-landed variant-alias and pending-preview commits (one contiguous changelog run).
- Added `ToolRenderer.provisionalPendingPreview` in `renderers.ts` and consulted it from `ToolExecutionComponent.isTranscriptBlockCommitStable`, replacing the 15.11.6 blanket gate that marked every collapsed pending preview commit-unstable.
- Marked only the tail-window previews the result render re-anchors as provisional: the edit streamed-diff tail (`edit/renderer.ts`), bash/ssh command caps (`createShellRenderer`, `sshToolRenderer`), and eval cells with interleaved outputs (`evalToolRenderer`).
- Restored mid-stream scrollback commits for every other pending preview, so tool calls taller than the viewport (e.g. a task call's context/assignment markdown) no longer read as cut off until the result lands.
- Added a regression test in `tool-live-region-scrollback.test.ts` scroll-appending a tall collapsed streaming task call into native scrollback mid-stream.
- Resolved retired effort-tier variant ids in `model-resolver.ts` through the hand-table aliases (`resolveVariantAlias`, `resolveBareVariantAlias`) plus the `X-thinking` → `X` grammar (`stripThinkingVariantToken`), with exact matches always winning while a raw id is live and explicit `:effort` suffixes transferring unchanged.
- Re-keyed models.yml `modelOverrides` and rate-limit selector suppressions from raw member ids onto the collapsed model in `model-registry.ts` (`normalizeSuppressedSelector`, lazy `hasLiveModel` checks so live raw ids keep their own overrides).
- Collapsed custom/config provider model lists at registry rebuild via `collapseBuiltModelVariants`, folding config-defined `X`/`X-thinking` twins into one entry.
- Extended `model-registry.test.ts` and `model-resolver.test.ts` with effort-tier variant collapsing and alias-resolution coverage.
- Added `SnapcompactShapePreview` (`snapcompact-shape-preview.ts`), rendering the highlighted shape variant over the sample text in `snapcompact-shape-preview-doc.md`; `auto` resolves to the active model's winning shape.
- Wired the preview as a footer component below the interactive rows in `settings-selector.ts`.
- Passed `modelApi`, `imageBudget`, and `requestRender` through `SelectorController` so the preview can rasterize against the session model and request repaints.
- Added snapcompact.shape setting with auto and variant options in agent config.
- Implemented resolveShape support for forced variants, auto provider winners, and repricing.
- Threaded resolved shape into compaction and inline-image flows for pricing and rendering.
- Updated fuzzy-search indexing to track compact word starts and required phrase matches to start at word boundaries.
- Adjusted token scoring and compact/phrase matching rules so exact and boundary-aware matches now rank above noisy substring matches.
- Filtered session search results to drop weak pure-fuzzy matches unless they contained literal tokens, and updated ranking tests for the new behavior.
- Added `handleUsageResetCommand` support to list and redeem usage reset credits.
- Refactored `/usage` into `show` and `reset` subcommands and removed `/reset-usage`.
- Handled ACP/TUI `/usage` flows so `show` reports usage and `reset` redeems credits.
- Kept selected theme setting values dirty-colored while selected in the UI list.
- Added `parentToolCallId` and `index` fields to subagent session records and propagated them from task lifecycle and progress events.
- Reworked subagent ordering to sort by parent-group order, spawn index, and stable creation order so out-of-order updates no longer reshuffled active HUD rows.
- Updated task execution to pass spawn indices through sync and async paths, switched the `tool.task` icon to Octicons tasklist, and added a registry test for out-of-order progress ordering.
- Added commit-stability signaling to transcript blocks and marked tool previews as unstable until expanded or finalized.
- Updated transcript scrollback promotion to derive live commit state only for blocks reporting commit-stable rows.
- Added tests ensuring provisional pending edit previews are never committed while durable live rows still promote after the stability window.
- Removed the `note` operation from `TodoOp` validation and the `text` field from todo operation payloads.
- Deleted `TodoItem` note storage and stripped note-aware logic from execution, markdown round-trip helpers, summaries, and rendering.
- Updated todo docs and prompt schema text to remove `note` usage and keep the public operation set aligned.
- Added AuthStorage.listResetCredits to query each stored Codex account from the reset-credits endpoint and return live availability plus active or error state.
- Exposed the new status fetch through AgentSession and switched reset-usage selectors and commands to consume it via toResetUsageAccounts.
- Updated reset-usage UI and slash-command output to show per-account errors and updated empty-state messaging when resets could not be loaded.
v15.11.4 introduced stateful previous_response_id chaining on the
official OpenAI endpoint. The in-provider retry classifier matched only
the generic stale-id phrasing ('previous response ... not found |
invalid | expired | stale'), missing the Zero Data Retention 400
'Previous response cannot be used for this organization due to Zero
Data Retention.'. The error therefore bypassed the categorical-disable
path, so the chain was reset (not disabled), the next successful turn
re-armed it, and every other turn 400'd in a loop.
Add a dedicated isOpenAIResponsesZeroDataRetentionError detector and a
markOpenAIResponsesChainZeroDataRetention helper that disables chaining
on the first hit (skipping the three-strike circuit breaker). The
in-call retry now drops 'store: true' from the replay so the request is
semantically valid for ZDR orgs, and reasoning continuity is preserved
by the existing include: ['reasoning.encrypted_content'] flag.
AgentSession.#isStaleOpenAIResponsesReplayError gains the ZDR phrasing
too, so any ZDR error that does bubble past the provider retry resets
the Responses session and retries at zero backoff instead of falling
back to a different model.
Fixes#2341
Rephrased and condensed prompts for the browser, eval, irc, and read tools. The updates focus on clarity, conciseness, and improved parseability for agent consumption by:
- Consolidating instructions and helper descriptions.
- Removing redundant phrasing and examples.
- Eliminating extraneous sections that added cognitive load.
- Added usage snapshot persistence in sqlite with hour-bucket upsert behavior.
- Added listUsageHistory query support with optional provider and sinceMs filters.
- Added usage CLI history mode with `--history` and `--days` and trend rendering.
- Added changelog documentation for usage trend inspection and no-history exit behavior.
- Added an anchored Subagents HUD renderer that formats active subagent sessions as `Id: description` rows.
- Integrated a new interactive-mode container and observer-driven render path so the HUD updates with session events and clears when no subagents are active.
- Exported `formatTaskId` for shared HUD formatting and added tests covering active-only filtering, fallback task/progress text, and line truncation.
- Updated task call and result headers to use the dispatch glyph during running async calls instead of spinner-style states.
- Switched running and pending agent rows to a static done-dot marker and reused the dot with foreground color settling when rows complete.
- Updated task rendering tests to validate the new glyphs and ensure running/pending rows do not emit spinner or pending symbols.
- Replaced global search query mutation logic with a dedicated `Input` instance so typing now supports cursor movement, word deletion, and other editor hotkeys.
- Updated search banner rendering to display the embedded input line with live cursor while preserving match-count formatting and prompt width handling.
- Extended `Input` with a configurable prompt and end-of-value cursor placement, then updated memory-search tests for the new hotkey-driven editor behavior.
- Updated the snapcompact system prompt note to state that rasterized tool results are provided in PNG frames below.
- Instructed the model that image delivery is deliberate harness behavior and should not be treated as a tool malfunction.
- Recorded the updated behavior in the changelog fixed notes for this release.
- Added explicit request-debug path helpers in `packages/ai/src/utils/request-debug.ts` for one-shot dumps.
- Added `/debug dump-next-request`, `/debug dump-request`, and `/debug next-request` subcommands to arm next-request dumps.
- Changed `debug` handling in `packages/coding-agent/src/slash-commands/builtin-registry.ts` to execute args instead of always opening selector.
- Fixed explicit request-debug mode to resolve `~`/relative paths and create parent directories before logging.
- Fixed one-shot request-debug mode to consume its target after one call and overwrite existing response logs.
- Added fuzzy search preprocessing to normalize case, camelCase, and punctuation.
- Added token and phrase scoring for exact, prefix, substring, acronym, and compact hits.
- Added fuzzyRank to return {item, score}-ranked matches and delegated fuzzyFilter to it.
- Updated settings search tabs to sort by best score, keep original-order ties, and mute unmatched tabs.
- Stored the active kitty keyboard enable sequence in the terminal and re-applied it after entering the alternate screen, then emitted `\x1b[<u` before leaving the alternate screen.
- Updated StdinBuffer parsing so `\x1b\x1b` waits for CSI/SS3 followers, splits `\x1b\x1b[` SGR mouse reports into a standalone Esc plus report, and reduced `PARTIAL_HOLD_MAX_MS` to 150ms.
- Detached async task progress rows stopped running a redraw driver and task progress rendering switched running/pending rows to static task-icon text.
- Background task snapshots were frozen once blocks left the transcript live region, preventing later partial snapshots from repainting commit-eligible rows.
- Updated task-progress and detached-background-task tests to validate static task rows and the new freeze behavior.
- Extracted `limitMatchesActiveAccount`/`reportMatchesActiveAccount` into `slash-commands/helpers/active-oauth-account.ts` as the single definition of the report-to-account matching rules, including projectId matching against `limit.scope.projectId`/metadata.
- Dropped the duplicated `ActiveAccountIdentity`/`OAuthAccessResolver` shims, `as unknown` session casts, the dead `getOAuthAccountId` fallback, and the email-vs-scope-accountId comparison from `command-controller.ts` and `usage-report.ts`.
- Replaced the async per-provider `resolveActiveAccountsForReports` map with one synchronous typed `authStorage.getOAuthAccountIdentity()` call per render, gated to the session's current provider.
- Re-exported `OAuthAccountIdentity` from `session/auth-storage.ts` and added `active-oauth-account.test.ts` covering the matching rules.
- Moved the /usage active-account entry from the released 15.9.4 section into `## [Unreleased]` under `### Changed`.
- Dropped the Ollama `ollama-chat` discovery entry that described code not present in this branch.
- Replaced the bare `catch {}` blocks in `runEmbedding()` (`beam/helpers.ts`), `getLocalModel()`, and the local-model path of `embed()` (`embeddings.ts`) with structured `logger.debug` entries carrying the error plus per-site context (item count, model name); failure semantics stay best-effort.
- Threaded `MnemopiOptions.debug` through `resolveRuntimeOptions()` in `memory.ts` so `mnemopiDebugEnabled()` escalates these logs to `warn` when `mnemopi.debug` is set.
- Passed `mnemopi.debug` from coding-agent settings into provider options (`debug` added to the `MnemopiProviderOptions` pick in `src/mnemopi/config.ts`).
- Added `embedding-failure-logging.test.ts` asserting the logged context on a real local-model load failure and the debug→warn escalation; the changelog entry landed with the previous commit (same contiguous run).
Fixes#2322: Mnemopi: runEmbedding() and getLocalModel() silently swallow errors — mnemopi.debug produces no diagnostic output
- Added `mnemopi.polyphonicRecall` and `mnemopi.enhancedRecall` boolean settings (default off) to `settings-schema.ts` and `loadMnemopiConfig`, applied via the new `configureRecallFeatures()` in `createScopedResources`.
- Made `polyphonicRecallEnabled()`, `enhancedRecallEnabled()`, and `isEnhancedRecallEnabled()` fall back to the configured defaults while `MNEMOPI_POLYPHONIC_RECALL` / `MNEMOPI_ENHANCED_RECALL` env vars still win when set.
- Exported `configureRecallFeatures`/`RecallFeatureFlags` from the package root and `core` barrels; documented the settings in `docs/mnemosyne-memory-backend.md`.
- Added `recall-feature-flags.test.ts` covering defaults, config enablement, and env precedence; the mnemopi changelog hunk also carries the adjacent #2322 entry (same contiguous run).
Fixes#2323: Mnemopi: MNEMOPI_POLYPHONIC_RECALL and MNEMOPI_ENHANCED_RECALL not configurable via config.yml
- Running job labels now shimmer in the TUI to provide a dynamic visual indicator of activity.
- Adjusted cache key to account for shimmer animation, ensuring it updates at 30fps instead of the 12.5fps spinner cadence.
- Suppressed job ID display when the job label is identical to its ID, avoiding redundant information.
- Ensured flattened chain rows under a `└─` branch are anchored by a vertical line one level right of the suppressed gutter (below the branch head's content), never in the `└─` corner column itself, resolving visual issues #2298 and #2325.
Fixes#2325.
- Removed obsolete compaction-threading audit test coverage.
- Updated issue-825-repro and skill-queue tests for object-based queued message text entries.
- Expanded queued-message tests to verify image restoration via clearQueue paths.
- Stored queued steering/follow-up attachments in `QueuedDisplayEntry` so restore APIs retain image data.
- Changed `AgentSession.clearQueue` and `popLastQueuedMessage` to return `{ text, images }` payloads.
- Restored queued text and images into editor image buffers when steer submission fails or queue items are restored.