- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
- Replaced `providers.parallelFetch` with a new `providers.fetch` enum in settings and added migration cleanup for the legacy key.
- Updated `renderHtmlToText` to follow configured reader preference with ordered fallback attempts and remote-reader timeout handling before local conversion.
- Updated YouTube and fetch tests to use `providers.fetch` and cover Jina-first stall fallback behavior.
- Lazy-loaded OTEL SDK, HTML export, TTSR, and autoresearch modules.
- Made resolveMemoryBackend async to import backends on demand.
- Replaced backend resolution with direct settings reads for rekey checks.
- Migrated Effort and THINKING_EFFORTS imports to @oh-my-pi/pi-ai/effort in CLI args and launch command files.
- Split model-registry dependencies across focused @oh-my-pi/pi-ai submodules instead of the root barrel export.
- Added a WeakMap cache in compileEquivalenceConfig to reuse compiled configs.
- Raised QUALIFIED_NAMESPACE_SUFFIX_CACHE_CAP and HEURISTIC_CANDIDATES_CACHE_CAP from 256 to 4096.
- Reworked resolveCanonicalIdForModel to build officialMatches during candidate iteration.
- Replaced WeakMap model cache with provider/id string keys for stable reuse.
- Returned official model ids directly when matched, before heuristics.
- Collapsed non-message token path to system prompt and tool schema totals.
- Added `TUI.resetDisplay()` to force an immediate full-frame replay including native scrollback.
- Moved the persistent model selector default from Ctrl+L to Alt+M, preserving existing user remaps.
- Reserved Alt+M so extensions cannot shadow the model selector shortcut.
- Replaced `¶path#hash` prefix with `[path#hash]` delimiters across parser, tokenizer, and grammar.
- Updated prompts, docs, and recovery paths to the new bracketed form.
The Ctrl+Q follow-up fallback should only be removed when a recognized
keybinding action claims that chord. Unknown or stale config entries are
ignored by the keybinding manager and should not suppress a working
follow-up default.
Addresses code review on #1905.
- Made "auto" the default, hiding MCP tools past 40-tool threshold.
- Centralized discovery mode resolution in shared mode helper.
- Activated search tool in createAgentSession once full registry exists.
- Added image-reference rendering to make `[Image #N]` placeholders clickable in chat.
- Added MIME-aware image blob materialization with extensioned sidecar paths.
- Added clickable path, line, and URL hyperlinks for read, search, and fetch outputs.
- Hardened OSC8 hyperlink emission with URI validation and control-byte/idempotency checks.
Add retry.modelFallback so users can keep automatic retry enabled while preventing retry recovery from switching through configured fallback model chains.
The default remains enabled, preserving existing fallback behavior. When disabled, retry still honors retry-after delays and retry limits while staying on the primary model.
Op: correct
Restores: spec:retry can stay enabled without automatic model switching
When a user had already bound ctrl+q to another action, the new
follow-up default still claimed the same chord and whichever handler was
registered last in InputController silently won. KeybindingsManager now
filters the new Ctrl+Q follow-up default when user config claims that
chord for another action, while preserving explicit follow-up remaps.
Addresses code review on #1905.
Windows Terminal does not deliver a distinct Ctrl+Enter event to console
apps, so the existing `app.message.followUp` default never fired there.
Ctrl+Q (the chord GitHub Copilot CLI uses for the same action) is added
as the primary default and works in every terminal we ship to; Ctrl+Enter
is kept as a secondary chord so users on Kitty/iTerm2/WezTerm/Ghostty
(where it does deliver) keep the existing muscle memory.
Fixes#1903
Mid-session edits to hindsight.bankId / bankIdPrefix / scoping kept the
active HindsightSessionState pinned to the bank selected at session
start, so retain/recall/reflect calls landed in the stale bank. Settings
hooks now fire onHindsightScopeChanged; the backend rebuilds the
primary state against the recomputed scope, disposing the previous one
after flushing its queue so queued tool-initiated retains still land in
the bank they were enqueued for.
Also:
- Renamed ensureBankMission to ensureBankExists. The old version
skipped creation entirely when bankMission was blank, so the first
mental-model POST (auto-seed) could land against a never-PUT bank.
Bank creation is now idempotent and unconditional, and runs before
mental-model bootstrap.
- Fixed AgentSession.dispose to flush the retain queue BEFORE clearing
the session state pointer. Reversed, HindsightRetainQueue.#doFlush's
identity guard would see the cleared pointer and drop the spliced
batch with a 'session vanished' warning.
- Snapshotted hindsightScopeCallbacks before iterating because each
rebuild subscribes a fresh callback inside the same fire; iterating
the live Set would spin.
Fixes#1902
Add `Model.omitMaxOutputTokens` (`models.yml` model definitions and
`modelOverrides` accept the same field). When set, the openai-responses
and openai-completions providers stop emitting `max_output_tokens` /
`max_tokens` / `max_completion_tokens` so the upstream API applies its
own default cap. Catalog `maxTokens` is still honoured for local
budgeting (compaction, context promotion); only the wire field is
suppressed.
Restores the pre-v15.8.3 behaviour for Ollama proxies fronting cloud
catalogs (GLM, Kimi, DeepSeek): OMP cannot discover their true output
limit, so users previously set `maxTokens` to the Ollama context window
to unlock the model. v15.8.3 began sending that value as
`max_output_tokens`, triggering HTTP 400 from the upstream provider.
`applyCommonResponsesSamplingParams` now takes the model object instead
of a bare provider string so the wire-suppression flag is available to
the sampling-params builder.
Fixes#1881
Centralized append-only context auto-mode resolution so SDK sessions, interactive sessions, and the status display use the same provider checks. Auto mode now enables for Xiaomi/SGLang hosts and explicit stored-request compat signals while preserving DeepSeek behavior.
Added focused regression coverage for Xiaomi Token Plan SGLang HiCache endpoints and explicit on/off behavior.
Fixes#1851
Allowed OpenAI-compatible providers to opt into Anthropic cache_control markers via compat.cacheControlFormat while preserving OpenRouter Anthropic defaults.
Fixes#1845
- Renamed `TodoWriteTool` to `TodoTool` and its source/prompt files.
- Updated tool registration, schema, renderers, and gating to `todo`.
- Adjusted cursor provider native tool names and tests to match.
- Renamed strike-animation constants and `todo-error-reminder` type.
- Added `textSizing` terminal capability and `tui.textSizing` setting for Kitty OSC66 scaling.
- Updated OSC66 range and visible-width handling to preserve grapheme slice behavior.
- Updated markdown rendering to apply heading text sizing only when enabled.
- Fixed DECCARA trailing-background fill tracking and DECRPM status 3/4 recognition.
- Added kitty-graphics module with Unicode placeholder and temp-file APIs.
- Refined terminal image handling to use rendered placeholder lines and detect placeholder output.
- Raised the max inline-image default from 3 to 8 in schema and image constants.
- Refined terminal capability flow for Kitty support, including screen-to-scrollback and text-sizing handling.
- Added integration tests for kitty placeholders, virtual placement, and temp-file rendering.
- Added `ImageBudget` keeping only N recent images live, demoting older ones to text via a full redraw plus Kitty graphics purge.
- Changed Kitty images to transmit-once + placement, so repaints re-send only tiny placement sequences instead of base64.
- Added `tui.maxInlineImages` setting (default 3, 0 disables) wired through the coding-agent renderers.
Loaded cached startup models for special built-in providers alongside standard provider descriptors so boot-time model resolution can see cached Google Antigravity, Gemini CLI, and OpenAI Codex discoveries before refresh.\n\nFixes #1721
Kept Ctrl+V as a default clipboard-image paste shortcut on Windows while preserving Alt+V as the Windows Terminal-safe fallback.
Updated keybinding docs and regression coverage for the platform-specific default.
Fixes#1708
Replace the sunset V0 Search API with V1 (POST /api/v1/search) under the
existing `kagi` provider id instead of shipping a parallel `kagi-v1`
provider. Credentials still resolve through the shared AuthStorage broker
(Bearer token, KAGI_API_KEY, /login kagi), and recency now maps to a
UTC-deterministic filters.after date.
- Merge V1 client into src/web/kagi.ts (categorized result buckets, direct
answer, related/adjacent questions)
- Keep classifyProviderHttpError mapping for auth/quota signals
- Drop the kagi-v1 entries from the provider registry, order, type union,
and settings schema
- Consolidate tests into web-search-kagi.test.ts
- Added `Settings.reloadForCwd` to mutate the live instance in place, so `/move` and cross-project resume pick up the destination project's `.claude/settings.yml` and path-scoped `enabledModels`/`disabledProviders`.
- Wired `reloadForCwd` into `applyCwdChange` (interactive mode) and the `--resume` startup path so settings always follow the active working directory.
- Added tests covering path-scoped re-resolution, no-op on same directory, and disk-backed project layer load/drop.
- Added `DiagnosticsLedger` to track diagnostics already surfaced per file, suppressing repeats within a session.
- Wired dedup into both edit and write tools via a `transformDiagnostics` hook on the writethrough pipeline.
- Added `lsp.diagnosticsDeduplicate` setting (default: true) to control the behavior.
- Add new keybinding `app.clipboard.pasteTextRaw` with Ctrl+Shift+V / Alt+Shift+V defaults
- Implement `readTextFromClipboard()` utility supporting macOS, Windows, Linux (Wayland/X11), and Termux
- Integrate raw paste handler into CustomEditor and InputController
- Text is inserted without formatting collapse to preserve whitespace and newlines
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
- Removed `ModelRegistry.create` and updated construction sites to use `new ModelRegistry` directly.
- Dropped the async `ConfigFile` migration warmup path and switched migration handling to the unified `#ensureMigrated` logic used by relocate and load flows.
- Adjusted model registry tests to instantiate `ModelRegistry` through the constructor.
F6: defer the JSON -> YAML migration out of ConfigFile's constructor and
add an async path so the boot sequence stops blocking the event loop on
sync I/O. New ConfigFile.tryLoadAsync/loadAsync/loadOrDefaultAsync/
getMtimeMsAsync and static ConfigFile.warmup. The migration is now
idempotent (per-process cache) so relocate() does not re-run it.
ModelRegistry.create(authStorage, modelsPath?) is a new async factory
that runs the warmup before the sync constructor's bundled-model load.
Production call sites (main.ts, sdk.ts, task/executor.ts, commit
pipelines, SDK example) all switched. Sync new ModelRegistry(...)
constructor is still supported for tests.
F2: rewrite MemorySessionStorage's mirror as { chunks: string[]; byteLen;
mtimeMs } so writeLineSync appends a single chunk in O(1) instead of
read-modify-writing the entire file (which was O(N) per append, O(N^2)
per session). statSync now reports true UTF-8 byte length instead of
character count. readTextPrefix walks chunks until the byte budget is
exhausted instead of materialising the full mirror.
Also rolls in per-package CHANGELOG entries for F1-F8.
F1 (coding-agent): close the withFileLock mkdir-vs-writeLockInfo race that
let a losing contender wipe the winner's freshly-created lock directory.
Every lock now carries a per-process UUID token; releaseLock verifies the
token before fs.rm, and isLockStale no longer treats an info-less but
fresh dir (or a dir that vanished mid-check) as stale.
F4 (coding-agent): sanitize tabs and truncate oversized error strings in
formatErrorMessage so error renderings that embed file content
(apply_patch, hashline, etc.) cannot break terminal alignment or
overflow the line width.
F7 (ai): support named-tool routing on Google providers. Widens
GoogleSharedStreamOptions.toolChoice and GoogleGeminiCliOptions.toolChoice
to accept { mode: 'ANY'; allowedFunctionNames }. mapGoogleToolChoice
now converts ToolChoice { type: 'tool'|'function', name } to the wire
shape (mirroring mapAnthropicToolChoice). buildGoogleGenerateContentParams
and the gemini-cli request serializer honor the allow-list.