Merge remote-tracking branch 'upstream/main' into feat/error-notify
# Conflicts: # packages/coding-agent/test/agent-session-retry-diagnostics.test.ts
This commit is contained in:
@@ -4,15 +4,15 @@
|
||||
|
||||
### Added
|
||||
|
||||
- Added Anthropic fallback content block support in agent-loop assistant-message snapshotting so the block round-trips through session persistence and IRC event fanout unchanged. ([#4177](https://github.com/can1357/oh-my-pi/issues/4177))
|
||||
- Added support for Anthropic fallback content blocks in agent-loop assistant messages, ensuring they are preserved across session persistence and event fanout.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed an issue where skipped tool results in queued messages were incorrectly treated as completed work, preventing necessary retries.
|
||||
- Fixed an issue where skipped tool results in queued messages were incorrectly treated as completed, preventing necessary retries.
|
||||
- Improved branch summaries to preserve informative tool results from abandoned branches while filtering out redundant output.
|
||||
- Fixed interruptible tool waits to properly abort on host-provided IRC interrupts in addition to user steering.
|
||||
- Fixed schema validation errors for closed union tools by correctly injecting intent tracing into each variant.
|
||||
- Fixed compaction reserve-budget provenance: an explicit `reserveTokens` equal to the built-in default is now honored instead of being replaced by the proportional small-window fallback, and the fallback reserve is clamped to at least one token so tiny context windows keep a threshold below the window size.
|
||||
- Fixed token compaction reserve-budget logic to honor explicit reserveTokens values equal to the built-in default, and clamped the fallback reserve to at least one token for very small context windows.
|
||||
|
||||
## [16.2.4] - 2026-06-28
|
||||
|
||||
|
||||
@@ -4,19 +4,19 @@
|
||||
|
||||
### Added
|
||||
|
||||
- Added opt-in support for Anthropic's server-side fallback beta chain (`server-side-fallback-2026-06-01`) on the `anthropic-messages` provider. When `AnthropicOptions.fallbacks` is set, the request carries the `fallbacks` field and the beta header, and the response parser honors mid-stream `fallback` content blocks and `usage.iterations` — promoting the served model on a `fallback_message` iteration and pricing per-attempt at the served model's cache-read rate for the fallback attempt's input tokens (per the [fallback billing cookbook §4](https://platform.claude.com/cookbook/fable-5-fallback-billing-guide)). Non-Anthropic providers and non-opted-in requests are fully inert. `transformMessages` centrally drops persisted `fallback` blocks on cross-provider hops and non-official Anthropic replays so a stored fallback turn never wedges downstream converters. `SimpleStreamOptions.fallbacks` and `AssistantMessage.content` now include the fallback surface. ([#4177](https://github.com/can1357/oh-my-pi/issues/4177))
|
||||
- Added opt-in support for Anthropic's server-side fallback beta (`server-side-fallback-2026-06-01`) on the `anthropic-messages` provider, including support for `AnthropicOptions.fallbacks`, mid-stream fallback content blocks, fallback billing/usage iterations, and automatic filtering of fallback blocks during cross-provider message transformations.
|
||||
|
||||
### Changed
|
||||
|
||||
- Clarified CoreWeave Serverless Inference login instructions to persist `COREWEAVE_PROJECT` in the user's shell startup file.
|
||||
- Updated CoreWeave Serverless Inference login instructions to clarify persisting `COREWEAVE_PROJECT` in shell startup files.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed an issue where broker usage fetch failures were not cached, causing sequential ranking passes to repeatedly hit the broker when it is down.
|
||||
- Fixed an issue where broker usage fetch failures were not cached, causing redundant network requests during sequential ranking passes when the broker is offline.
|
||||
- Fixed Xiaomi MiMo API key validation to use the supported `mimo-v2.5` model.
|
||||
- Fixed certificate verification errors for custom gateways behind private CA bundles by applying `NODE_EXTRA_CA_CERTS` to all provider fetches (including OpenAI-compatible, Codex, Ollama, Azure, and Google).
|
||||
- Fixed Claude Fable demoted-thinking replay to use markdown-italic assistant prose instead of `<thinking>` tags, avoiding reasoning-extraction-shaped context after model switches.
|
||||
- Fixed OpenAI Responses replay emitting locally rebuilt assistant item IDs without their required reasoning items, preventing `function_call` / `message` replay 400s from poisoned history. ([#4173](https://github.com/can1357/oh-my-pi/issues/4173))
|
||||
- Fixed certificate verification errors for custom gateways behind private CA bundles by ensuring `NODE_EXTRA_CA_CERTS` is respected across all provider fetches (including OpenAI-compatible, Codex, Ollama, Azure, and Google).
|
||||
- Fixed Claude Fable demoted-thinking replay to use markdown-italic assistant prose instead of `<thinking>` tags, preventing context issues after model switches.
|
||||
- Fixed OpenAI Responses replay errors (400 Bad Request) caused by missing reasoning items in locally rebuilt assistant item IDs during history replay.
|
||||
|
||||
## [16.2.13] - 2026-07-01
|
||||
|
||||
|
||||
@@ -4,16 +4,14 @@
|
||||
|
||||
### Breaking Changes
|
||||
|
||||
- Renamed the `requiresJuiceZeroHack` compat flag to `requiresReasoningSuppressionPrompt` (in `OpenAICompat` and `ResolvedOpenAIResponsesCompat`) and dropped the unused `"juice-zero-developer-message"` member from `OpenAIReasoningDisableMode`: the `# Juice: 0` wire hack is gone; the flag now gates the no-reasoning suppression prompt.
|
||||
- Renamed the `requiresJuiceZeroHack` compatibility flag to `requiresReasoningSuppressionPrompt` (affecting `OpenAICompat` and `ResolvedOpenAIResponsesCompat`) and removed the unused `"juice-zero-developer-message"` member from `OpenAIReasoningDisableMode`.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed the Xiaomi provider's default model to use the supported mimo-v2.5 model.
|
||||
- Fixed model discovery (`/models` probes) failing behind private-CA gateways: the discovery fetch fallback now honors `NODE_EXTRA_CA_CERTS`, matching provider chat requests. The models.dev metadata fetch and Ollama native `/api/tags`+`/api/show` probes, missed by the original sweep, now use the same wrapped fetch.
|
||||
- Fixed CoreWeave Serverless Inference project-header detection to ensure blank OpenAI-Project overrides do not block the COREWEAVE_PROJECT fallback.
|
||||
- Fixed LiteLLM MiniMax M3 discovery to remove reseller-only (3x usage) display suffixes.
|
||||
- Fixed users with a warm LiteLLM model cache keeping stale reseller display-name suffixes for up to 24 hours by bumping the cache namespace to rich-v2.
|
||||
- Updated the Responses compatibility flag documentation to clarify the no-reasoning fallback behavior.
|
||||
- Fixed the Xiaomi provider's default model to use the supported `mimo-v2.5` model.
|
||||
- Fixed model discovery probes (including Ollama and metadata fetches) failing behind private-CA gateways by ensuring they honor `NODE_EXTRA_CA_CERTS`.
|
||||
- Fixed CoreWeave Serverless Inference project-header detection to ensure blank OpenAI-Project overrides do not block the `COREWEAVE_PROJECT` fallback.
|
||||
- Fixed LiteLLM MiniMax M3 discovery to remove reseller-only display suffixes, and invalidated the model cache to ensure stale suffixes are cleared immediately.
|
||||
|
||||
## [16.2.13] - 2026-07-01
|
||||
|
||||
@@ -197,7 +195,7 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed Claude 4.6 routing on the `google-antigravity` (and `google-gemini-cli`) Cloud Code Assist providers, whose backend exposes the models asymmetrically: `claude-sonnet-4-6` has no `-thinking` twin and `claude-opus-4-6` has only the `-thinking` twin. The shared `thinkingPair` family was routing thinking efforts on `claude-sonnet-4-6` to a non-existent `claude-sonnet-4-6-thinking` wire id (404 `Requested entity was not found`); replaced both 4.6 entries with bespoke single-wire families that declare the dead ids as `retiredMembers` so `reconcileRetiredRouting` re-points stale bundled-catalog and SQLite-cache rows away from the 404 wire id. Refreshed the bundled `models.json` Sonnet 4.6 entry whose stored `effortRouting` still targeted the dead `-thinking` id. Added `claude-sonnet-4-6` and `claude-opus-4-6-thinking` entries to `ANTIGRAVITY_MODEL_WIRE_PROFILES` capped at the backend's 64000-output-token limit (over-cap requests 400'd with `Request contains an invalid argument`); `modelEnum` is now optional on `AntigravityModelWireProfile` since the Claude wire ids are accepted without a captured `labels.model_enum`. ([#3067](https://github.com/can1357/oh-my-pi/issues/3067))
|
||||
- Fixed Claude 4.6 routing on the `google-antigravity` (and `google-gemini-cli`) Cloud Code Assist providers, whose backend exposes the models asymmetrically: `claude-sonnet-4-6` has no `-thinking` twin and `claude-opus-4-6` has only the `-thinking` twin. The shared `thinkingPair` family was routing thinking efforts on `claude-sonnet-4-6` to a non-existent `claude-sonnet-4-6-thinking` wire id (404 `Requested entity was not found`); replaced both 4.6 entries with bespoke single-wire families so every effort and off resolve to the live wire id. Added `claude-sonnet-4-6` and `claude-opus-4-6-thinking` entries to `ANTIGRAVITY_MODEL_WIRE_PROFILES` capped at the backend's 64000-output-token limit (over-cap requests 400'd with `Request contains an invalid argument`); `modelEnum` is now optional on `AntigravityModelWireProfile` since the Claude wire ids are accepted without a captured `labels.model_enum`. ([#3067](https://github.com/can1357/oh-my-pi/issues/3067))
|
||||
|
||||
## [16.1.3] - 2026-06-19
|
||||
|
||||
|
||||
@@ -10,84 +10,32 @@
|
||||
|
||||
- Added `providers.anthropic.serverSideFallback` (default off; UI in the "Model → Retry & Fallback" group). When enabled, Claude Fable 5 / Mythos 5 requests carry `fallbacks: [{ model: "claude-opus-4-8" }]` via Anthropic's server-side-fallback beta chain so classifier-blocked turns are retried on Opus 4.8 without breaking the current call. Opt-in only — leaving it off preserves the pre-fallback behavior. ([#4177](https://github.com/can1357/oh-my-pi/issues/4177))
|
||||
- Added `task.softRequestBudgetNotice` (default off) to opt into the subagent soft-budget wrap-up steering notice while keeping the 1.5x graceful abort guard active.
|
||||
- Added `providers.anthropic.serverSideFallback` configuration option to opt into Anthropic's server-side-fallback beta chain, allowing Claude Fable 5 / Mythos 5 requests to automatically retry on Opus 4.8 when blocked by classifiers.
|
||||
- Added `task.softRequestBudgetNotice` configuration option to enable subagent soft-budget wrap-up steering notices while keeping the graceful abort guard active.
|
||||
|
||||
### Changed
|
||||
|
||||
- Updated the tester subagent prompt to explicitly forbid testing default configuration values and allow skipping tests entirely for trivial changes
|
||||
- Optimized session loading and rendering performance, including a 10x speedup for smooth streaming reveals on large messages, 35% faster session resumes for large files using native streaming JSONL parsing, and reduced overhead for edit-patch fallbacks.
|
||||
- Significantly optimized session loading and rendering performance, including a 10x speedup for streaming reveals on large messages, 35% faster session resumes for large files using native streaming JSONL parsing, and reduced overhead for edit-patch fallbacks.
|
||||
- Improved TUI responsiveness and reduced CPU usage during long-running tool sessions by throttling status-line redraws and optimizing subagent persistence checks.
|
||||
- Updated the advisor system prompt and documentation to accurately reflect WATCHDOG.yml tool grants.
|
||||
- Updated the tester subagent prompt to allow skipping tests for trivial changes.
|
||||
- Improved DuckDuckGo web search error clarity and documented datacenter/shared-egress limitations in provider settings.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed sequential patch/edit failures leaving dirty buffers in multi-file edits by flushing the LSP writethrough queue upfront on early abort paths
|
||||
- Fixed the subagent task preview erroneously showing "ctrl+o: Expand" hints on output lines capped inside the already-expanded view
|
||||
|
||||
- Fixed the cross-file write batching regression in `apply_patch` multi-file operations by reverting to flushing only on the last file write or explicitly on early failure paths, and refactored error counting to use clean booleans.
|
||||
- Fixed collab teardown (`/collab stop`, unrecoverable relay drop) cancelling a hook selector/editor the host user was actively typing in: teardown resolved pending guest asks the same way as a guest cancel, so the mirrored race dismissed the local dialog and dropped its input. Teardown now settles guest asks as `unavailable`, and the local dialog keeps running and wins with its eventual answer.
|
||||
- Fixed RPC mode deferred shutdown (`pi.shutdown()`) killing the process while a background-dispatched `bash` command was still running: the response frame is now written before exit, and a shutdown requested mid-bash fires once the command settles even when no further client frames arrive.
|
||||
- Fixed collab-guest transcript viewer rendering host-delivered errors raw: multi-line stacks broke the frame's row accounting and absolute host paths leaked to guests; errors are now collapsed to one sanitized, truncated row.
|
||||
- Fixed streaming tool-arg previews capturing string fields (e.g. `content`) from nested objects and injecting them as top-level args mid-stream; only top-level keys are read incrementally now.
|
||||
- Fixed task.maxConcurrency being breachable when a queued spawn was cancelled: the spawn path could release a semaphore permit it never acquired, letting a later task start while the cap was saturated.
|
||||
- Fixed session exit diagnostics recording signal and crash exits (SIGTERM, SIGHUP, uncaught exceptions) as a normal "dispose": the postmortem teardown now threads the real reason into session disposal.
|
||||
- Fixed the subagent yield-label guard ignoring JTD discriminator (oneOf) output schemas, which let stale incremental labels pass into successful results when final validation was skipped after retries.
|
||||
- Fixed grep/ast_grep search scopes rejecting `www.` and collapsed-scheme (`https:/host`) URL spellings that the read tool accepts; unresolvable URL-shaped scopes now fail with an explicit external-URL error instead of "Path not found".
|
||||
- Fixed model discovery ignoring `NODE_EXTRA_CA_CERTS`: the model registry's default fetch now applies the extra-CA wrapper, so `/models` probes work behind private-CA gateways like provider chat requests.
|
||||
- Fixed ctrl+p role-model cycling getting stuck on one transition and skipping every other role: a session-branch traversal regression returned entries leaf-to-root, so the cycle (and session model restore) read the oldest recorded model change instead of the newest.
|
||||
- Fixed ctrl+p cycling from a stale slot after the model was switched through another surface (alt+m, /model, retry fallback): the recorded role is now trusted only while its resolved model is still the active model, falling back to matching by model.
|
||||
- Fixed the `apply_patch` envelope to reject `*** Add File` / `*** Move to` targeting a pre-existing file upfront instead of silently overwriting it. The JSON `patch` mode's `op: "create"` intentionally remains the documented full-file overwrite (rename stays non-overwriting in both modes).
|
||||
- Fixed multi-file apply_patch to stop at the first failing file, surface applied vs. skipped paths, and correctly report the error to the agent loop.
|
||||
- Fixed process termination (SIGTERM, SIGHUP, uncaught exceptions) skipping editor draft saves, session shutdown events, and background job cleanup.
|
||||
- Fixed /quit and /exit commands blocking session closure by introducing a shutdown budget and backgrounding remaining tasks.
|
||||
- Fixed git and GitHub CLI subprocesses hanging on interactive prompts by forcing non-interactive environments, adding timeouts, and capping output.
|
||||
- Fixed RPC mode abort_bash being blocked by running bash commands by dispatching bash in the background.
|
||||
- Fixed task.maxConcurrency and task.maxRecursionDepth limits being bypassed by sub-spawn paths, ensuring limits are dynamically resized and respected.
|
||||
- Fixed the edit tool inflating session files by pruning extremely large file snapshots from tool-result details.
|
||||
- Fixed edit-tool Markdown list guidance so hashline parser errors and the model-facing prompt teach `+- item` escaping instead of steering agents toward full-file `write` fallbacks. ([#4179](https://github.com/can1357/oh-my-pi/issues/4179))
|
||||
- Fixed workstation OS detection rendering "Kernel: unknown" on macOS 15+.
|
||||
- Fixed /copy code and /copy cmd commands being treated as normal prompts instead of copying the requested blocks.
|
||||
- Fixed interactive bash status line not updating after directory changes (cd).
|
||||
- Fixed session title refreshes ignoring user TITLE_SYSTEM.md overrides during replans, and prevented auto-generated titles from incorrectly preserving all-caps text from user messages.
|
||||
- Fixed the live todo HUD going stale during long tool-use loops by adding mid-run reminders for incomplete items.
|
||||
- Fixed /shake and mid-stream chat rebuilds erasing active LLM output.
|
||||
- Fixed RpcClient failing to restart after being stopped or failing on initial startup.
|
||||
- Fixed /collab web guests being unable to answer ask tool questions by routing host UI requests through writable collab peers.
|
||||
- Fixed search and AST tools accepting external read URLs by materializing fetched URL text through the read cache before path resolution.
|
||||
- Fixed Tavily web search to retry without recency filters if no content is returned.
|
||||
- Fixed extension validation failures for omp install pi-lean-ctx by exposing legacy tool factories.
|
||||
- Fixed legacy `createReadTool`/`createReadToolDefinition` ignoring `autoResizeImages`; the option now maps onto the `images.autoResize` setting of the underlying read tool.
|
||||
- Fixed legacy `createGrepTool`/`createGrepToolDefinition` silently ignoring options: an explicit `context` parameter is now forwarded to the built-in grep (as symmetric before/after context), the unsupported `limit` parameter is no longer advertised in the tool schema, and supplying legacy `operations` throws a descriptive error at creation time instead of silently searching the local filesystem.
|
||||
- Fixed visibility of the focused option in the multi-select ask picker on certain color themes.
|
||||
- Fixed TUI row overlapping and duplication in the eval tool's live subagent progress tree under heavy concurrency.
|
||||
- Fixed session resumes after silent exits by recording pre-tool start markers and shutdown diagnostics.
|
||||
- Fixed status-line redraw crashes when tool-call arguments contain BigInt values.
|
||||
- Fixed terminal scrollback rows retaining old colors after theme switches.
|
||||
- Fixed Esc key behavior in the TUI to clear unrecoverable input instead of preserving drafts.
|
||||
- Fixed marketplace-installed plugins appearing redundantly in both the npm list and the extension status provider.
|
||||
- Fixed parent and peer IRC message delivery delays.
|
||||
- Fixed provider/model:auto entries in modelRoles collapsing to inherit and losing their auto state on reload.
|
||||
- Fixed plan execution prompts to avoid embedding the full plan, referencing the local plan file instead.
|
||||
- Fixed subagent live progress leaking raw tool output into the parent TUI.
|
||||
- Fixed eval subagents with custom output schemas receiving stale incremental yield labels.
|
||||
- Fixed hidden-thinking live status rows rendering as glyph-only lines by adding a persistent label.
|
||||
- Fixed /compact summary divider placement to keep it in the live scrollable region.
|
||||
- Fixed the time_spent status-line segment ticking continuously during idle sessions.
|
||||
- Fixed browser tool schema validation to require the code argument for run calls.
|
||||
- Improved robustness of MCP authentication error detection and header-based server discovery.
|
||||
- Fixed reliable detection of 401/403 authorization failures during Smithery commands and HTTP RPCs.
|
||||
- Improved streaming preview responsiveness for write, edit, and eval tools by decoding streamed string arguments incrementally.
|
||||
- Added retry-path diagnostics for assistant-tail removal and scheduled continuations after transient provider errors.
|
||||
- Fixed CJK history rendering issues across repeated compactions.
|
||||
- Fixed user-invoked skills failing to identify themselves or resolve relative paths across various execution paths.
|
||||
- Fixed type errors introduced by the merge sweep: restored the ask row-budget priority field, narrowed dereferenced schema property access, and updated stale test API usage.
|
||||
- Fixed git clone and fetch being killed by the 5-minute local-command timeout; network transfers now use a separate 30-minute deadline, overridable per call.
|
||||
- Fixed the TUI collab guest (omp join) silently dropping host ask/selector UI requests; they now present through the standard dialog flow and round-trip responses, with cancellation and resync replay handled.
|
||||
- Fixed transcript rebuilds (theme change, /shake, focus replay) showing stale streamed write/edit/eval content by sharing the partial-JSON decode between the live streaming path and every rebuild path.
|
||||
- Fixed an explicitly configured compaction.reserveTokens equal to the built-in default being silently replaced by the proportional small-window fallback; the setting now defaults to unset and explicit values are always honored.
|
||||
- Fixed user-configured LiteLLM discovery providers keeping stale reseller display-name suffixes for up to 24 hours after upgrade by invalidating the warm model cache.
|
||||
- Fixed `mergeTaskBranches` and `applyNestedPatches` leaving stage 1/2/3 unmerged entries in `.git/index` when the post-merge stash pop conflicted with the cherry-picked HEAD. The corrupted index survived indefinitely and every subsequent overlay-isolated task inherited it through the lower layer, causing `captureRepoDeltaPatch` to emit `diff --cc` output that `git apply` rejects with "No valid patches in input". The stash restore now runs behind a 3-way preflight check (`git apply --3way --check`) and a `reset --hard HEAD` cleanup fallback; the stash entry is preserved for manual recovery on conflict, and the merged commits still land on HEAD. ([#4175](https://github.com/can1357/oh-my-pi/issues/4175))
|
||||
- Fixed macOS `Command+V` image pastes in Ghostty by binding the Kitty `super+v` key event to the image-paste action alongside `Ctrl+V`. ([#4178](https://github.com/can1357/oh-my-pi/issues/4178))
|
||||
- Fixed several issues with the `apply_patch` and edit tools, including preventing dirty buffers on early aborts, rejecting overwrites of pre-existing files, stopping at the first failing file in multi-file operations, and pruning extremely large file snapshots to prevent session inflation.
|
||||
- Fixed process termination (SIGTERM, SIGHUP, uncaught exceptions) to ensure editor drafts are saved, sessions shut down cleanly, and background jobs are cleaned up.
|
||||
- Fixed `/quit` and `/exit` commands blocking session closure by introducing a shutdown budget and backgrounding remaining tasks.
|
||||
- Fixed git clone and fetch operations being killed prematurely by applying a separate 30-minute deadline for network transfers instead of the standard 5-minute local-command timeout.
|
||||
- Fixed git index corruption issues in `mergeTaskBranches` and `applyNestedPatches` when post-merge stash pops conflicted with the cherry-picked HEAD.
|
||||
- Fixed collaboration (`/collab`) issues, including preventing teardowns from cancelling active host dialogs, sanitizing host-delivered errors for guests, and ensuring guest UI requests are correctly routed and rendered.
|
||||
- Fixed RPC mode deferred shutdown (`pi.shutdown()`) and `abort_bash` commands from being blocked by active background bash processes.
|
||||
- Fixed model discovery ignoring `NODE_EXTRA_CA_CERTS` and resolved a rare Bun garbage collection segfault during discovery.
|
||||
- Fixed `task.maxConcurrency` and `task.maxRecursionDepth` limits being bypassed by sub-spawn paths or when queued spawns were cancelled.
|
||||
- Fixed various TUI and rendering issues, including macOS Ghostty `Command+V` image pasting, CJK history rendering across compactions, status-line redraw crashes with BigInt values, and overlapping rows in the subagent progress tree.
|
||||
- Fixed `/copy code` and `/copy cmd` commands being treated as normal prompts instead of copying the requested blocks.
|
||||
- Fixed legacy tool compatibility issues for `createReadTool`, `createGrepTool`, and extension validation failures for `omp install pi-lean-ctx`.
|
||||
- Improved robustness of MCP authentication error detection, Smithery command authorization failures, and streaming preview responsiveness for write, edit, and eval tools.
|
||||
- Fixed llama.cpp router/preset mode reporting 128k context in the status bar for every preset picked from `/model` regardless of the preset's configured `--ctx-size`. The router-level `/v1/models` only carries `meta.n_ctx` after a preset's child instance is loaded, and router `/props` reports a dummy `n_ctx: 0`, so cold-picked presets fell through to the 128k discovery default. Discovery now also reads `--ctx-size` (or `-c`) from each entry's `status.args` rendered CLI vector and, as a fallback, `ctx-size = N` from `status.preset` INI; the same fallback applies on the runtime refresh that runs when `/model` switches to a preset ([#4190](https://github.com/can1357/oh-my-pi/issues/4190)).
|
||||
|
||||
## [16.2.13] - 2026-07-01
|
||||
|
||||
@@ -483,7 +431,7 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed Ctrl+Z hanging the terminal after any tool call had run: the TUI tore down (`ui.stop()`) but the process kept running in `Sl+` state, leaving the user with a dead terminal recoverable only via `kill -9`. The embedded `brush-core` shell behind every bash tool call installs a tokio SIGTSTP listener on `Process::wait` (`crates/vendor/brush-core/src/sys/unix/signal.rs::tstp_signal_listener` → `tokio::signal::unix::signal(SIGTSTP)`); per tokio's contract, the first call for a SignalKind permanently replaces the kernel-default handler for the lifetime of the process. So the first bash invocation — even `/usr/bin/true` — silently overrode SIGTSTP's "stop" default, and `InputController.handleCtrlZ`'s subsequent `process.kill(0, "SIGTSTP")` was swallowed by tokio. The handler now sends `SIGSTOP` (uncatchable, unblockable, unignorable) to the foreground process group, so the kernel parks omp regardless of installed handlers and the shell sees the whole job stop even when omp runs behind a wrapper (`npx`, `pnpm exec`, `bunx`, …) or as one stage of a pipeline. MCP stdio servers now spawn detached into their own session — they're insulated both from terminal job-control signals (which used to stop their process trees and leave the JSONL read loop blocked on silent pipes) and from the new pgid=0 suspend itself ([#3461](https://github.com/can1357/oh-my-pi/issues/3461)).
|
||||
- Fixed Ctrl+Z hanging the terminal after any tool call had run: the TUI tore down (`ui.stop()`) but the process kept running in `Sl+` state, leaving the user with a dead terminal recoverable only via `kill -9`. The embedded `brush-core` shell behind every bash tool call installs a tokio SIGTSTP listener on `Process::wait` (`crates/brush-core-vendored/src/sys/unix/signal.rs::tstp_signal_listener` → `tokio::signal::unix::signal(SIGTSTP)`); per tokio's contract, the first call for a SignalKind permanently replaces the kernel-default handler for the lifetime of the process. So the first bash invocation — even `/usr/bin/true` — silently overrode SIGTSTP's "stop" default, and `InputController.handleCtrlZ`'s subsequent `process.kill(0, "SIGTSTP")` was swallowed by tokio. The handler now sends `SIGSTOP` (uncatchable, unblockable, unignorable) to the foreground process group, so the kernel parks omp regardless of installed handlers and the shell sees the whole job stop even when omp runs behind a wrapper (`npx`, `pnpm exec`, `bunx`, …) or as one stage of a pipeline. MCP stdio servers now spawn detached into their own session — they're insulated both from terminal job-control signals (which used to stop their process trees and leave the JSONL read loop blocked on silent pipes) and from the new pgid=0 suspend itself ([#3461](https://github.com/can1357/oh-my-pi/issues/3461)).
|
||||
- Fixed image-only composer submissions while the agent is streaming being treated as empty input, which dropped the image or aborted the active turn when another message was queued. Pending pasted images now count as submit content for Enter and Ctrl+Enter follow-ups. ([#3467](https://github.com/can1357/oh-my-pi/issues/3467))
|
||||
- Fixed `omp gallery --state` accepting lifecycle tokens that did not match displayed state labels and rendering unknown state values as `· undefined`; displayed labels now work as aliases, invalid values fail with a valid-token list, and failed gallery fixtures visibly render failures. ([#3473](https://github.com/can1357/oh-my-pi/issues/3473))
|
||||
- Fixed the bash tool's snapshotted `mise()` shell function dying with `command: command not found:` because `$__MISE_EXE` was empty in the replay shell. `generateSnapshotScript` captured the function via `declare -f`/`typeset -f` but only ever re-exported `PATH`, so every other env var the rc file set (notably the `*_EXE` sidecar `mise activate` exports) was lost; the function body then expanded `command "$__MISE_EXE" "$@"` to `command "" …` and died with exit 127. The snapshot script now scans captured function bodies for `$VAR` / `${VAR…}` references and re-emits `export NAME='value'` for each referenced var that is currently set (with a denylist for shell-internal names like `PATH`/`HOME`/`BASH_*`/`LC_*` plus a likely-secret denylist for `*TOKEN*`/`*SECRET*`/`*API_KEY*`/`*PASSWORD*`/`*PRIVATE_KEY*`/`*ACCESS_KEY*`/`*CREDENTIAL*`/`*SESSION_KEY*`), the snapshot script `umask 077`s itself and the JS caller chmods the snapshot file/dir to `0600`/`0700` so the new export pass can't leak secrets into a shared tmp dir. Fixes mise, asdf shims, direnv-style helpers, and other activation idioms that pair a function with a helper env var. `getShellConfigFile` now also honours `env.HOME` (falling back to `os.homedir()`) so sandboxed callers can target a non-default rc. ([#3470](https://github.com/can1357/oh-my-pi/issues/3470))
|
||||
@@ -3382,12 +3330,6 @@
|
||||
|
||||
## [15.2.1] - 2026-05-21
|
||||
|
||||
### Added
|
||||
|
||||
- Added `.omp-plugin/marketplace.json` as a preferred marketplace catalog path. `fetchMarketplace` now searches `.omp-plugin/marketplace.json` before `.claude-plugin/marketplace.json` for every local and cloned source. Lets a single marketplace repository publish a tool-specific catalog (e.g. an omp-only superset of a shared Claude Code marketplace) without forcing the omp/Claude distinction into per-plugin tagging. Mirrors the `package.json#omp.extensions` precedence pattern; the `.claude-plugin/marketplace.json` fallback keeps every existing marketplace loading unchanged.
|
||||
|
||||
## [15.2.1] - 2026-05-21
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed compaction routing to the wrong provider when `modelRoles.default` is set to a different model than the active chat. Auto- and manual compaction now prefer the active session's model and only fall back to role-based candidates when the current model has no usable credentials. Previously, an Anthropic chat with `modelRoles.default = openai/gpt-5` would compact through OpenAI (including the remote-compaction endpoint), even though the live conversation never used OpenAI.
|
||||
@@ -11493,6 +11435,72 @@ pi --extension ./safety.ts -e ./todo.ts
|
||||
|
||||
- Expanded keybinding documentation to list all 32 supported symbol keys with notes on ctrl+symbol behavior ([#450](https://github.com/badlogic/pi-mono/pull/450) by [@kaofelix](https://github.com/kaofelix))
|
||||
|
||||
## [0.34.0] - 2026-01-04
|
||||
|
||||
### Added
|
||||
|
||||
- Hook API: `before_agent_start` handlers can now return `systemPromptAppend` to dynamically append text to the system prompt for that turn. Multiple hooks' appends are concatenated.
|
||||
- Hook API: `before_agent_start` handlers can now return multiple messages (all are injected, not just the first)
|
||||
- New example hook: `tools.ts` - Interactive `/tools` command to enable/disable tools with session persistence
|
||||
- New example hook: `pirate.ts` - Demonstrates `systemPromptAppend` to make the agent speak like a pirate
|
||||
- Tool registry now contains all built-in tools (read, bash, edit, write, grep, find, ls) even when `--tools` limits the initially active set. Hooks can enable any tool from the registry via `pi.setActiveTools()`.
|
||||
- System prompt now automatically rebuilds when tools change via `setActiveTools()`, updating tool descriptions and guidelines to match the new tool set
|
||||
- Hook errors now display full stack traces for easier debugging
|
||||
|
||||
### Changed
|
||||
|
||||
- Removed image placeholders after copy & paste, replaced with inserting image file paths directly. ([#442](https://github.com/badlogic/pi-mono/pull/442) by [@mitsuhiko](https://github.com/mitsuhiko))
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed potential text decoding issues in bash executor by using streaming TextDecoder instead of Buffer.toString()
|
||||
- External editor (Ctrl-G) now shows full pasted content instead of `[paste #N ...]` placeholders ([#444](https://github.com/badlogic/pi-mono/pull/444) by [@aliou](https://github.com/aliou))
|
||||
|
||||
## [0.33.0] - 2026-01-04
|
||||
|
||||
### Breaking Changes
|
||||
|
||||
- **Key detection functions removed from `@mariozechner/pi-tui`**: All `isXxx()` key detection functions (`isEnter()`, `isEscape()`, `isCtrlC()`, etc.) have been removed. Use `matchesKey(data, keyId)` instead (e.g., `matchesKey(data, "enter")`, `matchesKey(data, "ctrl+c")`). This affects hooks and custom tools that use `ctx.ui.custom()` with keyboard input handling. ([#405](https://github.com/badlogic/pi-mono/pull/405))
|
||||
|
||||
### Added
|
||||
|
||||
- Clipboard image paste support via `Ctrl+V`. Images are saved to a temp file and attached to the message. Works on macOS, Windows, and Linux (X11). ([#419](https://github.com/badlogic/pi-mono/issues/419))
|
||||
- Configurable keybindings via `~/.pi/agent/keybindings.json`. All keyboard shortcuts (editor navigation, deletion, app actions like model cycling, etc.) can now be customized. Supports multiple bindings per action. ([#405](https://github.com/badlogic/pi-mono/pull/405) by [@hjanuschka](https://github.com/hjanuschka))
|
||||
- `/quit` and `/exit` slash commands to gracefully exit the application. Unlike double Ctrl+C, these properly await hook and custom tool cleanup handlers before exiting. ([#426](https://github.com/badlogic/pi-mono/pull/426) by [@ben-vargas](https://github.com/ben-vargas))
|
||||
|
||||
### Fixed
|
||||
|
||||
- Subagent example README referenced incorrect filename `subagent.ts` instead of `index.ts` ([#427](https://github.com/badlogic/pi-mono/pull/427) by [@Whamp](https://github.com/Whamp))
|
||||
|
||||
## [0.32.3] - 2026-01-03
|
||||
|
||||
### Fixed
|
||||
|
||||
- `--list-models` no longer shows Google Vertex AI models without explicit authentication configured
|
||||
- JPEG/GIF/WebP images not displaying in terminals using Kitty graphics protocol (Kitty, Ghostty, WezTerm). The protocol requires PNG format, so non-PNG images are now converted before display.
|
||||
- Version check URL typo preventing update notifications from working ([#423](https://github.com/badlogic/pi-mono/pull/423) by [@skuridin](https://github.com/skuridin))
|
||||
- Large images exceeding Anthropic's 5MB limit now retry with progressive quality/size reduction ([#424](https://github.com/badlogic/pi-mono/pull/424) by [@mitsuhiko](https://github.com/mitsuhiko))
|
||||
|
||||
## [0.32.2] - 2026-01-03
|
||||
|
||||
### Added
|
||||
|
||||
- `$ARGUMENTS` syntax for custom slash commands as alternative to `$@` for all arguments joined. Aligns with patterns used by Claude, Codex, and OpenCode. Both syntaxes remain fully supported. ([#418](https://github.com/badlogic/pi-mono/pull/418) by [@skuridin](https://github.com/skuridin))
|
||||
|
||||
### Changed
|
||||
|
||||
- **Slash commands and hook commands now work during streaming**: Previously, using a slash command or hook command while the agent was streaming would crash with "Agent is already processing". Now:
|
||||
- Hook commands execute immediately (they manage their own LLM interaction via `pi.sendMessage()`)
|
||||
- File-based slash commands are expanded and queued via steer/followUp
|
||||
- `steer()` and `followUp()` now expand file-based slash commands and error on hook commands (hook commands cannot be queued)
|
||||
- `prompt()` accepts new `streamingBehavior` option (`"steer"` or `"followUp"`) to specify queueing behavior during streaming
|
||||
- RPC `prompt` command now accepts optional `streamingBehavior` field
|
||||
([#420](https://github.com/badlogic/pi-mono/issues/420))
|
||||
|
||||
### Fixed
|
||||
|
||||
- Slash command argument substitution now processes positional arguments (`$1`, `$2`, etc.) before all-arguments (`$@`, `$ARGUMENTS`) to prevent recursive substitution when argument values contain dollar-digit patterns like `$100`. ([#418](https://github.com/badlogic/pi-mono/pull/418) by [@skuridin](https://github.com/skuridin))
|
||||
|
||||
## [0.32.1] - 2026-01-03
|
||||
|
||||
### Added
|
||||
|
||||
@@ -158,6 +158,15 @@ type LlamaCppDiscoveredModelRuntimeMetadata = {
|
||||
type LlamaCppModelListEntry = {
|
||||
id: string;
|
||||
runtimeContextWindow?: number;
|
||||
/**
|
||||
* `--ctx-size` extracted from the entry's `status.args` (rendered CLI arg
|
||||
* vector) or `status.preset` INI. Populated for llama-server router-mode
|
||||
* presets so unloaded models surface the user's configured window instead
|
||||
* of falling through to the 128K default — the router-level `/props`
|
||||
* reports a dummy `n_ctx: 0` and `meta.n_ctx` is only merged in after a
|
||||
* child instance loads (issue #4190).
|
||||
*/
|
||||
configuredContextWindow?: number;
|
||||
trainingContextWindow?: number;
|
||||
};
|
||||
|
||||
@@ -274,10 +283,60 @@ function parseLlamaCppModelList(payload: unknown): LlamaCppModelListEntry[] {
|
||||
if (!isRecord(item) || typeof item.id !== "string" || !item.id) {
|
||||
return [];
|
||||
}
|
||||
return [{ id: item.id, ...extractLlamaCppModelContextWindows(item) }];
|
||||
return [
|
||||
{
|
||||
id: item.id,
|
||||
...extractLlamaCppModelContextWindows(item),
|
||||
configuredContextWindow: extractLlamaCppConfiguredContextWindow(item),
|
||||
},
|
||||
];
|
||||
});
|
||||
}
|
||||
|
||||
// llama-server's `to_args()` renders the long form `--ctx-size` (never `-c`),
|
||||
// but tolerate the short form and the embedded `--flag=value` shape so a
|
||||
// hand-rolled forwarder cannot silently downgrade the discovered window.
|
||||
const LLAMA_CPP_CTX_SIZE_FLAGS = new Set(["--ctx-size", "-c"]);
|
||||
|
||||
function extractLlamaCppCtxSizeFromArgs(value: unknown): number | undefined {
|
||||
if (!Array.isArray(value)) {
|
||||
return undefined;
|
||||
}
|
||||
for (let i = 0; i < value.length; i++) {
|
||||
const raw = value[i];
|
||||
if (typeof raw !== "string") continue;
|
||||
const eq = raw.indexOf("=");
|
||||
const flag = eq >= 0 ? raw.slice(0, eq) : raw;
|
||||
if (!LLAMA_CPP_CTX_SIZE_FLAGS.has(flag)) continue;
|
||||
const rawValue = eq >= 0 ? raw.slice(eq + 1) : value[i + 1];
|
||||
const parsed = toPositiveNumberOrUndefined(rawValue);
|
||||
if (parsed !== undefined) return parsed;
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
// `common_preset::to_ini()` emits one option per line as `<long-arg-without-dashes> = <value>`,
|
||||
// so `ctx-size = 8192` is the exact wire form (issue #4190).
|
||||
function extractLlamaCppCtxSizeFromIni(value: unknown): number | undefined {
|
||||
if (typeof value !== "string") {
|
||||
return undefined;
|
||||
}
|
||||
const match = value.match(/(?:^|\n)\s*ctx-size\s*=\s*(-?\d+)\s*(?:$|\n)/);
|
||||
return match ? toPositiveNumberOrUndefined(match[1]) : undefined;
|
||||
}
|
||||
|
||||
function extractLlamaCppConfiguredContextWindow(item: Record<string, unknown>): number | undefined {
|
||||
const status = item.status;
|
||||
if (!isRecord(status)) {
|
||||
return undefined;
|
||||
}
|
||||
const fromArgs = extractLlamaCppCtxSizeFromArgs(status.args);
|
||||
if (fromArgs !== undefined) {
|
||||
return fromArgs;
|
||||
}
|
||||
return extractLlamaCppCtxSizeFromIni(status.preset);
|
||||
}
|
||||
|
||||
function extractLlamaCppInputCapabilities(payload: Record<string, unknown>): ("text" | "image")[] | undefined {
|
||||
const modalities = payload.modalities;
|
||||
if (!isRecord(modalities)) {
|
||||
@@ -473,6 +532,7 @@ export async function discoverLlamaCppModels(
|
||||
if (!id) continue;
|
||||
const contextWindow =
|
||||
item.runtimeContextWindow ??
|
||||
item.configuredContextWindow ??
|
||||
serverMetadata?.contextWindow ??
|
||||
item.trainingContextWindow ??
|
||||
DISCOVERY_DEFAULT_CONTEXT_WINDOW;
|
||||
@@ -529,7 +589,11 @@ export async function discoverLlamaCppModelRuntimeMetadata(
|
||||
if (!entry) {
|
||||
return undefined;
|
||||
}
|
||||
const contextWindow = entry.runtimeContextWindow ?? serverMetadata?.contextWindow ?? entry.trainingContextWindow;
|
||||
const contextWindow =
|
||||
entry.runtimeContextWindow ??
|
||||
entry.configuredContextWindow ??
|
||||
serverMetadata?.contextWindow ??
|
||||
entry.trainingContextWindow;
|
||||
if (contextWindow === undefined) {
|
||||
return undefined;
|
||||
}
|
||||
|
||||
@@ -81,6 +81,8 @@ export class ChatTranscriptBuilder {
|
||||
#readArgs = new Map<string, Record<string, unknown>>();
|
||||
#readGroup: ReadToolGroupComponent | null = null;
|
||||
#pendingUsage: Usage | undefined;
|
||||
#pendingUsageDuration: number | undefined;
|
||||
#pendingUsageTtft: number | undefined;
|
||||
#lastAssistantUsage: Usage | undefined;
|
||||
#waitingPoll: ToolExecutionComponent | null = null;
|
||||
#todoSnapshot: ToolExecutionComponent | null = null;
|
||||
@@ -127,6 +129,8 @@ export class ChatTranscriptBuilder {
|
||||
this.#readArgs.clear();
|
||||
this.#readGroup = null;
|
||||
this.#pendingUsage = undefined;
|
||||
this.#pendingUsageDuration = undefined;
|
||||
this.#pendingUsageTtft = undefined;
|
||||
this.#lastAssistantUsage = undefined;
|
||||
this.#waitingPoll = null;
|
||||
this.#todoSnapshot = null;
|
||||
@@ -192,8 +196,12 @@ export class ChatTranscriptBuilder {
|
||||
if (!this.#pendingUsage) return;
|
||||
this.#readGroup?.seal();
|
||||
this.#readGroup = null;
|
||||
this.container.addChild(createUsageRowBlock(this.#pendingUsage));
|
||||
this.container.addChild(
|
||||
createUsageRowBlock(this.#pendingUsage, this.#pendingUsageDuration, this.#pendingUsageTtft),
|
||||
);
|
||||
this.#pendingUsage = undefined;
|
||||
this.#pendingUsageDuration = undefined;
|
||||
this.#pendingUsageTtft = undefined;
|
||||
}
|
||||
|
||||
#appendChatMessage(message: AgentMessage): void {
|
||||
@@ -343,6 +351,8 @@ export class ChatTranscriptBuilder {
|
||||
}
|
||||
|
||||
this.#pendingUsage = settings.get("display.showTokenUsage") ? message.usage : undefined;
|
||||
this.#pendingUsageDuration = message.duration;
|
||||
this.#pendingUsageTtft = message.ttft;
|
||||
}
|
||||
|
||||
#appendToolResult(message: Extract<AgentMessage, { role: "toolResult" }>): void {
|
||||
|
||||
@@ -379,7 +379,7 @@ const tokenRateSegment: StatusLineSegment = {
|
||||
const { tokensPerSecond } = ctx.usageStats;
|
||||
if (!tokensPerSecond) return { content: "", visible: false };
|
||||
|
||||
const content = withIcon(theme.icon.output, `${tokensPerSecond.toFixed(1)}/s`);
|
||||
const content = withIcon(theme.icon.throughput, `${tokensPerSecond.toFixed(1)}/s`);
|
||||
return { content: theme.fg("statusLineOutput", content), visible: true };
|
||||
},
|
||||
};
|
||||
@@ -498,7 +498,7 @@ const cacheReadSegment: StatusLineSegment = {
|
||||
const { cacheRead } = ctx.usageStats;
|
||||
if (!cacheRead) return { content: "", visible: false };
|
||||
|
||||
const parts = [theme.icon.cache, theme.icon.output, formatNumber(cacheRead)].filter(Boolean);
|
||||
const parts = [theme.icon.cache, formatNumber(cacheRead)].filter(Boolean);
|
||||
const content = parts.join(" ");
|
||||
return { content: theme.fg("statusLineSpend", content), visible: true };
|
||||
},
|
||||
@@ -510,7 +510,7 @@ const cacheWriteSegment: StatusLineSegment = {
|
||||
const { cacheWrite } = ctx.usageStats;
|
||||
if (!cacheWrite) return { content: "", visible: false };
|
||||
|
||||
const parts = [theme.icon.cache, theme.icon.input, formatNumber(cacheWrite)].filter(Boolean);
|
||||
const parts = [theme.icon.cache, formatNumber(cacheWrite)].filter(Boolean);
|
||||
const content = parts.join(" ");
|
||||
return { content: theme.fg("statusLineOutput", content), visible: true };
|
||||
},
|
||||
|
||||
@@ -3,13 +3,27 @@ import { Container, Spacer, Text } from "@oh-my-pi/pi-tui";
|
||||
import { formatNumber } from "@oh-my-pi/pi-utils";
|
||||
import { theme } from "../../modes/theme/theme";
|
||||
|
||||
export function createUsageRowBlock(usage: Usage): Container {
|
||||
/** Below this the rate is nonsense (cached/instant responses yield absurd tok/s). */
|
||||
const MIN_DURATION_MS = 100;
|
||||
|
||||
export function createUsageRowBlock(usage: Usage, durationMs?: number, ttftMs?: number): Container {
|
||||
const totalInput = usage.input + usage.cacheWrite;
|
||||
const parts: string[] = [];
|
||||
parts.push(`${theme.icon.input} ${formatNumber(totalInput)}`);
|
||||
parts.push(`${theme.icon.output} ${formatNumber(usage.output)}`);
|
||||
if (usage.cacheRead > 0) {
|
||||
parts.push(`cache: ${formatNumber(usage.cacheRead)}`);
|
||||
parts.push(`${theme.icon.cache} ${formatNumber(usage.cacheRead)}`);
|
||||
}
|
||||
if (ttftMs && ttftMs > 0) {
|
||||
parts.push(`${theme.icon.time} ${(ttftMs / 1000).toFixed(1)}s`);
|
||||
}
|
||||
if (durationMs && durationMs > MIN_DURATION_MS && usage.output > 0) {
|
||||
// Throughput excludes TTFT — generation time is duration minus time-to-first-token.
|
||||
const genMs = durationMs - (ttftMs ?? 0);
|
||||
if (genMs > MIN_DURATION_MS) {
|
||||
const tokPerSec = (usage.output / genMs) * 1000;
|
||||
parts.push(`${theme.icon.throughput} ${tokPerSec.toFixed(1)}/s`);
|
||||
}
|
||||
}
|
||||
const block = new Container();
|
||||
block.addChild(new Spacer(1));
|
||||
|
||||
@@ -786,7 +786,9 @@ export class EventController {
|
||||
this.#lastAssistantComponent = this.ctx.streamingComponent;
|
||||
this.#lastAssistantComponent.markTranscriptBlockFinalized();
|
||||
if (settings.get("display.showTokenUsage")) {
|
||||
this.ctx.chatContainer.addChild(createUsageRowBlock(event.message.usage));
|
||||
this.ctx.chatContainer.addChild(
|
||||
createUsageRowBlock(event.message.usage, event.message.duration, event.message.ttft),
|
||||
);
|
||||
}
|
||||
this.ctx.streamingComponent = undefined;
|
||||
this.ctx.streamingMessage = undefined;
|
||||
|
||||
@@ -115,6 +115,7 @@ export type SymbolKey =
|
||||
| "icon.cacheMiss"
|
||||
| "icon.input"
|
||||
| "icon.output"
|
||||
| "icon.throughput"
|
||||
| "icon.host"
|
||||
| "icon.session"
|
||||
| "icon.package"
|
||||
@@ -319,6 +320,7 @@ const UNICODE_SYMBOLS: SymbolMap = {
|
||||
"icon.cacheMiss": "⊘",
|
||||
"icon.input": "⤵",
|
||||
"icon.output": "⤴",
|
||||
"icon.throughput": "⚡",
|
||||
"icon.host": "🖥",
|
||||
"icon.session": "🆔",
|
||||
"icon.package": "📦",
|
||||
@@ -597,6 +599,8 @@ const NERD_SYMBOLS: SymbolMap = {
|
||||
"icon.input": "\uf090",
|
||||
// pick: | alt: →
|
||||
"icon.output": "\uf08b",
|
||||
// pick: (nf-fa-tachometer) | alt: ⚡ ↬
|
||||
"icon.throughput": "\uf0e4",
|
||||
// pick: | alt:
|
||||
"icon.host": "\uf109",
|
||||
// pick: | alt:
|
||||
@@ -827,10 +831,11 @@ const ASCII_SYMBOLS: SymbolMap = {
|
||||
"icon.ghost": "@",
|
||||
"icon.agents": "AG",
|
||||
"icon.job": "bg",
|
||||
"icon.output": "out:",
|
||||
"icon.throughput": "tok/s:",
|
||||
"icon.cache": "cache",
|
||||
"icon.cacheMiss": "!",
|
||||
"icon.input": "in:",
|
||||
"icon.output": "out:",
|
||||
"icon.host": "host",
|
||||
"icon.session": "id",
|
||||
"icon.package": "[P]",
|
||||
@@ -1820,6 +1825,7 @@ export class Theme {
|
||||
cacheMiss: this.#symbols["icon.cacheMiss"],
|
||||
input: this.#symbols["icon.input"],
|
||||
output: this.#symbols["icon.output"],
|
||||
throughput: this.#symbols["icon.throughput"],
|
||||
host: this.#symbols["icon.host"],
|
||||
session: this.#symbols["icon.session"],
|
||||
package: this.#symbols["icon.package"],
|
||||
|
||||
@@ -297,12 +297,16 @@ export class UiHelpers {
|
||||
// read run so the row sits under it. Mirrors the live path, where the read
|
||||
// group is created during streaming and the row is appended below it.
|
||||
let pendingUsage: Usage | undefined;
|
||||
let pendingUsageDuration: number | undefined;
|
||||
let pendingUsageTtft: number | undefined;
|
||||
const flushPendingUsage = () => {
|
||||
if (!pendingUsage) return;
|
||||
readGroup?.seal();
|
||||
readGroup = null;
|
||||
this.ctx.chatContainer.addChild(createUsageRowBlock(pendingUsage));
|
||||
this.ctx.chatContainer.addChild(createUsageRowBlock(pendingUsage, pendingUsageDuration, pendingUsageTtft));
|
||||
pendingUsage = undefined;
|
||||
pendingUsageDuration = undefined;
|
||||
pendingUsageTtft = undefined;
|
||||
};
|
||||
// Rebuild-time mirror of the event controller's displaceable-poll
|
||||
// bookkeeping: a `job` poll that found every watched job still running is
|
||||
@@ -455,6 +459,8 @@ export class UiHelpers {
|
||||
}
|
||||
}
|
||||
pendingUsage = this.ctx.settings.get("display.showTokenUsage") ? message.usage : undefined;
|
||||
pendingUsageDuration = message.duration;
|
||||
pendingUsageTtft = message.ttft;
|
||||
} else if (message.role === "toolResult") {
|
||||
const pendingReadComponent = this.ctx.pendingTools.get(message.toolCallId);
|
||||
const isReadGroupResult =
|
||||
|
||||
@@ -1,12 +1,3 @@
|
||||
<system-reminder>
|
||||
You have completed several tool calls since the last `{{toolRefs.todo}}` update, and {{incompleteCount}} todo item{{#if plural}}s{{/if}} still {{#if plural}}remain{{else}}remains{{/if}} pending or in_progress:
|
||||
{{#each phases}}
|
||||
- {{name}}
|
||||
{{#each tasks}}
|
||||
- {{content}} ({{status}})
|
||||
{{/each}}
|
||||
{{/each}}
|
||||
|
||||
If any are now done, call `{{toolRefs.todo}}` with `op: "done"` so the live HUD matches reality. Keep todos in sync as you work — do not batch every transition at the end of the run.
|
||||
(Mid-run reminder {{attempt}}/{{maxAttempts}})
|
||||
Gentle reminder: {{incompleteCount}} todo item{{#if plural}}s are{{else}} is{{/if}} still open. If you finished a task since the last `{{toolRefs.todo}}` update, mark it done now so progress stays visible; otherwise just keep working.
|
||||
</system-reminder>
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
Checkpoint completed. The checkpoint's exploratory branch was rewound; the branch summary and retained report below are now the context.
|
||||
|
||||
Do not call `rewind` again for this checkpoint. Continue from this retained report.
|
||||
|
||||
Report:
|
||||
{{report}}
|
||||
@@ -9,5 +9,6 @@ Requirements:
|
||||
- You MUST call this before yielding if a checkpoint is active.
|
||||
|
||||
Behavior:
|
||||
- If no checkpoint is active, this tool errors.
|
||||
- On success, the session rewinds and keeps your report as retained context.
|
||||
- If no checkpoint is active, this tool errors. If the checkpoint already rewound, continue from the retained report instead of retrying.
|
||||
- On success, the session rewinds, keeps your report as retained context, and closes the checkpoint.
|
||||
- A successful rewind is final for that checkpoint; repeat calls error.
|
||||
|
||||
@@ -1578,6 +1578,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
|
||||
activateDiscoveredTools: toolNames => session.activateDiscoveredTools(toolNames),
|
||||
getCheckpointState: () => session.getCheckpointState(),
|
||||
setCheckpointState: state => session.setCheckpointState(state ?? undefined),
|
||||
getLastCompletedRewind: () => session.getLastCompletedRewind(),
|
||||
getToolChoiceQueue: () => session.toolChoiceQueue,
|
||||
buildToolChoice: name => {
|
||||
const m = session.model;
|
||||
|
||||
@@ -64,7 +64,6 @@ import {
|
||||
renderHandoffPrompt,
|
||||
resolveBudgetReserveTokens,
|
||||
resolveThresholdTokens,
|
||||
type SessionEntry,
|
||||
type SessionMessageEntry,
|
||||
type ShakeConfig,
|
||||
type ShakeRegion,
|
||||
@@ -257,6 +256,7 @@ import planModeReferencePrompt from "../prompts/system/plan-mode-reference.md" w
|
||||
import planModeToolDecisionReminderPrompt from "../prompts/system/plan-mode-tool-decision-reminder.md" with {
|
||||
type: "text",
|
||||
};
|
||||
import rewindReportTemplate from "../prompts/system/rewind-report.md" with { type: "text" };
|
||||
import sideChannelNoToolsReminder from "../prompts/system/side-channel-no-tools.md" with { type: "text" };
|
||||
import thinkingLoopRedirectTemplate from "../prompts/system/thinking-loop-redirect.md" with { type: "text" };
|
||||
import toolCallLoopRedirectTemplate from "../prompts/system/tool-call-loop-redirect.md" with { type: "text" };
|
||||
@@ -298,7 +298,7 @@ import {
|
||||
import { assertEditableFile } from "../tools/auto-generated-guard";
|
||||
import { releaseTabsForOwner } from "../tools/browser/tab-supervisor";
|
||||
import { normalizeToolNames } from "../tools/builtin-names";
|
||||
import type { CheckpointState } from "../tools/checkpoint";
|
||||
import type { CheckpointState, CompletedRewindState } from "../tools/checkpoint";
|
||||
import { outputMeta, wrapToolWithMetaNotice } from "../tools/output-meta";
|
||||
import { normalizeLocalScheme, resolveToCwd } from "../tools/path-utils";
|
||||
import { isAutoQaEnabled } from "../tools/report-tool-issue";
|
||||
@@ -350,7 +350,7 @@ import {
|
||||
import type { BuildSessionContextOptions, SessionContext } from "./session-context";
|
||||
import { getLatestCompactionEntry, getRestorableSessionModels } from "./session-context";
|
||||
import { formatSessionDumpText } from "./session-dump-format";
|
||||
import type { BranchSummaryEntry, CompactionEntry, NewSessionOptions } from "./session-entries";
|
||||
import type { BranchSummaryEntry, CompactionEntry, NewSessionOptions, SessionEntry } from "./session-entries";
|
||||
import { EPHEMERAL_MODEL_CHANGE_ROLE } from "./session-entries";
|
||||
import { formatSessionHistoryMarkdown } from "./session-history-format";
|
||||
import { cleanupEmptyMoveSession, type SessionManager } from "./session-manager";
|
||||
@@ -363,14 +363,31 @@ import { YieldQueue } from "./yield-queue";
|
||||
const SESSION_STOP_CONTINUATION_CAP = 8;
|
||||
|
||||
/**
|
||||
* Consecutive tool-use turns without the agent touching the `todo` tool that
|
||||
* trip the mid-run reconciliation nudge. Picked so a single batch of edits +
|
||||
* verification (~3-5 turns each) does not trigger it, but a sustained loop of
|
||||
* work without flipping any todos does. Without this nudge, long runs drive
|
||||
* the live todo HUD to `0/N` until the final stop, then batch-flip to `N/N`
|
||||
* (issue #3651).
|
||||
* Mutating tool results (`bash`/`eval`/`edit`/`write`/`ast_edit`) without the
|
||||
* agent touching the `todo` tool that trip the mid-run reconciliation nudge.
|
||||
* Read-only exploration (grep/read/glob/lsp) never ticks this: an agent
|
||||
* researching for a long stretch has nothing to flip. Picked so a normal
|
||||
* fix-verify loop (~3-6 mutations) never sees the nudge, but a sustained run
|
||||
* of landed work without flipping any todos does. Without this nudge, long
|
||||
* runs drive the live todo HUD to `0/N` until the final stop, then batch-flip
|
||||
* to `N/N` (issue #3651).
|
||||
*/
|
||||
const MID_RUN_TODO_NUDGE_TURN_THRESHOLD = 8;
|
||||
const MID_RUN_TODO_NUDGE_MUTATION_THRESHOLD = 12;
|
||||
/** Mid-run nudges per prompt cycle. Deliberately tighter than
|
||||
* `todo.reminders.max` (the stop-time budget): this is a gentle hidden hint,
|
||||
* not an escalation ladder. */
|
||||
const MID_RUN_TODO_NUDGE_MAX_PER_CYCLE = 2;
|
||||
/** Tool results that count as landed work for the mid-run todo nudge. */
|
||||
const MID_RUN_TODO_NUDGE_MUTATING_TOOLS: Record<string, true> = {
|
||||
bash: true,
|
||||
eval: true,
|
||||
edit: true,
|
||||
write: true,
|
||||
ast_edit: true,
|
||||
};
|
||||
/** `customType` for the hidden mid-run todo nudge; `display: false`, so it reaches
|
||||
* the model but never renders in the TUI or transcript. */
|
||||
const MID_RUN_TODO_NUDGE_MESSAGE_TYPE = "mid-run-todo-nudge";
|
||||
|
||||
/** Abort reason for the Gemini reasoning-header runaway interrupt. Surfaced on the
|
||||
* discarded assistant turn only; never reaches the model. */
|
||||
@@ -382,6 +399,49 @@ const GEMINI_TOOL_REMINDER_TYPE = "gemini-tool-call-reminder";
|
||||
const THINKING_LOOP_REDIRECT_TYPE = "thinking-loop-redirect";
|
||||
const TOOL_CALL_LOOP_REDIRECT_TYPE = "tool-call-loop-redirect";
|
||||
|
||||
function customMessageContentText(content: string | (TextContent | ImageContent)[]): string {
|
||||
if (typeof content === "string") return content;
|
||||
const parts: string[] = [];
|
||||
for (const part of content) {
|
||||
if (part.type === "text") parts.push(part.text);
|
||||
}
|
||||
return parts.join("\n");
|
||||
}
|
||||
|
||||
function stringProperty(value: object, key: string): string | undefined {
|
||||
const field = Object.getOwnPropertyDescriptor(value, key)?.value;
|
||||
return typeof field === "string" ? field : undefined;
|
||||
}
|
||||
|
||||
function reportFromRewindReportContent(content: string): string {
|
||||
const marker = "\nReport:\n";
|
||||
const index = content.lastIndexOf(marker);
|
||||
const report = index >= 0 ? content.slice(index + marker.length) : content;
|
||||
return report.trim();
|
||||
}
|
||||
|
||||
function completedRewindFromEntry(entry: SessionEntry): CompletedRewindState | undefined {
|
||||
if (entry.type !== "custom_message" || entry.customType !== "rewind-report") return undefined;
|
||||
const details = entry.details;
|
||||
if (!details || typeof details !== "object") return undefined;
|
||||
const startedAt = stringProperty(details, "startedAt");
|
||||
const rewoundAt = stringProperty(details, "rewoundAt");
|
||||
if (!startedAt || !rewoundAt) return undefined;
|
||||
const report =
|
||||
stringProperty(details, "report")?.trim() ||
|
||||
reportFromRewindReportContent(customMessageContentText(entry.content));
|
||||
return report.length > 0 ? { report, startedAt, rewoundAt } : undefined;
|
||||
}
|
||||
|
||||
function isSuccessfulCheckpointEntry(entry: SessionEntry): boolean {
|
||||
return (
|
||||
entry.type === "message" &&
|
||||
entry.message.role === "toolResult" &&
|
||||
entry.message.toolName === "checkpoint" &&
|
||||
entry.message.isError !== true
|
||||
);
|
||||
}
|
||||
|
||||
// A side-channel assistant response is signed for the hidden prompt/history that
|
||||
// produced it. If we persist that response under a different user turn, native
|
||||
// replay anchors become invalid; keep only visible, non-cryptographic content.
|
||||
@@ -1506,14 +1566,18 @@ export class AgentSession {
|
||||
*/
|
||||
#todoReminderAwaitingProgress = false;
|
||||
/**
|
||||
* Consecutive tool-use turns the agent has completed without invoking the
|
||||
* `todo` tool. Drives {@link #takeMidRunTodoNudge} so the live HUD stays in
|
||||
* sync with actual progress instead of flipping `0/N -> N/N` only at the
|
||||
* very end of a long run (issue #3651). Reset to 0 on any assistant turn
|
||||
* that calls `todo`, on a successful nudge fire (cooldown), and at every
|
||||
* Successful mutating tool results (bash/eval/edit/write/ast_edit) since the
|
||||
* agent last touched the `todo` tool. Drives {@link #takeMidRunTodoNudge} so
|
||||
* the live HUD stays in sync with actual progress instead of flipping
|
||||
* `0/N -> N/N` only at the very end of a long run (issue #3651). Read-only
|
||||
* tools and errored results never tick it. Reset to 0 on any `todo` tool
|
||||
* result, on a nudge fire (cooldown), on a stop-time reminder, and at every
|
||||
* new-prompt / clear / handoff lifecycle boundary.
|
||||
*/
|
||||
#toolTurnsSinceLastTodoTouch = 0;
|
||||
#mutationsSinceLastTodoTouch = 0;
|
||||
/** Mid-run nudges fired this prompt cycle; capped by
|
||||
* {@link MID_RUN_TODO_NUDGE_MAX_PER_CYCLE}, reset with the counter above. */
|
||||
#midRunNudgeCount = 0;
|
||||
#todoPhases: TodoPhase[] = [];
|
||||
#replanTitleRefreshInFlight: Promise<void> | undefined = undefined;
|
||||
/** Resolved TITLE_SYSTEM.md override applied to every automatic session-title
|
||||
@@ -1693,6 +1757,7 @@ export class AgentSession {
|
||||
#pruneToolDescriptions = false;
|
||||
#checkpointState: CheckpointState | undefined = undefined;
|
||||
#pendingRewindReport: string | undefined = undefined;
|
||||
#lastCompletedRewind: CompletedRewindState | undefined = undefined;
|
||||
#rewoundToolResultIds = new Set<string>();
|
||||
#lastSuccessfulYieldToolCallId: string | undefined = undefined;
|
||||
/**
|
||||
@@ -2104,6 +2169,8 @@ export class AgentSession {
|
||||
this.#advisorEnabled = this.settings.get("advisor.enabled") as boolean;
|
||||
if (this.#advisorEnabled) this.#buildAdvisorRuntime();
|
||||
|
||||
this.#rehydrateLastCompletedRewind();
|
||||
|
||||
// Always subscribe to agent events for internal handling
|
||||
// (session persistence, hooks, auto-compaction, retry logic)
|
||||
this.#unsubscribeAgent = this.agent.subscribe(this.#handleAgentEvent);
|
||||
@@ -3230,18 +3297,16 @@ export class AgentSession {
|
||||
// trip a spurious nudge against stale state, and a turn that just hit
|
||||
// the threshold could fail to nudge until a later turn (issue #3651).
|
||||
// Pure in-memory math — no ordering requirement vs persistence or
|
||||
// session-event fan-out.
|
||||
if (event.type === "message_end" && event.message.role === "assistant") {
|
||||
const tooledTurn = event.message.content.some(content => content.type === "toolCall");
|
||||
if (tooledTurn) {
|
||||
const touchedTodo = event.message.content.some(
|
||||
content => content.type === "toolCall" && content.name === "todo",
|
||||
);
|
||||
if (touchedTodo) {
|
||||
this.#toolTurnsSinceLastTodoTouch = 0;
|
||||
} else {
|
||||
this.#toolTurnsSinceLastTodoTouch++;
|
||||
}
|
||||
// session-event fan-out. Keyed on toolResult (not the assistant toolCall
|
||||
// turn) so planned-but-aborted or permission-denied calls never count,
|
||||
// and only successful mutating tools tick — read-only exploration is
|
||||
// not progress an agent could mark done.
|
||||
if (event.type === "message_end" && event.message.role === "toolResult") {
|
||||
const { toolName, isError } = event.message;
|
||||
if (toolName === "todo") {
|
||||
this.#mutationsSinceLastTodoTouch = 0;
|
||||
} else if (!isError && MID_RUN_TODO_NUDGE_MUTATING_TOOLS[toolName]) {
|
||||
this.#mutationsSinceLastTodoTouch++;
|
||||
}
|
||||
}
|
||||
// Plan-mode internal transition: stamp `SILENT_ABORT_MARKER` on the
|
||||
@@ -3541,6 +3606,7 @@ export class AgentSession {
|
||||
startedAt: details?.startedAt ?? new Date().toISOString(),
|
||||
};
|
||||
this.#pendingRewindReport = undefined;
|
||||
this.#lastCompletedRewind = undefined;
|
||||
}
|
||||
if (toolName === "rewind" && !isError && this.#checkpointState) {
|
||||
const detailReport = typeof details?.report === "string" ? details.report.trim() : "";
|
||||
@@ -6651,13 +6717,31 @@ export class AgentSession {
|
||||
this.agent.setTools(activeTools);
|
||||
}
|
||||
|
||||
#rehydrateLastCompletedRewind(): void {
|
||||
let completed: CompletedRewindState | undefined;
|
||||
for (const entry of this.sessionManager.getBranch()) {
|
||||
if (isSuccessfulCheckpointEntry(entry)) {
|
||||
completed = undefined;
|
||||
continue;
|
||||
}
|
||||
completed = completedRewindFromEntry(entry) ?? completed;
|
||||
}
|
||||
this.#lastCompletedRewind = completed;
|
||||
}
|
||||
|
||||
getCheckpointState(): CheckpointState | undefined {
|
||||
return this.#checkpointState;
|
||||
}
|
||||
|
||||
getLastCompletedRewind(): CompletedRewindState | undefined {
|
||||
return this.#lastCompletedRewind;
|
||||
}
|
||||
|
||||
setCheckpointState(state: CheckpointState | undefined): void {
|
||||
this.#checkpointState = state;
|
||||
if (!state) {
|
||||
if (state) {
|
||||
this.#lastCompletedRewind = undefined;
|
||||
} else {
|
||||
this.#pendingRewindReport = undefined;
|
||||
}
|
||||
}
|
||||
@@ -7189,7 +7273,8 @@ export class AgentSession {
|
||||
// Reset todo reminder count on new user prompt
|
||||
this.#todoReminderCount = 0;
|
||||
this.#todoReminderAwaitingProgress = false;
|
||||
this.#toolTurnsSinceLastTodoTouch = 0;
|
||||
this.#mutationsSinceLastTodoTouch = 0;
|
||||
this.#midRunNudgeCount = 0;
|
||||
this.#emptyStopRetryCount = 0;
|
||||
this.#unexpectedStopRetryCount = 0;
|
||||
// A new prompt cycle starts: drop any sticky yield-termination from the
|
||||
@@ -8246,7 +8331,8 @@ export class AgentSession {
|
||||
|
||||
this.#todoReminderCount = 0;
|
||||
this.#todoReminderAwaitingProgress = false;
|
||||
this.#toolTurnsSinceLastTodoTouch = 0;
|
||||
this.#mutationsSinceLastTodoTouch = 0;
|
||||
this.#midRunNudgeCount = 0;
|
||||
this.#planReferenceSent = false;
|
||||
this.#planReferencePath = "local://PLAN.md";
|
||||
this.#resetAdvisorSessionState();
|
||||
@@ -9652,7 +9738,8 @@ export class AgentSession {
|
||||
this.#scheduledHiddenNextTurnGeneration = undefined;
|
||||
this.#todoReminderCount = 0;
|
||||
this.#todoReminderAwaitingProgress = false;
|
||||
this.#toolTurnsSinceLastTodoTouch = 0;
|
||||
this.#mutationsSinceLastTodoTouch = 0;
|
||||
this.#midRunNudgeCount = 0;
|
||||
|
||||
// Inject the handoff document as a custom message
|
||||
const handoffContent = createHandoffContext(handoffText);
|
||||
@@ -10380,8 +10467,16 @@ export class AgentSession {
|
||||
});
|
||||
this.sessionManager.branchWithSummary(null, report, { startedAt: checkpointState.startedAt });
|
||||
}
|
||||
const details = { startedAt: checkpointState.startedAt, rewoundAt: new Date().toISOString() };
|
||||
this.sessionManager.appendCustomMessageEntry("rewind-report", report, false, details, "agent");
|
||||
const rewoundAt = new Date().toISOString();
|
||||
const details = { report, startedAt: checkpointState.startedAt, rewoundAt };
|
||||
this.sessionManager.appendCustomMessageEntry(
|
||||
"rewind-report",
|
||||
prompt.render(rewindReportTemplate, { report }),
|
||||
false,
|
||||
details,
|
||||
"agent",
|
||||
);
|
||||
this.#lastCompletedRewind = { report, startedAt: checkpointState.startedAt, rewoundAt };
|
||||
|
||||
if (activeMessages) {
|
||||
for (const message of activeMessages) {
|
||||
@@ -10670,8 +10765,8 @@ export class AgentSession {
|
||||
// A stop-time reminder starts a fresh reminder runway. Without resetting
|
||||
// the mid-run counter here, a run that stopped just below the threshold
|
||||
// would spend its stale pre-reminder count and fire "Mid-run reminder 2/3"
|
||||
// after only one post-reminder tool turn.
|
||||
this.#toolTurnsSinceLastTodoTouch = 0;
|
||||
// after only a little post-reminder work.
|
||||
this.#mutationsSinceLastTodoTouch = 0;
|
||||
this.#todoReminderAwaitingProgress = true;
|
||||
// Inject reminder and persist it so the JSONL transcript matches model context.
|
||||
this.agent.appendMessage(reminderMessage);
|
||||
@@ -10681,22 +10776,23 @@ export class AgentSession {
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the next mid-run todo reconciliation nudge when the agent has gone
|
||||
* {@link MID_RUN_TODO_NUDGE_TURN_THRESHOLD} tool-use turns without invoking
|
||||
* the `todo` tool and incomplete items remain. Returns the developer-role
|
||||
* reminder when it should fire, or `null` to skip. Called once per turn via
|
||||
* the aside provider; mutates internal counters when it fires so the caller
|
||||
* does not need to track delivery state.
|
||||
* Build the next mid-run todo reconciliation nudge when the agent has landed
|
||||
* {@link MID_RUN_TODO_NUDGE_MUTATION_THRESHOLD} mutating tool results without
|
||||
* invoking the `todo` tool and incomplete items remain. Returns the hidden
|
||||
* (`display: false`) custom message when it should fire, or `null` to skip.
|
||||
* Called once per turn via the aside provider; mutates internal counters when
|
||||
* it fires so the caller does not need to track delivery state.
|
||||
*
|
||||
* Companion to {@link #checkTodoCompletion}, which only fires when the agent
|
||||
* stops with text. Without this, a long tool-use loop drives the live HUD
|
||||
* to `0/N` until the final stop, then batch-flips to `N/N` (issue #3651).
|
||||
* Shares `#todoReminderCount` / `#todoReminderAwaitingProgress` with the
|
||||
* stop-time path so the per-cycle reminder budget is unified.
|
||||
* Deliberately a SEPARATE concept from {@link #checkTodoCompletion}'s
|
||||
* stop-time reminder: this is a gentle model-only hint (no `todo_reminder`
|
||||
* event, no TUI render, no escalation counter, own per-cycle budget), while
|
||||
* the stop-time reminder is the user-visible escalation ladder. Without this
|
||||
* nudge, long runs drive the live HUD to `0/N` until the final stop, then
|
||||
* batch-flip to `N/N` (issue #3651).
|
||||
*/
|
||||
#takeMidRunTodoNudge(): AgentMessage | null {
|
||||
if (this.#toolTurnsSinceLastTodoTouch < MID_RUN_TODO_NUDGE_TURN_THRESHOLD) return null;
|
||||
if (this.#todoReminderAwaitingProgress) return null;
|
||||
if (this.#mutationsSinceLastTodoTouch < MID_RUN_TODO_NUDGE_MUTATION_THRESHOLD) return null;
|
||||
if (this.#midRunNudgeCount >= MID_RUN_TODO_NUDGE_MAX_PER_CYCLE) return null;
|
||||
if (!this.settings.get("todo.enabled")) return null;
|
||||
if (!this.settings.get("todo.reminders")) return null;
|
||||
// Plan-mode runs are authoring a plan file, not implementing it; todos
|
||||
@@ -10709,54 +10805,33 @@ export class AgentSession {
|
||||
// schema — the request would fabricate an unknown tool call.
|
||||
if (!this.getActiveToolNames().includes("todo")) return null;
|
||||
|
||||
const remindersMax = this.settings.get("todo.reminders.max");
|
||||
if (this.#todoReminderCount >= remindersMax) return null;
|
||||
|
||||
const phases = this.getTodoPhases();
|
||||
if (phases.length === 0) return null;
|
||||
const incompleteByPhase = phases
|
||||
.map(phase => ({
|
||||
name: phase.name,
|
||||
tasks: phase.tasks
|
||||
.filter(
|
||||
(task): task is TodoItem & { status: "pending" | "in_progress" } =>
|
||||
task.status === "pending" || task.status === "in_progress",
|
||||
)
|
||||
.map(task => ({ content: task.content, status: task.status })),
|
||||
}))
|
||||
.filter(phase => phase.tasks.length > 0);
|
||||
const incomplete = incompleteByPhase.flatMap(phase => phase.tasks);
|
||||
const incomplete = this.getTodoPhases()
|
||||
.flatMap(phase => phase.tasks)
|
||||
.filter(task => task.status === "pending" || task.status === "in_progress");
|
||||
if (incomplete.length === 0) return null;
|
||||
|
||||
// Reset the turn counter so the nudge has another full runway before the
|
||||
// next mid-run fire; the shared #todoReminderCount cap keeps total reminders
|
||||
// (mid-run + stop-time) bounded per cycle.
|
||||
this.#toolTurnsSinceLastTodoTouch = 0;
|
||||
this.#todoReminderCount++;
|
||||
this.#todoReminderAwaitingProgress = true;
|
||||
const attempt = this.#todoReminderCount;
|
||||
// Reset the mutation counter so the nudge has another full runway before
|
||||
// the next fire; #midRunNudgeCount caps total nudges per prompt cycle.
|
||||
this.#mutationsSinceLastTodoTouch = 0;
|
||||
this.#midRunNudgeCount++;
|
||||
|
||||
const { toolRefs } = this.#buildEagerPreludeContext();
|
||||
const reminder = prompt.render(midRunTodoNudgePrompt, {
|
||||
toolRefs,
|
||||
incompleteCount: incomplete.length,
|
||||
plural: incomplete.length !== 1,
|
||||
phases: incompleteByPhase,
|
||||
attempt,
|
||||
maxAttempts: remindersMax,
|
||||
});
|
||||
|
||||
logger.debug("Mid-run todo nudge fired", { incomplete: incomplete.length, attempt });
|
||||
void this.#emitSessionEvent({
|
||||
type: "todo_reminder",
|
||||
todos: incomplete,
|
||||
attempt,
|
||||
maxAttempts: remindersMax,
|
||||
logger.debug("Mid-run todo nudge fired", {
|
||||
incomplete: incomplete.length,
|
||||
nudge: this.#midRunNudgeCount,
|
||||
});
|
||||
|
||||
return {
|
||||
role: "developer",
|
||||
content: [{ type: "text", text: reminder }],
|
||||
role: "custom",
|
||||
customType: MID_RUN_TODO_NUDGE_MESSAGE_TYPE,
|
||||
content: reminder,
|
||||
display: false,
|
||||
attribution: "agent",
|
||||
timestamp: Date.now(),
|
||||
};
|
||||
@@ -13909,6 +13984,8 @@ export class AgentSession {
|
||||
? this.#getSessionDefaultSelectedMCPToolNames(previousSessionFile)
|
||||
: undefined;
|
||||
|
||||
const previousLastCompletedRewind = this.#lastCompletedRewind;
|
||||
|
||||
this.agent.clearAllQueues();
|
||||
this.#pendingNextTurnMessages = [];
|
||||
this.#scheduledHiddenNextTurnGeneration = undefined;
|
||||
@@ -13928,6 +14005,7 @@ export class AgentSession {
|
||||
this.#didSessionMessagesChange(previousSessionContext.messages, sessionContext.messages);
|
||||
const fallbackSelectedMCPToolNames = this.#getSessionDefaultSelectedMCPToolNames(sessionPath);
|
||||
await this.#restoreMCPSelectionsForSessionContext(sessionContext, { fallbackSelectedMCPToolNames });
|
||||
this.#rehydrateLastCompletedRewind();
|
||||
|
||||
// Emit session_switch event to hooks
|
||||
if (this.#extensionRunner) {
|
||||
@@ -14069,6 +14147,7 @@ export class AgentSession {
|
||||
this.agent.replaceQueues(previousSteeringMessages, previousFollowUpMessages);
|
||||
this.#pendingNextTurnMessages = previousPendingNextTurnMessages;
|
||||
this.#scheduledHiddenNextTurnGeneration = previousScheduledHiddenNextTurnGeneration;
|
||||
this.#lastCompletedRewind = previousLastCompletedRewind;
|
||||
if (previousModel) {
|
||||
this.agent.setModel(previousModel);
|
||||
}
|
||||
@@ -14419,6 +14498,7 @@ export class AgentSession {
|
||||
const displayContext = deobfuscateSessionContext(stateContext, this.#obfuscator);
|
||||
await this.#restoreMCPSelectionsForSessionContext(displayContext);
|
||||
this.agent.replaceMessages(displayContext.messages);
|
||||
this.#rehydrateLastCompletedRewind();
|
||||
this.#resetAdvisorSessionState();
|
||||
this.#syncTodoPhasesFromBranch();
|
||||
this.#closeCodexProviderSessionsForHistoryRewrite();
|
||||
|
||||
@@ -17,6 +17,15 @@ export interface CheckpointState {
|
||||
startedAt: string;
|
||||
}
|
||||
|
||||
export interface CompletedRewindState {
|
||||
/** Report retained after a successful rewind. */
|
||||
report: string;
|
||||
/** Timestamp for the checkpoint that was rewound. */
|
||||
startedAt: string;
|
||||
/** Timestamp when the rewind completed. */
|
||||
rewoundAt: string;
|
||||
}
|
||||
|
||||
const checkpointSchema = type({
|
||||
goal: type("string").describe("investigation goal"),
|
||||
});
|
||||
@@ -123,7 +132,12 @@ export class RewindTool implements AgentTool<typeof rewindSchema, RewindToolDeta
|
||||
throw new ToolError("Checkpoint not available in subagents.");
|
||||
}
|
||||
if (!this.session.getCheckpointState?.()) {
|
||||
throw new ToolError("No active checkpoint.");
|
||||
if (this.session.getLastCompletedRewind?.()) {
|
||||
throw new ToolError(
|
||||
"Checkpoint already completed; continue from the retained rewind report instead of calling rewind again.",
|
||||
);
|
||||
}
|
||||
throw new ToolError("No active checkpoint. Create a checkpoint before calling rewind.");
|
||||
}
|
||||
const report = params.report.trim();
|
||||
if (report.length === 0) {
|
||||
|
||||
@@ -40,7 +40,7 @@ import { AstGrepTool } from "./ast-grep";
|
||||
import { BashTool } from "./bash";
|
||||
import { BrowserTool } from "./browser";
|
||||
import { type BuiltinToolName, normalizeToolNames } from "./builtin-names";
|
||||
import { type CheckpointState, CheckpointTool, RewindTool } from "./checkpoint";
|
||||
import { type CheckpointState, CheckpointTool, type CompletedRewindState, RewindTool } from "./checkpoint";
|
||||
import { DebugTool } from "./debug";
|
||||
import { EvalTool } from "./eval";
|
||||
import { resolveEvalBackends } from "./eval-backends";
|
||||
@@ -332,6 +332,8 @@ export interface ToolSession {
|
||||
getCheckpointState?: () => CheckpointState | undefined;
|
||||
/** Set or clear active checkpoint state. */
|
||||
setCheckpointState?: (state: CheckpointState | null) => void;
|
||||
/** Get the most recent completed rewind, if this session just rewound a checkpoint. */
|
||||
getLastCompletedRewind?: () => CompletedRewindState | undefined;
|
||||
|
||||
/** Per-session snapshot store of file contents as last shown to the model
|
||||
* by `read`/`search`. Used by hashline anchor-stale recovery to
|
||||
|
||||
@@ -10,6 +10,7 @@ import { AgentSession } from "@oh-my-pi/pi-coding-agent/session/agent-session";
|
||||
import { AuthStorage } from "@oh-my-pi/pi-coding-agent/session/auth-storage";
|
||||
import { convertToLlm } from "@oh-my-pi/pi-coding-agent/session/messages";
|
||||
import { SessionManager } from "@oh-my-pi/pi-coding-agent/session/session-manager";
|
||||
import { RewindTool, type ToolSession } from "@oh-my-pi/pi-coding-agent/tools";
|
||||
import { TempDir } from "@oh-my-pi/pi-utils";
|
||||
|
||||
const checkpointSchema = z.object({ goal: z.string() });
|
||||
@@ -44,6 +45,7 @@ const rewindTool: AgentTool<typeof rewindSchema, { report: string; rewound: bool
|
||||
type Harness = {
|
||||
session: AgentSession;
|
||||
authStorage: AuthStorage;
|
||||
extraSessions: AgentSession[];
|
||||
tempDir: TempDir;
|
||||
};
|
||||
|
||||
@@ -52,6 +54,9 @@ const activeHarnesses: Harness[] = [];
|
||||
afterEach(async () => {
|
||||
while (activeHarnesses.length > 0) {
|
||||
const harness = activeHarnesses.pop();
|
||||
for (const extraSession of harness?.extraSessions ?? []) {
|
||||
await extraSession.dispose();
|
||||
}
|
||||
await harness?.session.dispose();
|
||||
harness?.authStorage.close();
|
||||
harness?.tempDir.removeSync();
|
||||
@@ -98,7 +103,7 @@ async function createHarness(responses: MockResponse[]): Promise<Harness & { moc
|
||||
modelRegistry,
|
||||
toolRegistry: new Map(tools.map(tool => [tool.name, tool])),
|
||||
});
|
||||
const harness = { session, authStorage, tempDir };
|
||||
const harness = { session, authStorage, tempDir, extraSessions: [] };
|
||||
activeHarnesses.push(harness);
|
||||
return { ...harness, mock };
|
||||
}
|
||||
@@ -115,6 +120,16 @@ function expectLastAssistant(messages: AgentMessage[]): AssistantMessage {
|
||||
if (message?.role !== "assistant") throw new Error("Expected last message to be assistant");
|
||||
return message;
|
||||
}
|
||||
function createToolSession(overrides: Partial<ToolSession> = {}): ToolSession {
|
||||
return {
|
||||
cwd: "/tmp/test",
|
||||
hasUI: true,
|
||||
getSessionFile: () => null,
|
||||
getSessionSpawns: () => "*",
|
||||
settings: Settings.isolated(),
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
describe("AgentSession checkpoint rewind branch context", () => {
|
||||
it("rebuilds active history through branch_summary before the post-rewind assistant turn", async () => {
|
||||
@@ -153,6 +168,13 @@ describe("AgentSession checkpoint rewind branch context", () => {
|
||||
);
|
||||
expect(summaryIndex).toBeGreaterThan(-1);
|
||||
expect(reportIndex).toBeGreaterThan(summaryIndex);
|
||||
const reportMessage = finalCall.context.messages[reportIndex];
|
||||
if (!reportMessage) throw new Error("Expected rewind report context");
|
||||
const reportText = messageText(reportMessage);
|
||||
expect(reportText).toContain("Checkpoint completed.");
|
||||
expect(reportText).toContain("Do not call `rewind` again");
|
||||
expect(reportText).toContain(report);
|
||||
|
||||
expect(
|
||||
finalCall.context.messages.some(message => message.role === "toolResult" && message.toolName === "rewind"),
|
||||
).toBe(false);
|
||||
@@ -166,4 +188,87 @@ describe("AgentSession checkpoint rewind branch context", () => {
|
||||
expect(finalThinking?.thinking).toBe("answer after rewind");
|
||||
expect(finalThinking?.thinkingSignature).toBe("sig_after_rewind");
|
||||
});
|
||||
|
||||
it("rehydrates completed rewind state from the retained report on resume", async () => {
|
||||
const report = "findings: retained after resume";
|
||||
const harness = await createHarness([
|
||||
{
|
||||
content: [{ type: "toolCall", id: "call_checkpoint", name: "checkpoint", arguments: { goal: "inspect" } }],
|
||||
stopReason: "toolUse",
|
||||
},
|
||||
{
|
||||
content: [{ type: "toolCall", id: "call_rewind", name: "rewind", arguments: { report } }],
|
||||
stopReason: "toolUse",
|
||||
},
|
||||
{
|
||||
content: ["DONE"],
|
||||
stopReason: "stop",
|
||||
},
|
||||
]);
|
||||
|
||||
await harness.session.prompt("investigate with a checkpoint");
|
||||
|
||||
const reloadedMock = createMockModel({ responses: [] });
|
||||
const reloadedSettings = Settings.isolated({
|
||||
"compaction.enabled": false,
|
||||
"retry.enabled": false,
|
||||
"todo.enabled": false,
|
||||
"todo.eager": "default",
|
||||
"todo.reminders": false,
|
||||
});
|
||||
reloadedSettings.setModelRole("default", `${reloadedMock.provider}/${reloadedMock.id}`);
|
||||
const reloadedTools = [checkpointTool as AgentTool, rewindTool as AgentTool];
|
||||
const reloadedAgent = new Agent({
|
||||
getApiKey: () => "test-key",
|
||||
initialState: {
|
||||
model: reloadedMock,
|
||||
systemPrompt: ["Test"],
|
||||
tools: reloadedTools,
|
||||
messages: harness.session.sessionManager.buildSessionContext().messages,
|
||||
},
|
||||
convertToLlm,
|
||||
streamFn: reloadedMock.stream,
|
||||
});
|
||||
const reloadedSession = new AgentSession({
|
||||
agent: reloadedAgent,
|
||||
sessionManager: harness.session.sessionManager,
|
||||
settings: reloadedSettings,
|
||||
modelRegistry: new ModelRegistry(
|
||||
harness.authStorage,
|
||||
path.join(harness.tempDir.path(), "models-reloaded.yml"),
|
||||
),
|
||||
toolRegistry: new Map(reloadedTools.map(tool => [tool.name, tool])),
|
||||
});
|
||||
harness.extraSessions.push(reloadedSession);
|
||||
|
||||
expect(reloadedSession.getLastCompletedRewind()).toEqual({
|
||||
report,
|
||||
startedAt: "2026-01-01T00:00:00.000Z",
|
||||
rewoundAt: expect.any(String),
|
||||
});
|
||||
const tool = new RewindTool(
|
||||
createToolSession({
|
||||
getLastCompletedRewind: () => reloadedSession.getLastCompletedRewind(),
|
||||
}),
|
||||
);
|
||||
await expect(tool.execute("repeat_rewind", { report: "retry" })).rejects.toThrow(
|
||||
"Checkpoint already completed; continue from the retained rewind report instead of calling rewind again.",
|
||||
);
|
||||
});
|
||||
|
||||
it("tells the model to continue when rewind is repeated after completion", async () => {
|
||||
const tool = new RewindTool(
|
||||
createToolSession({
|
||||
getLastCompletedRewind: () => ({
|
||||
report: "findings retained",
|
||||
startedAt: "2026-01-01T00:00:00.000Z",
|
||||
rewoundAt: "2026-01-01T00:01:00.000Z",
|
||||
}),
|
||||
}),
|
||||
);
|
||||
|
||||
await expect(tool.execute("repeat_rewind", { report: "retry" })).rejects.toThrow(
|
||||
"Checkpoint already completed; continue from the retained rewind report instead of calling rewind again.",
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,154 +0,0 @@
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "bun:test";
|
||||
import * as path from "node:path";
|
||||
import { scheduler } from "node:timers/promises";
|
||||
import { Agent } from "@oh-my-pi/pi-agent-core";
|
||||
import type { Api, AssistantMessage, Model } from "@oh-my-pi/pi-ai";
|
||||
import * as AIError from "@oh-my-pi/pi-ai/error";
|
||||
import { createMockModel } from "@oh-my-pi/pi-ai/providers/mock";
|
||||
import { AssistantMessageEventStream } from "@oh-my-pi/pi-ai/utils/event-stream";
|
||||
import { ModelRegistry } from "@oh-my-pi/pi-coding-agent/config/model-registry";
|
||||
import { Settings } from "@oh-my-pi/pi-coding-agent/config/settings";
|
||||
import { AgentSession } from "@oh-my-pi/pi-coding-agent/session/agent-session";
|
||||
import { AuthStorage } from "@oh-my-pi/pi-coding-agent/session/auth-storage";
|
||||
import { SessionManager } from "@oh-my-pi/pi-coding-agent/session/session-manager";
|
||||
import { logger, TempDir } from "@oh-my-pi/pi-utils";
|
||||
|
||||
function emptyUsage(): AssistantMessage["usage"] {
|
||||
return {
|
||||
input: 0,
|
||||
output: 0,
|
||||
cacheRead: 0,
|
||||
cacheWrite: 0,
|
||||
totalTokens: 0,
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 },
|
||||
};
|
||||
}
|
||||
|
||||
function transientErrorStream(model: Model<Api>): AssistantMessageEventStream {
|
||||
const stream = new AssistantMessageEventStream();
|
||||
queueMicrotask(() => {
|
||||
const message: AssistantMessage = {
|
||||
role: "assistant",
|
||||
content: [],
|
||||
api: model.api,
|
||||
provider: model.provider,
|
||||
model: model.id,
|
||||
usage: emptyUsage(),
|
||||
stopReason: "error",
|
||||
errorMessage: "socket closed",
|
||||
errorId: AIError.create(AIError.Flag.Transient),
|
||||
timestamp: Date.now(),
|
||||
};
|
||||
stream.push({ type: "error", reason: "error", error: message });
|
||||
});
|
||||
return stream;
|
||||
}
|
||||
|
||||
function successStream(model: Model<Api>): AssistantMessageEventStream {
|
||||
const stream = new AssistantMessageEventStream();
|
||||
queueMicrotask(() => {
|
||||
const message: AssistantMessage = {
|
||||
role: "assistant",
|
||||
content: [{ type: "text", text: "recovered" }],
|
||||
api: model.api,
|
||||
provider: model.provider,
|
||||
model: model.id,
|
||||
usage: emptyUsage(),
|
||||
stopReason: "stop",
|
||||
timestamp: Date.now(),
|
||||
};
|
||||
stream.push({ type: "done", reason: "stop", message });
|
||||
});
|
||||
return stream;
|
||||
}
|
||||
|
||||
async function waitFor(predicate: () => boolean, timeoutMs = 500): Promise<void> {
|
||||
const deadline = Date.now() + timeoutMs;
|
||||
while (Date.now() < deadline) {
|
||||
if (predicate()) return;
|
||||
await Bun.sleep(1);
|
||||
}
|
||||
throw new Error("Timed out waiting for retry diagnostics");
|
||||
}
|
||||
|
||||
function isRemovalMissDebugCall(call: unknown[]): boolean {
|
||||
const payload = call[1];
|
||||
return (
|
||||
call[0] === "agent active context assistant removal missed" &&
|
||||
typeof payload === "object" &&
|
||||
payload !== null &&
|
||||
"reason" in payload &&
|
||||
payload.reason === "auto-retry"
|
||||
);
|
||||
}
|
||||
|
||||
describe("AgentSession retry diagnostics", () => {
|
||||
let tempDir: TempDir;
|
||||
let authStorage: AuthStorage;
|
||||
let session: AgentSession | undefined;
|
||||
|
||||
beforeEach(async () => {
|
||||
tempDir = TempDir.createSync("@pi-retry-diagnostics-");
|
||||
authStorage = await AuthStorage.create(path.join(tempDir.path(), "auth.db"));
|
||||
authStorage.setRuntimeApiKey("openrouter", "openrouter-test-key");
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
if (session) {
|
||||
await session.dispose();
|
||||
session = undefined;
|
||||
}
|
||||
authStorage.close();
|
||||
tempDir.removeSync();
|
||||
vi.restoreAllMocks();
|
||||
});
|
||||
|
||||
it("logs the active-context state when retry removal misses the assistant tail", async () => {
|
||||
const model = createMockModel({ provider: "openrouter", id: "glm-test" }).model;
|
||||
const calls: string[] = [];
|
||||
const agent = new Agent({
|
||||
getApiKey: requestedModel => `${requestedModel.provider}-test-key`,
|
||||
initialState: { model, systemPrompt: ["Test"], tools: [], messages: [] },
|
||||
streamFn: requestedModel => {
|
||||
calls.push(`${requestedModel.provider}/${requestedModel.id}`);
|
||||
if (calls.length === 1) return transientErrorStream(requestedModel);
|
||||
return successStream(requestedModel);
|
||||
},
|
||||
});
|
||||
const settings = Settings.isolated({
|
||||
"compaction.enabled": false,
|
||||
"retry.enabled": true,
|
||||
"retry.baseDelayMs": 0,
|
||||
"retry.maxRetries": 1,
|
||||
"retry.modelFallback": false,
|
||||
"todo.enabled": false,
|
||||
});
|
||||
settings.setModelRole("default", `${model.provider}/${model.id}`);
|
||||
session = new AgentSession({
|
||||
agent,
|
||||
sessionManager: SessionManager.inMemory(),
|
||||
settings,
|
||||
modelRegistry: new ModelRegistry(authStorage),
|
||||
});
|
||||
agent.subscribe(event => {
|
||||
if (event.type !== "message_end" || event.message.role !== "assistant") return;
|
||||
if (event.message.stopReason !== "error") return;
|
||||
const messages = agent.state.messages;
|
||||
const last = messages.at(-1);
|
||||
if (last?.role !== "assistant") return;
|
||||
agent.replaceMessages([...messages.slice(0, -1), { ...last, timestamp: last.timestamp + 1 }]);
|
||||
});
|
||||
vi.spyOn(scheduler, "wait").mockResolvedValue(undefined);
|
||||
const debugSpy = vi.spyOn(logger, "debug").mockImplementation(() => {});
|
||||
|
||||
const promptPromise = session.prompt("trigger transient error").catch(() => undefined);
|
||||
await waitFor(() => debugSpy.mock.calls.some(isRemovalMissDebugCall));
|
||||
|
||||
expect(debugSpy.mock.calls).toContainEqual([
|
||||
"agent active context assistant removal missed",
|
||||
expect.objectContaining({ reason: "auto-retry", lastRole: "assistant" }),
|
||||
]);
|
||||
await session.abort();
|
||||
await promptPromise;
|
||||
});
|
||||
});
|
||||
@@ -7,22 +7,27 @@ import { ModelRegistry } from "@oh-my-pi/pi-coding-agent/config/model-registry";
|
||||
import { Settings } from "@oh-my-pi/pi-coding-agent/config/settings";
|
||||
import { AgentSession, type AgentSessionEvent } from "@oh-my-pi/pi-coding-agent/session/agent-session";
|
||||
import { AuthStorage } from "@oh-my-pi/pi-coding-agent/session/auth-storage";
|
||||
import type { CustomMessage } from "@oh-my-pi/pi-coding-agent/session/messages";
|
||||
import { SessionManager } from "@oh-my-pi/pi-coding-agent/session/session-manager";
|
||||
import { TodoTool, type ToolSession } from "@oh-my-pi/pi-coding-agent/tools";
|
||||
import { TempDir } from "@oh-my-pi/pi-utils";
|
||||
|
||||
/**
|
||||
* Regression coverage for issue #3651: the only structured "reconcile your
|
||||
* todos" reminder used to fire at a text-only `agent_end`. A model running a
|
||||
* long tool-use loop therefore got no nudge until the very last turn, then
|
||||
* batch-flipped every task `done`. The contract this defends:
|
||||
* Regression coverage for issue #3651 and its redesign: the mid-run todo
|
||||
* reconciliation nudge keeps the live HUD honest during long runs, but is a
|
||||
* gentle MODEL-ONLY hint — deliberately separate from the user-visible
|
||||
* stop-time reminder ladder. The contract this defends:
|
||||
*
|
||||
* 1. After {@link MID_RUN_TODO_NUDGE_TURN_THRESHOLD} consecutive tool-use
|
||||
* turns without invoking the `todo` tool, the aside provider injects a
|
||||
* `<system-reminder>` for the next turn AND emits a `todo_reminder` event.
|
||||
* 2. Sub-threshold counts do NOT inject anything.
|
||||
* 3. Any `todo` tool call inside the run resets the counter, so an interleaved
|
||||
* todo turn keeps the nudge silent.
|
||||
* 1. Only SUCCESSFUL MUTATING tool results (bash/eval/edit/write/ast_edit)
|
||||
* tick the counter. Read-only exploration (grep/read/glob/lsp) and
|
||||
* errored results never do.
|
||||
* 2. At {@link MID_RUN_TODO_NUDGE_MUTATION_THRESHOLD} mutations without a
|
||||
* `todo` call, the aside provider injects a hidden custom message
|
||||
* (`display: false`) — NO `todo_reminder` event, nothing renders.
|
||||
* 3. A `todo` tool result resets the counter.
|
||||
* 4. At most {@link MID_RUN_TODO_NUDGE_MAX_PER_CYCLE} nudges fire per
|
||||
* prompt cycle.
|
||||
* 5. The counter update lands synchronously with the message_end emit.
|
||||
*
|
||||
* Drives the aside provider directly: the production agent loop polls it
|
||||
* between tool-use turns (mid-work boundary in `agent-loop.ts`), so calling it
|
||||
@@ -38,7 +43,9 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
let reminderEvents: Array<Extract<AgentSessionEvent, { type: "todo_reminder" }>>;
|
||||
let asideProvider: (() => AsideMessage[] | Promise<AsideMessage[]>) | undefined;
|
||||
|
||||
const THRESHOLD = 8; // mirrors MID_RUN_TODO_NUDGE_TURN_THRESHOLD
|
||||
const THRESHOLD = 12; // mirrors MID_RUN_TODO_NUDGE_MUTATION_THRESHOLD
|
||||
const MAX_PER_CYCLE = 2; // mirrors MID_RUN_TODO_NUDGE_MAX_PER_CYCLE
|
||||
const NUDGE_TYPE = "mid-run-todo-nudge"; // mirrors MID_RUN_TODO_NUDGE_MESSAGE_TYPE
|
||||
|
||||
function toolUseAssistant(toolName: string): AssistantMessage {
|
||||
const id = `call_${toolName}_${Date.now()}_${Math.random()}`;
|
||||
@@ -62,10 +69,6 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
};
|
||||
}
|
||||
|
||||
function emitToolUseTurn(toolName: string): void {
|
||||
session.agent.emitExternalEvent({ type: "message_end", message: toolUseAssistant(toolName) });
|
||||
}
|
||||
|
||||
function textOnlyAssistant(): AssistantMessage {
|
||||
return {
|
||||
role: "assistant",
|
||||
@@ -85,6 +88,7 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
timestamp: Date.now(),
|
||||
};
|
||||
}
|
||||
|
||||
async function emitTextOnlyStop(): Promise<void> {
|
||||
const msg = textOnlyAssistant();
|
||||
session.agent.emitExternalEvent({ type: "message_end", message: msg });
|
||||
@@ -92,9 +96,10 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
session.agent.emitExternalEvent({ type: "agent_end", messages: [msg] });
|
||||
}
|
||||
|
||||
function emitToolResult(toolName: string): void {
|
||||
/** Production-shaped tool round trip: assistant toolCall turn + toolResult. */
|
||||
function emitToolResult(toolName: string, opts?: { isError?: boolean }): void {
|
||||
const toolCallId = `call_${toolName}_${Date.now()}_${Math.random()}`;
|
||||
emitToolUseTurn(toolName);
|
||||
session.agent.emitExternalEvent({ type: "message_end", message: toolUseAssistant(toolName) });
|
||||
const content: TextContent[] = [{ type: "text", text: "ok" }];
|
||||
session.agent.emitExternalEvent({
|
||||
type: "message_end",
|
||||
@@ -103,7 +108,7 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
toolCallId,
|
||||
toolName,
|
||||
content,
|
||||
isError: false,
|
||||
isError: opts?.isError ?? false,
|
||||
timestamp: Date.now(),
|
||||
},
|
||||
});
|
||||
@@ -114,24 +119,25 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
* chain on `#messageEndPersistenceTail`. After a batch of synchronous emits
|
||||
* the counter only catches up once every queued persist task drains, so
|
||||
* tests yield a full event-loop tick before draining asides.
|
||||
*
|
||||
* Real-timer exception (ts-no-test-timers): `Bun.sleep(0)` is a single
|
||||
* event-loop tick, not a tuned duration — the private persistence tail
|
||||
* exposes no drain promise to await, and fake timers cannot flush it.
|
||||
*/
|
||||
async function settle(): Promise<void> {
|
||||
await Bun.sleep(0);
|
||||
}
|
||||
|
||||
async function drainAsides(): Promise<Array<{ role: string; text: string }>> {
|
||||
async function drainNudges(): Promise<CustomMessage[]> {
|
||||
if (!asideProvider) throw new Error("aside provider was never captured");
|
||||
const thunks = await asideProvider();
|
||||
const out: Array<{ role: string; text: string }> = [];
|
||||
const out: CustomMessage[] = [];
|
||||
for (const entry of thunks) {
|
||||
const message = typeof entry === "function" ? entry() : entry;
|
||||
if (!message) continue;
|
||||
if (message.role !== "developer") continue;
|
||||
const content = message.content;
|
||||
if (!Array.isArray(content)) continue;
|
||||
for (const part of content) {
|
||||
if (part.type === "text") out.push({ role: message.role, text: part.text });
|
||||
}
|
||||
if (message.role !== "custom") continue;
|
||||
if ((message as CustomMessage).customType !== NUDGE_TYPE) continue;
|
||||
out.push(message as CustomMessage);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
@@ -213,51 +219,73 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
vi.restoreAllMocks();
|
||||
});
|
||||
|
||||
it("stays silent until the threshold of non-todo tool-use turns is reached", async () => {
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolUseTurn("edit");
|
||||
it("read-only exploration never ticks the counter, no matter how long", async () => {
|
||||
for (let i = 0; i < THRESHOLD * 3; i++) emitToolResult(i % 2 === 0 ? "grep" : "read");
|
||||
|
||||
await settle();
|
||||
const messages = await drainAsides();
|
||||
expect(messages).toEqual([]);
|
||||
expect(await drainNudges()).toEqual([]);
|
||||
expect(reminderEvents).toEqual([]);
|
||||
});
|
||||
|
||||
it("injects a developer-role reminder once the threshold is reached", async () => {
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolUseTurn("edit");
|
||||
it("stays silent below the mutation threshold", async () => {
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolResult("edit");
|
||||
|
||||
await settle();
|
||||
const messages = await drainAsides();
|
||||
expect(messages.length).toBe(1);
|
||||
const text = messages[0]?.text ?? "";
|
||||
expect(text).toContain("<system-reminder>");
|
||||
// Surfaces every incomplete task by content, not just a count.
|
||||
expect(text).toContain("Sweep call sites");
|
||||
expect(text).toContain("Update tests");
|
||||
expect(text).toContain("Polish docs");
|
||||
// Carries the mid-run framing so the agent does not treat it as a stop-time prompt.
|
||||
expect(text).toContain("Mid-run reminder 1/3");
|
||||
expect(await drainNudges()).toEqual([]);
|
||||
expect(reminderEvents).toEqual([]);
|
||||
});
|
||||
|
||||
expect(reminderEvents.length).toBe(1);
|
||||
expect(reminderEvents[0]?.attempt).toBe(1);
|
||||
expect(reminderEvents[0]?.maxAttempts).toBe(3);
|
||||
expect(reminderEvents[0]?.todos.length).toBe(3);
|
||||
it("injects a hidden custom nudge at the threshold — no event, no render", async () => {
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolResult("edit");
|
||||
|
||||
await settle();
|
||||
const nudges = await drainNudges();
|
||||
expect(nudges.length).toBe(1);
|
||||
const nudge = nudges[0];
|
||||
// Hidden from the TUI/transcript, visible to the model only.
|
||||
expect(nudge?.display).toBe(false);
|
||||
const text = typeof nudge?.content === "string" ? nudge.content : "";
|
||||
expect(text).toContain("<system-reminder>");
|
||||
expect(text).toContain("3 todo items");
|
||||
// Gentle hint, not the stop-time escalation ladder: no per-task
|
||||
// enumeration, no attempt counter.
|
||||
expect(text).not.toContain("Sweep call sites");
|
||||
expect(text).not.toMatch(/reminder \d\/\d/i);
|
||||
|
||||
// SEPARATE concept from the stop-time reminder: no todo_reminder event,
|
||||
// so nothing renders a TodoReminderComponent or reaches extensions.
|
||||
expect(reminderEvents).toEqual([]);
|
||||
|
||||
// Counter reset: another full runway is required before the next nudge,
|
||||
// so an immediate poll right after firing must NOT re-inject.
|
||||
const followUp = await drainAsides();
|
||||
expect(followUp).toEqual([]);
|
||||
expect(await drainNudges()).toEqual([]);
|
||||
});
|
||||
|
||||
it("errored mutating results do not tick the counter", async () => {
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolResult("bash", { isError: true });
|
||||
|
||||
await settle();
|
||||
expect(await drainNudges()).toEqual([]);
|
||||
});
|
||||
|
||||
it("does not nudge when a `todo` call has reset the counter mid-window", async () => {
|
||||
// Seven non-todo turns get us within one of the threshold...
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolUseTurn("edit");
|
||||
// ...then a todo call resets the counter; the remaining runway is fresh.
|
||||
emitToolUseTurn("todo");
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolUseTurn("edit");
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolResult("write");
|
||||
emitToolResult("todo");
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolResult("write");
|
||||
|
||||
await settle();
|
||||
const messages = await drainAsides();
|
||||
expect(messages).toEqual([]);
|
||||
expect(await drainNudges()).toEqual([]);
|
||||
expect(reminderEvents).toEqual([]);
|
||||
});
|
||||
|
||||
it("caps nudges per prompt cycle", async () => {
|
||||
let fired = 0;
|
||||
for (let cycle = 0; cycle < MAX_PER_CYCLE + 2; cycle++) {
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolResult("edit");
|
||||
await settle();
|
||||
fired += (await drainNudges()).length;
|
||||
}
|
||||
expect(fired).toBe(MAX_PER_CYCLE);
|
||||
expect(reminderEvents).toEqual([]);
|
||||
});
|
||||
|
||||
@@ -265,57 +293,51 @@ describe("AgentSession mid-run todo reconciliation nudge", () => {
|
||||
// Regression for the review on PR #3652: pre-fix the counter update sat
|
||||
// after `await messageEndPersistence.persist(...)`, so the live counter
|
||||
// only caught up once microtasks drained. A poll between the emit burst
|
||||
// and the persistence chain settling would observe stale state — a turn
|
||||
// that JUST flipped a todo could still trip the nudge against the
|
||||
// pre-reset counter. With the hoisted (synchronous) update, the
|
||||
// production-shaped contract holds even when the aside poll runs in the
|
||||
// same JS task as the emit, before any microtask gets a chance to fire.
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolUseTurn("edit");
|
||||
// and the persistence chain settling would observe stale state. With the
|
||||
// hoisted (synchronous) update, the production-shaped contract holds even
|
||||
// when the aside poll runs in the same JS task as the emit.
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolResult("edit");
|
||||
|
||||
if (!asideProvider) throw new Error("aside provider was never captured");
|
||||
const result = asideProvider();
|
||||
if (result instanceof Promise) throw new Error("aside provider unexpectedly returned a Promise");
|
||||
const messagesAfterThreshold = result
|
||||
const nudges = result
|
||||
.map(entry => (typeof entry === "function" ? entry() : entry))
|
||||
.filter((m): m is NonNullable<typeof m> => Boolean(m))
|
||||
.filter(m => m.role === "developer");
|
||||
// The threshold-hit fire is the proof point: pre-hoist, the eight
|
||||
// increments are all queued microtasks, so this sync poll would see
|
||||
// counter=0 and skip the nudge entirely.
|
||||
expect(messagesAfterThreshold.length).toBe(1);
|
||||
.filter(m => m.role === "custom" && (m as CustomMessage).customType === NUDGE_TYPE);
|
||||
expect(nudges.length).toBe(1);
|
||||
});
|
||||
|
||||
it("stays silent when `todo` is not in the active-tool list, even if `todo.enabled` is still on", async () => {
|
||||
// Regression for the review on PR #3652: an explicit active-tool list
|
||||
// (or discovery-mode filtering) can drop `todo` from the slate while the
|
||||
// setting flag stays true and an incomplete persisted/user-edited todo
|
||||
// list survives. Asking the model to call a tool that is not in its
|
||||
// schema would produce fabricated/unknown tool calls or loop on
|
||||
// impossible reminders. Mirror {@link #createEagerTodoPrelude}'s guard.
|
||||
// An explicit active-tool list (or discovery-mode filtering) can drop
|
||||
// `todo` from the slate while the setting flag stays true. Asking the
|
||||
// model to call a tool that is not in its schema would produce
|
||||
// fabricated/unknown tool calls. Mirror {@link #createEagerTodoPrelude}.
|
||||
await session.setActiveToolsByName([]);
|
||||
expect(session.getActiveToolNames()).not.toContain("todo");
|
||||
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolUseTurn("edit");
|
||||
for (let i = 0; i < THRESHOLD; i++) emitToolResult("edit");
|
||||
await settle();
|
||||
const messages = await drainAsides();
|
||||
expect(messages).toEqual([]);
|
||||
expect(await drainNudges()).toEqual([]);
|
||||
expect(reminderEvents).toEqual([]);
|
||||
});
|
||||
|
||||
it("does not spend the pre-stop tool-turn count immediately after a stop-time reminder", async () => {
|
||||
it("does not spend the pre-stop mutation count immediately after a stop-time reminder", async () => {
|
||||
vi.spyOn(session.agent, "continue").mockResolvedValue();
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolUseTurn("edit");
|
||||
for (let i = 0; i < THRESHOLD - 1; i++) emitToolResult("edit");
|
||||
|
||||
await settle();
|
||||
await emitTextOnlyStop();
|
||||
await session.waitForIdle();
|
||||
// The stop-time path is the user-visible ladder: it emits the event.
|
||||
expect(reminderEvents.length).toBe(1);
|
||||
expect(reminderEvents[0]?.attempt).toBe(1);
|
||||
|
||||
// The stop-time reminder reset the mutation counter, so one more landed
|
||||
// mutation (crossing the stale pre-reminder threshold) must stay silent.
|
||||
emitToolResult("edit");
|
||||
await settle();
|
||||
const messages = await drainAsides();
|
||||
expect(messages).toEqual([]);
|
||||
expect(await drainNudges()).toEqual([]);
|
||||
expect(reminderEvents.length).toBe(1);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,106 +0,0 @@
|
||||
import { beforeAll, describe, expect, it } from "bun:test";
|
||||
import type { SegmentContext } from "@oh-my-pi/pi-coding-agent/modes/components/status-line/segments";
|
||||
import { renderSegment } from "@oh-my-pi/pi-coding-agent/modes/components/status-line/segments";
|
||||
import { initTheme, theme } from "@oh-my-pi/pi-coding-agent/modes/theme/theme";
|
||||
|
||||
beforeAll(async () => {
|
||||
await initTheme();
|
||||
});
|
||||
|
||||
function createCtx(usage: Partial<SegmentContext["usageStats"]>): SegmentContext {
|
||||
return {
|
||||
session: {
|
||||
state: {},
|
||||
isFastModeEnabled: () => false,
|
||||
modelRegistry: { isUsingOAuth: () => false },
|
||||
sessionManager: undefined,
|
||||
} as unknown as SegmentContext["session"],
|
||||
width: 120,
|
||||
compactThinkingLevel: false,
|
||||
options: {},
|
||||
planMode: null,
|
||||
loopMode: null,
|
||||
goalMode: null,
|
||||
collab: null,
|
||||
usageStats: {
|
||||
input: 0,
|
||||
output: 0,
|
||||
cacheRead: 0,
|
||||
cacheWrite: 0,
|
||||
premiumRequests: 0,
|
||||
cost: 0,
|
||||
tokensPerSecond: null,
|
||||
...usage,
|
||||
},
|
||||
contextPercent: 0,
|
||||
contextTokens: 0,
|
||||
contextWindow: 0,
|
||||
autoCompactEnabled: false,
|
||||
subagentCount: 0,
|
||||
activeMs: 0,
|
||||
activeRepo: null,
|
||||
worktree: null,
|
||||
git: {
|
||||
branch: null,
|
||||
status: null,
|
||||
pr: null,
|
||||
},
|
||||
usage: null,
|
||||
};
|
||||
}
|
||||
|
||||
describe("issue #953 cache status line icons", () => {
|
||||
it("renders cache reads as cache output and cache writes as cache input", () => {
|
||||
const cacheRead = renderSegment("cache_read", createCtx({ cacheRead: 28_919_910 }));
|
||||
const cacheWrite = renderSegment("cache_write", createCtx({ cacheWrite: 1_759_992 }));
|
||||
|
||||
expect(cacheRead.visible).toBe(true);
|
||||
expect(cacheRead.content).toContain(theme.icon.cache);
|
||||
expect(cacheRead.content).toContain(theme.icon.output);
|
||||
expect(cacheRead.content).not.toContain(theme.icon.input);
|
||||
|
||||
expect(cacheWrite.visible).toBe(true);
|
||||
expect(cacheWrite.content).toContain(theme.icon.cache);
|
||||
expect(cacheWrite.content).toContain(theme.icon.input);
|
||||
expect(cacheWrite.content).not.toContain(theme.icon.output);
|
||||
});
|
||||
});
|
||||
|
||||
describe("cache_hit segment", () => {
|
||||
it("shows hit rate from cacheRead / (cacheRead + cacheWrite) when cacheWrite > 0", () => {
|
||||
const segment = renderSegment("cache_hit", createCtx({ cacheRead: 7_500, cacheWrite: 2_500 }));
|
||||
|
||||
expect(segment.visible).toBe(true);
|
||||
expect(segment.content).toContain(theme.icon.cache);
|
||||
// 7500 / (7500 + 2500) = 75%
|
||||
expect(segment.content).toContain("75.00%");
|
||||
});
|
||||
|
||||
it("shows hit rate from cacheRead / (cacheRead + input) when cacheWrite = 0 (DeepSeek fallback)", () => {
|
||||
// DeepSeek: cacheWrite=0, input=miss_tokens
|
||||
const segment = renderSegment("cache_hit", createCtx({ cacheRead: 6_000, cacheWrite: 0, input: 4_000 }));
|
||||
|
||||
expect(segment.visible).toBe(true);
|
||||
// 6000 / (6000 + 4000) = 60%
|
||||
expect(segment.content).toContain("60.00%");
|
||||
});
|
||||
|
||||
it("shows 100% when all input was cached (cacheWrite=0, input=0)", () => {
|
||||
const segment = renderSegment("cache_hit", createCtx({ cacheRead: 5_000, cacheWrite: 0, input: 0 }));
|
||||
|
||||
expect(segment.visible).toBe(true);
|
||||
expect(segment.content).toContain("100.00%");
|
||||
});
|
||||
|
||||
it("is hidden when cacheRead is 0", () => {
|
||||
const segment = renderSegment("cache_hit", createCtx({ cacheRead: 0, cacheWrite: 5_000 }));
|
||||
|
||||
expect(segment.visible).toBe(false);
|
||||
});
|
||||
|
||||
it("is hidden when there is no cache activity at all", () => {
|
||||
const segment = renderSegment("cache_hit", createCtx({ cacheRead: 0, cacheWrite: 0, input: 1_000 }));
|
||||
|
||||
expect(segment.visible).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -773,6 +773,147 @@ describe("ModelRegistry runtime discovery", () => {
|
||||
expect(registry.find("llama.cpp", "ctx-train")?.contextWindow).toBe(65536);
|
||||
expect(registry.find("llama.cpp", "unloaded")?.contextWindow).toBe(128000);
|
||||
});
|
||||
test("llama.cpp router discovery reads --ctx-size from each preset's status.args and status.preset", async () => {
|
||||
// llama-server in router mode advertises each preset via /v1/models but
|
||||
// meta.n_ctx / n_ctx_train are only populated after the child instance
|
||||
// loads. Router-level /props returns a dummy n_ctx: 0. Without the
|
||||
// status.args / status.preset fallbacks every preset would collapse to
|
||||
// the 128k global default (issue #4190).
|
||||
const fetchMock: FetchImpl = async input => {
|
||||
const url = String(input);
|
||||
if (url === "http://127.0.0.1:8080/models") {
|
||||
return new Response(
|
||||
JSON.stringify({
|
||||
object: "list",
|
||||
data: [
|
||||
{
|
||||
id: "long-preset",
|
||||
object: "model",
|
||||
status: {
|
||||
value: "unloaded",
|
||||
args: ["--model", "/models/l.gguf", "--ctx-size", "65536"],
|
||||
preset: "[long-preset]\nmodel = /models/l.gguf\nctx-size = 65536\n\n",
|
||||
},
|
||||
source: "preset",
|
||||
},
|
||||
{
|
||||
id: "short-preset",
|
||||
object: "model",
|
||||
status: {
|
||||
value: "unloaded",
|
||||
args: ["--model", "/models/s.gguf", "-c", "8192"],
|
||||
},
|
||||
source: "preset",
|
||||
},
|
||||
{
|
||||
id: "ini-only-preset",
|
||||
object: "model",
|
||||
status: {
|
||||
value: "unloaded",
|
||||
preset: "[ini-only-preset]\nmodel = /models/i.gguf\nctx-size = 32768\n\n",
|
||||
},
|
||||
source: "preset",
|
||||
},
|
||||
{
|
||||
id: "explicit-model-default",
|
||||
object: "model",
|
||||
// --ctx-size 0 means "loaded from model"; must NOT surface as 0.
|
||||
status: {
|
||||
value: "unloaded",
|
||||
args: ["--model", "/models/d.gguf", "--ctx-size", "0"],
|
||||
},
|
||||
source: "preset",
|
||||
},
|
||||
],
|
||||
}),
|
||||
{ status: 200, headers: { "Content-Type": "application/json" } },
|
||||
);
|
||||
}
|
||||
if (url === "http://127.0.0.1:8080/props") {
|
||||
// Verbatim shape of get_router_props() — n_ctx: 0 dummy.
|
||||
return new Response(
|
||||
JSON.stringify({
|
||||
role: "router",
|
||||
max_instances: 4,
|
||||
models_autoload: true,
|
||||
model_alias: "llama-server",
|
||||
model_path: "none",
|
||||
default_generation_settings: { params: {}, n_ctx: 0 },
|
||||
}),
|
||||
{ status: 200, headers: { "Content-Type": "application/json" } },
|
||||
);
|
||||
}
|
||||
throw new Error(`Unexpected URL: ${url}`);
|
||||
};
|
||||
const registry = new ModelRegistry(authStorage, modelsJsonPath, { fetch: fetchMock });
|
||||
await registry.refresh();
|
||||
expect(registry.find("llama.cpp", "long-preset")?.contextWindow).toBe(65536);
|
||||
expect(registry.find("llama.cpp", "short-preset")?.contextWindow).toBe(8192);
|
||||
expect(registry.find("llama.cpp", "ini-only-preset")?.contextWindow).toBe(32768);
|
||||
// `--ctx-size 0` falls through past the configured hint to the global default.
|
||||
expect(registry.find("llama.cpp", "explicit-model-default")?.contextWindow).toBe(128000);
|
||||
});
|
||||
|
||||
test("llama.cpp router preset refresh honors --ctx-size when the child hasn't been loaded yet", async () => {
|
||||
// Reporter's workflow: `/model` picks a preset. On its very first switch
|
||||
// the child hasn't been spawned yet (meta.n_ctx absent), but the
|
||||
// configured window is still what the user wants surfaced.
|
||||
writeModelCache(
|
||||
"llama.cpp",
|
||||
Date.now(),
|
||||
[
|
||||
buildModel({
|
||||
id: "cold-preset",
|
||||
name: "cold-preset",
|
||||
provider: "llama.cpp",
|
||||
api: "openai-responses",
|
||||
baseUrl: "http://127.0.0.1:8080",
|
||||
reasoning: false,
|
||||
input: ["text"],
|
||||
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
||||
contextWindow: 128000,
|
||||
maxTokens: 32768,
|
||||
}),
|
||||
],
|
||||
true,
|
||||
"",
|
||||
cacheDbPath,
|
||||
);
|
||||
const fetchMock: FetchImpl = async input => {
|
||||
const url = String(input);
|
||||
if (url === "http://127.0.0.1:8080/models") {
|
||||
return new Response(
|
||||
JSON.stringify({
|
||||
data: [
|
||||
{
|
||||
id: "cold-preset",
|
||||
status: {
|
||||
value: "unloaded",
|
||||
args: ["--model", "/models/c.gguf", "--ctx-size", "16384"],
|
||||
},
|
||||
},
|
||||
],
|
||||
}),
|
||||
{ status: 200, headers: { "Content-Type": "application/json" } },
|
||||
);
|
||||
}
|
||||
if (url === "http://127.0.0.1:8080/props") {
|
||||
return new Response(JSON.stringify({ default_generation_settings: { params: {}, n_ctx: 0 } }), {
|
||||
status: 200,
|
||||
headers: { "Content-Type": "application/json" },
|
||||
});
|
||||
}
|
||||
throw new Error(`Unexpected URL: ${url}`);
|
||||
};
|
||||
const registry = new ModelRegistry(authStorage, modelsJsonPath, { fetch: fetchMock });
|
||||
const stale = registry.find("llama.cpp", "cold-preset");
|
||||
if (!stale) throw new Error("cached llama.cpp model missing");
|
||||
expect(stale.contextWindow).toBe(128000);
|
||||
const refreshed = await registry.refreshSelectedModelMetadata(stale);
|
||||
expect(refreshed.contextWindow).toBe(16384);
|
||||
expect(refreshed.maxTokens).toBe(16384);
|
||||
expect(registry.find("llama.cpp", "cold-preset")?.contextWindow).toBe(16384);
|
||||
});
|
||||
|
||||
test("llama.cpp selected model refresh patches newly loaded meta n_ctx and unlimited output limit", async () => {
|
||||
writeModelCache(
|
||||
|
||||
@@ -307,6 +307,47 @@ describe("UiHelpers.renderInitialMessages — image replay", () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe("UiHelpers.renderSessionContext — error-stop tool calls", () => {
|
||||
it("keeps the synthetic assistant error result instead of replaying a later tool result", async () => {
|
||||
await Settings.init({ inMemory: true });
|
||||
const transcript = transcriptWith([
|
||||
{
|
||||
role: "assistant",
|
||||
content: [
|
||||
{
|
||||
type: "toolCall",
|
||||
id: "error-tool",
|
||||
name: "eval",
|
||||
arguments: { language: "py", code: "raise RuntimeError('boom')" },
|
||||
},
|
||||
],
|
||||
api: "anthropic-messages",
|
||||
provider: "anthropic",
|
||||
model: "claude-sonnet",
|
||||
usage: emptyUsage,
|
||||
stopReason: "error",
|
||||
errorMessage: "synthetic assistant stop error",
|
||||
timestamp: 1,
|
||||
},
|
||||
{
|
||||
role: "toolResult",
|
||||
toolCallId: "error-tool",
|
||||
toolName: "eval",
|
||||
content: [{ type: "text", text: "late tool result must not replace the assistant stop error" }],
|
||||
isError: false,
|
||||
timestamp: 2,
|
||||
},
|
||||
]);
|
||||
const { ctx, chatContainer } = makeRenderCtx(transcript);
|
||||
|
||||
new UiHelpers(ctx).renderInitialMessages();
|
||||
|
||||
const rendered = Bun.stripANSI(chatContainer.render(120).join("\n"));
|
||||
expect(rendered).toContain("synthetic assistant stop error");
|
||||
expect(rendered).not.toContain("late tool result must not replace the assistant stop error");
|
||||
});
|
||||
});
|
||||
|
||||
describe("UiHelpers.renderSessionContext — mid-stream tool call rebuild", () => {
|
||||
it("decodes streamed write content from partialJson, not the provider's stale parsed arguments", async () => {
|
||||
// A transcript rebuild (theme change, settings edit, focus replay) can land
|
||||
|
||||
@@ -4,10 +4,10 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed an issue in the mobile collaboration web UI where 'ask' questions were displayed without response controls.
|
||||
- Fixed a host re-send of the same editor 'ask' request clobbering a half-typed guest response; the draft now resets only when a new request arrives.
|
||||
- Fixed the agent transcript drawer hot-retrying forever when the host reports a terminal transcript error (such as an oversized row); the error now stops polling and is shown below any rows already loaded.
|
||||
- Fixed pre-welcome host `error` frames (such as a protocol-version rejection) being invisible until the welcome timeout; the session now ends immediately with the host's reason.
|
||||
- Fixed missing response controls for "ask" questions in the mobile collaboration web UI.
|
||||
- Fixed an issue where re-sending the same editor "ask" request would clear a guest's in-progress draft response.
|
||||
- Fixed infinite retry loops in the agent transcript drawer when encountering terminal errors, ensuring the error is displayed and polling stops.
|
||||
- Fixed a delay in displaying pre-welcome connection errors (such as protocol version rejections), allowing the session to terminate immediately with the host's error reason.
|
||||
|
||||
## [16.2.0] - 2026-06-27
|
||||
|
||||
@@ -77,16 +77,6 @@
|
||||
- Fixed mobile layout issues where the entire chat flow would overflow horizontally and text was rendered too large on iOS Safari (by setting `text-size-adjust: 100%`)
|
||||
- Made transcript rows stack vertically on small screens to optimize reading space, and prevented grid track expansion
|
||||
- Hid non-essential metadata (such as the model name, thinking level, and working directory path) and context gauge tracks on mobile headers to prevent overflow
|
||||
- Fixed mobile layout issues where the entire chat flow would overflow horizontally and text was rendered too large on iOS Safari (by setting `text-size-adjust: 100%`)
|
||||
- Made transcript rows stack vertically on small screens to optimize reading space, and prevented grid track expansion
|
||||
- Hid non-essential metadata (such as the model name, thinking level, and working directory path) and context gauge tracks on mobile headers to prevent overflow
|
||||
- Wrapped composer button labels to display icon-only on mobile devices for a more compact and readable layout
|
||||
- Made the connect screen, ended session card, and notification toasts fully responsive for smaller device viewports
|
||||
- Fixed mobile layout issues where the entire chat flow would overflow horizontally and text was rendered too large on iOS Safari (by setting `text-size-adjust: 100%`)
|
||||
- Made transcript rows stack vertically on small screens to optimize reading space, and prevented grid track expansion
|
||||
- Hid non-essential metadata (such as the model name, thinking level, and working directory path) and context gauge tracks on mobile headers to prevent overflow
|
||||
- Wrapped composer button labels to display icon-only on mobile devices for a more compact and readable layout
|
||||
- Made the connect screen, ended session card, and notification toasts fully responsive for smaller device viewports
|
||||
|
||||
## [15.13.1] - 2026-06-15
|
||||
|
||||
@@ -100,7 +90,17 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed mobile layout issues where the entire chat flow would overflow horizontally and text was rendered too large on iOS Safari (by setting `text-size-adjust: 100%`)
|
||||
- Pinned the app shell grid to a single `minmax(0, 1fr)` column so a long session title can no longer set a min-content floor that pushes the header, transcript, and composer wider than narrow or in-app mobile viewports; the title now ellipsizes instead of clipping every row's right edge
|
||||
- Made transcript rows stack vertically on small screens to optimize reading space, and prevented grid track expansion
|
||||
- Hid non-essential metadata (such as the model name, thinking level, and working directory path) and context gauge tracks on mobile headers to prevent overflow
|
||||
- Wrapped composer button labels to display icon-only on mobile devices for a more compact and readable layout
|
||||
- Made the connect screen, ended session card, and notification toasts fully responsive for smaller device viewports
|
||||
- Fixed mobile layout issues where the entire chat flow would overflow horizontally and text was rendered too large on iOS Safari (by setting `text-size-adjust: 100%`)
|
||||
- Made transcript rows stack vertically on small screens to optimize reading space, and prevented grid track expansion
|
||||
- Hid non-essential metadata (such as the model name, thinking level, and working directory path) and context gauge tracks on mobile headers to prevent overflow
|
||||
- Wrapped composer button labels to display icon-only on mobile devices for a more compact and readable layout
|
||||
- Made the connect screen, ended session card, and notification toasts fully responsive for smaller device viewports
|
||||
|
||||
## [15.12.4] - 2026-06-13
|
||||
|
||||
|
||||
@@ -2,13 +2,15 @@
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### Changed
|
||||
|
||||
- Optimized stale-anchor remap validation from quadratic to linear complexity, significantly improving performance on large files.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed an issue where snapshot tag collisions could cause line-anchored edits to be incorrectly applied to unrelated content.
|
||||
- Fixed an issue where snapshot tag collisions could cause line-anchored edits to be incorrectly applied to unrelated content, improving recovery and edit-preview safety.
|
||||
- Fixed tracking of edit anchors when earlier in-session insertions or deletions shift unchanged target lines.
|
||||
- Fixed recovery and edit-preview paths still treating a 16-bit snapshot tag as exact identity after collision-aware retention landed: ambiguous colliding tags now fall through to recovery/rejection (`byHashExact`) instead of applying anchors against the most-recent collider.
|
||||
- Reduced stale-anchor remap validation from quadratic to linear: duplicate-line detection and anchor-neighbor context are now precomputed once per pass instead of scanning the whole file per anchor.
|
||||
- Fixed hashline edit guidance for Markdown list rows by teaching `+- item` escaping in the model prompt and minus-row parser error. ([#4179](https://github.com/can1357/oh-my-pi/issues/4179))
|
||||
- Fixed hashline edit guidance and parsing errors for Markdown list rows.
|
||||
|
||||
## [16.2.8] - 2026-06-30
|
||||
|
||||
|
||||
@@ -4,13 +4,12 @@
|
||||
|
||||
### Added
|
||||
|
||||
- Added `workingDir` to `ShellRunResult` to allow hosts to synchronize the session's current working directory without executing a hidden probe command.
|
||||
- Added workingDir to ShellRunResult to allow hosts to synchronize the session's current working directory without executing a hidden probe command.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed an issue where panics in native worker tasks (such as grep, AST parsing, globbing, workspace listing, HTML-to-markdown conversion, fuzzy finding, and clipboard image reading) would abort the host process instead of properly rejecting the returned JavaScript Promise. Panics recovered this way are recorded in the native crash log (disk only, no stderr noise) so real native bugs still leave a diagnostic artifact.
|
||||
- Fixed the blocking-task panic recovery itself aborting the host when a panic payload's own destructor panics; the message is extracted first and the payload disposed without unwinding across the FFI boundary.
|
||||
- Fixed a crash on Windows under low memory/commit charge conditions when spawning worker threads for token counting or sorting operations.
|
||||
- Fixed an issue where panics in native worker tasks (such as grep, AST parsing, globbing, workspace listing, HTML-to-markdown conversion, fuzzy finding, and clipboard image reading) would abort the host process instead of properly rejecting the returned JavaScript Promise.
|
||||
- Fixed a crash on Windows under low memory or commit charge conditions when spawning worker threads for token counting or sorting operations.
|
||||
|
||||
## [16.2.11] - 2026-07-01
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed the bounded stdin escape parser scanning past its own cap: OSC/DCS/APC terminator lookups used unbounded `indexOf`, so one oversized unterminated sequence could still block the event loop; scans are now strictly bounded to the per-sequence byte cap.
|
||||
- Fixed a potential event loop hang caused by processing oversized, unterminated terminal escape sequences (OSC/DCS/APC).
|
||||
- Fixed an issue where large Windows terminal session restores could get truncated mid-frame during ConPTY full-paint resume.
|
||||
|
||||
## [16.2.13] - 2026-07-01
|
||||
@@ -12,25 +12,6 @@
|
||||
### Fixed
|
||||
|
||||
- Fixed fuzzy-search filtering for CJK and other non-ASCII queries by preserving Unicode letters and numbers during query normalization ([#4114](https://github.com/can1357/oh-my-pi/issues/4114)).
|
||||
- Bounded terminal input parsing so a malformed CSI/OSC/DCS/APC or a large non-bracketed paste no longer blocks the event loop. `StdinBuffer.extractCompleteSequences` now resolves each escape by a single linear scan with a per-type length cap (CSI 4 KiB, OSC/DCS/APC 16 MiB) and carries a resume-search offset so a chunked OSC 5522 payload stays O(total) instead of O(total²). `BracketedPasteHandler` gained a byte cap (default 64 MiB) that aborts paste mode and delivers the accumulated bytes when a lost end marker would otherwise hold memory forever — defense in depth for callers that bypass `StdinBuffer`. The `ProcessTerminal` data handler now takes a fast path when the sequence is not ESC-prefixed and no reassembly buffer is active, so a large non-bracketed paste skips six escape-probe regex tests per Unicode scalar ([#4073](https://github.com/can1357/oh-my-pi/issues/4073)).
|
||||
- Added adaptive render backpressure: a frame that overruns the 30 fps cadence now inflates the following frame's delay to at most twice its own cost (capped at 200 ms), preventing the render loop from busy-looping when a slow paint would otherwise fire the next frame at `setTimeout(0)`. ([#4145](https://github.com/can1357/oh-my-pi/issues/4145))
|
||||
|
||||
### Added
|
||||
|
||||
- Added regression coverage for `findCommittedPrefixResync` — the tui seam that re-anchors the committed prefix when a live block re-lays-out at settle. Locks in the earliest-audited-mismatch re-anchor, the hard-scan escape from tail-sample tolerance for a newly-permanent forced-overflow row, exempt-window drift silence, and shrink-into-prefix truncation, so a future refactor of the resync path cannot silently strand pending SSH placeholder chrome above the settled block ([#4124](https://github.com/can1357/oh-my-pi/issues/4124)).
|
||||
- Added `Editor.setTopBorderProvider()` so hosts can install a lazy top-border builder that runs once per painted frame instead of eagerly rebuilding after every state change. Falls back to the existing `setTopBorder()` slot when no provider is registered.
|
||||
|
||||
## [16.2.13] - 2026-07-01
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed fuzzy-search filtering for CJK and other non-ASCII queries by preserving Unicode letters and numbers during query normalization ([#4114](https://github.com/can1357/oh-my-pi/issues/4114)).
|
||||
|
||||
## [16.2.12] - 2026-07-01
|
||||
|
||||
### Fixed
|
||||
|
||||
- Optimized streaming markdown rendering to reuse already-rendered prefix lines and only render new content deltas, improving performance and reducing redraw flicker.
|
||||
|
||||
## [16.2.12] - 2026-07-01
|
||||
|
||||
@@ -44,34 +25,12 @@
|
||||
|
||||
- Fixed mid-prompt `/skill:<name>` autocomplete acceptance wiping the user's draft. The autocomplete now inserts the `/skill:<name> ` token at the cursor (replacing only the partial `/sk` slash token) and preserves prose typed before and after it, so a user can compose a prompt and reach for a skill without losing their train of thought ([#3913](https://github.com/can1357/oh-my-pi/issues/3913)).
|
||||
|
||||
## [16.2.10] - 2026-06-30
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed mid-prompt `/skill:<name>` autocomplete acceptance wiping the user's draft. The autocomplete now inserts the `/skill:<name> ` token at the cursor (replacing only the partial `/sk` slash token) and preserves prose typed before and after it, so a user can compose a prompt and reach for a skill without losing their train of thought ([#3913](https://github.com/can1357/oh-my-pi/issues/3913)).
|
||||
|
||||
## [16.2.9] - 2026-06-30
|
||||
|
||||
### Added
|
||||
|
||||
- Added `Editor.submit()` to allow programmatic composer submission, enabling integration with speech input and other automated flows.
|
||||
|
||||
## [16.2.9] - 2026-06-30
|
||||
|
||||
### Added
|
||||
|
||||
- Added `Editor.submit()` to allow programmatic composer submission, enabling integration with speech input and other automated flows.
|
||||
|
||||
## [16.2.7] - 2026-06-30
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed an issue where a fast double-Escape keypress was swallowed and ignored, preventing double-escape gestures and subsequent Escape key handlers from firing.
|
||||
|
||||
### Changed
|
||||
|
||||
- Sped up fuzzy filtering in selectors (model, settings, file/tree, hook, OAuth) by preparing the query once per filter instead of once per candidate, and memoizing the per-text search index across keystrokes. Incremental typing over a 400-item list drops ~57% (6.7ms → 2.9ms for an 8-keystroke session) with no change to match ranking and no first-keystroke regression; long texts (pasted prompts, transcripts) bypass the cache so memory stays bounded.
|
||||
|
||||
## [16.2.7] - 2026-06-30
|
||||
|
||||
### Fixed
|
||||
@@ -637,7 +596,7 @@
|
||||
|
||||
### Changed
|
||||
|
||||
- Changed native-scrollback safety defaults to treat unknown POSIX, SSH, and multiplexer-shaped terminals as ED3-risk for passive rendering; checkpoint replay now requires a positive at-tail viewport proof instead of assuming prompt submit makes host scrollback safe ([#1799](https://github.com/can1357/oh-my-pi/issues/1799)).
|
||||
- Changed native-scrollback safety defaults to treat unknown POSIX, SSH, and multiplexer-shaped terminals as ED3-risk for passive rendering; checkpoint replay now requires a positive at-tail viewport proof instead of assuming prompt submit makes host scrollback safe.
|
||||
- Changed synchronized-output defaults to a conservative opt-in profile: DEC 2026 paint wrappers stay disabled for remote/multiplexer/VTE/unknown terminals unless explicitly forced, while the autowrap guards remain active.
|
||||
|
||||
### Fixed
|
||||
|
||||
+31
-18
@@ -2836,6 +2836,7 @@ export class TUI extends Container {
|
||||
}
|
||||
window = this.#prepareLinesArray(window, width);
|
||||
}
|
||||
const cursorTrackingLineCount = hasVisibleOverlay ? Math.max(frame.length, windowTop + height) : frame.length;
|
||||
|
||||
const intent: RenderIntent = fullPaint
|
||||
? { kind: "fullPaint", clearScrollback: replaceRequested || geometryRebuild ? !isMultiplexerSession() : false }
|
||||
@@ -2866,6 +2867,7 @@ export class TUI extends Container {
|
||||
clearScrollback: intent.clearScrollback,
|
||||
chunkTo,
|
||||
windowTop,
|
||||
cursorTrackingLineCount,
|
||||
});
|
||||
this.#committedPrefix = rawFrame.slice(0, chunkTo);
|
||||
this.#updateCommittedAuditRows(
|
||||
@@ -2892,6 +2894,7 @@ export class TUI extends Container {
|
||||
prevHardwareCursorRow,
|
||||
forceWindowRewrite: this.#forceViewportRepaintOnNextRender || (geometryChanged && resizeRepaintsInPlace()),
|
||||
repaintVirtualScrollInPlace: hasVisibleOverlay,
|
||||
cursorTrackingLineCount,
|
||||
});
|
||||
for (let i = this.#committedPrefix.length; i < chunkTo; i++) {
|
||||
this.#committedPrefix.push(rawFrame[i] ?? "");
|
||||
@@ -3273,10 +3276,15 @@ export class TUI extends Container {
|
||||
cursorPos: { row: number; col: number } | null,
|
||||
purgeSequence: string,
|
||||
imageTransmitBuffer: string,
|
||||
options: { clearScrollback: boolean; chunkTo: number; windowTop: number },
|
||||
options: {
|
||||
clearScrollback: boolean;
|
||||
chunkTo: number;
|
||||
windowTop: number;
|
||||
cursorTrackingLineCount: number;
|
||||
},
|
||||
): void {
|
||||
this.#fullRedrawCount += 1;
|
||||
const { chunkTo, windowTop } = options;
|
||||
const { chunkTo, windowTop, cursorTrackingLineCount } = options;
|
||||
// Map the frame-space cursor into paint space: committed-prefix rows
|
||||
// keep their index, visible-window rows land after the prefix, and a
|
||||
// cursor in neither region (hidden behind the overlay gap) hides.
|
||||
@@ -3376,7 +3384,9 @@ export class TUI extends Container {
|
||||
buffer += this.#paintEndSequence;
|
||||
this.terminal.write(buffer);
|
||||
|
||||
const committedCursorState = paintCursorPos ? this.#targetHardwareCursorState(cursorPos, frame.length) : null;
|
||||
const committedCursorState = paintCursorPos
|
||||
? this.#targetHardwareCursorState(cursorPos, cursorTrackingLineCount)
|
||||
: null;
|
||||
const committedCursor = committedCursorState
|
||||
? {
|
||||
toRow: committedCursorState.row,
|
||||
@@ -3646,6 +3656,7 @@ export class TUI extends Container {
|
||||
prevHardwareCursorRow: number;
|
||||
forceWindowRewrite: boolean;
|
||||
repaintVirtualScrollInPlace: boolean;
|
||||
cursorTrackingLineCount: number;
|
||||
},
|
||||
): void {
|
||||
const {
|
||||
@@ -3655,6 +3666,7 @@ export class TUI extends Container {
|
||||
prevHardwareCursorRow,
|
||||
forceWindowRewrite,
|
||||
repaintVirtualScrollInPlace,
|
||||
cursorTrackingLineCount,
|
||||
} = options;
|
||||
const chunkFrom = this.#committedRows;
|
||||
const chunkLength = chunkTo - chunkFrom;
|
||||
@@ -3706,7 +3718,7 @@ export class TUI extends Container {
|
||||
}
|
||||
cursorFromRow = windowTop + lastChanged;
|
||||
}
|
||||
const cursorControl = this.#cursorControlSequence(cursorPos, frame.length, cursorFromRow);
|
||||
const cursorControl = this.#cursorControlSequence(cursorPos, cursorTrackingLineCount, cursorFromRow);
|
||||
buffer += cursorControl.seq;
|
||||
buffer += this.#paintEndSequence;
|
||||
this.terminal.write(buffer);
|
||||
@@ -3717,16 +3729,17 @@ export class TUI extends Container {
|
||||
}
|
||||
}
|
||||
|
||||
// In-window diff: nothing commits. While an overlay is visible, commits
|
||||
// are frozen; if the underlying windowTop moves, repaint in place rather
|
||||
// than falling through to a seam rewrite that would scroll native history
|
||||
// without appending to the commit tape.
|
||||
const virtualScrollRewrite = repaintVirtualScrollInPlace && scroll !== 0;
|
||||
if (chunkLength === 0 && (scroll === 0 || virtualScrollRewrite)) {
|
||||
if (forceWindowRewrite || virtualScrollRewrite) this.#fullRedrawCount += 1;
|
||||
let firstChanged = forceWindowRewrite || virtualScrollRewrite ? 0 : -1;
|
||||
let lastChanged = forceWindowRewrite || virtualScrollRewrite ? height - 1 : -1;
|
||||
if (!forceWindowRewrite && !virtualScrollRewrite) {
|
||||
// In-window diff: nothing commits. While an overlay is visible, repaint
|
||||
// the full viewport in place from a top-clamped cursor origin. Overlay
|
||||
// cursor-only frames can leave the tracked row behind the physical cursor;
|
||||
// a relative partial rewrite from that stale origin can CRLF on the bottom
|
||||
// row and scroll native history without appending to the commit tape.
|
||||
const overlayInPlaceRewrite = repaintVirtualScrollInPlace;
|
||||
if (chunkLength === 0 && (scroll === 0 || overlayInPlaceRewrite)) {
|
||||
if (forceWindowRewrite || overlayInPlaceRewrite) this.#fullRedrawCount += 1;
|
||||
let firstChanged = forceWindowRewrite || overlayInPlaceRewrite ? 0 : -1;
|
||||
let lastChanged = forceWindowRewrite || overlayInPlaceRewrite ? height - 1 : -1;
|
||||
if (!forceWindowRewrite && !overlayInPlaceRewrite) {
|
||||
const comparable = previousWindow.length === height;
|
||||
for (let r = 0; r < height; r++) {
|
||||
if (comparable && (window[r] ?? "") === (previousWindow[r] ?? "")) continue;
|
||||
@@ -3736,13 +3749,13 @@ export class TUI extends Container {
|
||||
}
|
||||
if (firstChanged === -1) {
|
||||
if (purgeSequence.length > 0) this.terminal.write(purgeSequence);
|
||||
this.#writeCursorPosition(cursorPos, frame.length);
|
||||
this.#writeCursorPosition(cursorPos, cursorTrackingLineCount);
|
||||
this.#previousWidth = width;
|
||||
this.#previousHeight = height;
|
||||
return;
|
||||
}
|
||||
let buffer = this.#paintBeginSequence + purgeSequence;
|
||||
if (virtualScrollRewrite) {
|
||||
if (overlayInPlaceRewrite) {
|
||||
// The cursor tracker can be stale after overlay-only frames. A large
|
||||
// CUU clamps at the viewport top without using absolute cursor home,
|
||||
// so the following full-window rewrite cannot overflow the bottom.
|
||||
@@ -3777,7 +3790,7 @@ export class TUI extends Container {
|
||||
buffer += `\x1b[${lastChanged - contentBottomScreenRow}A`;
|
||||
cursorFromRow = contentBottomRow;
|
||||
}
|
||||
const cursorControl = this.#cursorControlSequence(cursorPos, frame.length, cursorFromRow);
|
||||
const cursorControl = this.#cursorControlSequence(cursorPos, cursorTrackingLineCount, cursorFromRow);
|
||||
buffer += cursorControl.seq;
|
||||
buffer += this.#paintEndSequence;
|
||||
this.terminal.write(buffer);
|
||||
@@ -3807,7 +3820,7 @@ export class TUI extends Container {
|
||||
}
|
||||
const parkUp = height - 1 - (contentBottomRow - windowTop);
|
||||
if (parkUp > 0) buffer += `\x1b[${parkUp}A`;
|
||||
const cursorControl = this.#cursorControlSequence(cursorPos, frame.length, contentBottomRow);
|
||||
const cursorControl = this.#cursorControlSequence(cursorPos, cursorTrackingLineCount, contentBottomRow);
|
||||
buffer += cursorControl.seq;
|
||||
buffer += this.#paintEndSequence;
|
||||
this.terminal.write(buffer);
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
import { afterEach, beforeEach, describe, expect, it } from "bun:test";
|
||||
import { type Component, CURSOR_MARKER, TUI } from "@oh-my-pi/pi-tui";
|
||||
import { type Component, CURSOR_MARKER, type Focusable, type OverlayFocusOwner, TUI } from "@oh-my-pi/pi-tui";
|
||||
import { VirtualTerminal } from "./virtual-terminal";
|
||||
|
||||
class LineComponent implements Component {
|
||||
@@ -64,6 +64,54 @@ class CursorOnlyComponent implements Component {
|
||||
}
|
||||
}
|
||||
|
||||
class FocusedMutableOverlay implements Component, Focusable {
|
||||
focused = false;
|
||||
#text: string;
|
||||
|
||||
constructor(text: string) {
|
||||
this.#text = text;
|
||||
}
|
||||
|
||||
setText(text: string): void {
|
||||
this.#text = text;
|
||||
}
|
||||
|
||||
invalidate(): void {
|
||||
// No cached state
|
||||
}
|
||||
|
||||
render(_width: number): string[] {
|
||||
return [`${this.#text}${this.focused ? CURSOR_MARKER : ""}`];
|
||||
}
|
||||
}
|
||||
|
||||
class OverlayFocusDelegator implements Component, OverlayFocusOwner {
|
||||
#text: string;
|
||||
|
||||
constructor(
|
||||
text: string,
|
||||
private readonly ownedFocusTarget: Component,
|
||||
) {
|
||||
this.#text = text;
|
||||
}
|
||||
|
||||
setText(text: string): void {
|
||||
this.#text = text;
|
||||
}
|
||||
|
||||
ownsOverlayFocusTarget(component: Component): boolean {
|
||||
return component === this.ownedFocusTarget;
|
||||
}
|
||||
|
||||
invalidate(): void {
|
||||
// No cached state
|
||||
}
|
||||
|
||||
render(_width: number): string[] {
|
||||
return [this.#text];
|
||||
}
|
||||
}
|
||||
|
||||
function buildRows(count: number): string[] {
|
||||
return Array.from({ length: count }, (_v, i) => `row-${i}`);
|
||||
}
|
||||
@@ -159,6 +207,52 @@ describe("TUI overlays", () => {
|
||||
expect(term.getScrollBuffer().length).toBeLessThan(200);
|
||||
});
|
||||
|
||||
it("keeps the native viewport anchored when an overlay repaint follows a focused cursor below the frame tail", async () => {
|
||||
const term = new VirtualTerminal(24, 6, 100);
|
||||
const tui = new TUI(term, true);
|
||||
const base = new MutableContentComponent(buildRows(8));
|
||||
const cursorOverlay = new FocusedMutableOverlay("overlay-cursor");
|
||||
const statusOverlay = new OverlayFocusDelegator("status-before", cursorOverlay);
|
||||
tui.addChild(base);
|
||||
|
||||
try {
|
||||
tui.start();
|
||||
await flushRender(term);
|
||||
|
||||
tui.showOverlay(cursorOverlay, { row: 5, col: 0, width: 16 });
|
||||
tui.showOverlay(statusOverlay, { row: 0, col: 0, width: 16 });
|
||||
tui.setFocus(cursorOverlay);
|
||||
tui.requestRender();
|
||||
await flushRender(term);
|
||||
|
||||
base.setLines(["base-0", "base-1"]);
|
||||
tui.requestRender();
|
||||
await flushRender(term);
|
||||
expect(term.getCursor().row).toBe(5);
|
||||
|
||||
const before = term.getBufferPosition();
|
||||
const beforeScrollBufferLength = term.getScrollBuffer().length;
|
||||
|
||||
statusOverlay.setText("status-after");
|
||||
tui.requestRender();
|
||||
await flushRender(term);
|
||||
|
||||
expect(term.getBufferPosition()).toEqual(before);
|
||||
expect(term.getScrollBuffer()).toHaveLength(beforeScrollBufferLength);
|
||||
expect(term.getViewport().map(line => line.trimEnd())).toEqual([
|
||||
"status-after",
|
||||
"base-1",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"overlay-cursor",
|
||||
]);
|
||||
expect(term.getCursor().row).toBe(5);
|
||||
} finally {
|
||||
tui.stop();
|
||||
}
|
||||
});
|
||||
|
||||
it("clamps tall overlays without an explicit maxHeight to the available rows", async () => {
|
||||
const term = new VirtualTerminal(80, 24);
|
||||
const tui = new TUI(term);
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
### Added
|
||||
|
||||
- Added `wrapFetchForExtraCa` / `withExtraCaFetch` (moved from `@oh-my-pi/pi-ai` internals): a fetch wrapper that applies `NODE_EXTRA_CA_CERTS` to Bun's `RequestInit.tls.ca`, shared by provider streaming and catalog model discovery.
|
||||
- Added `wrapFetchForExtraCa` and `withExtraCaFetch` utility functions to apply `NODE_EXTRA_CA_CERTS` to Bun's `RequestInit.tls.ca` configuration.
|
||||
|
||||
## [16.2.9] - 2026-06-30
|
||||
|
||||
@@ -113,24 +113,16 @@
|
||||
|
||||
- Added profile-aware directory helpers and isolated profile state roots, while keeping the install ID shared across profiles.
|
||||
- Added a named-profile API to the `dirs` module — `setProfile()`, `getActiveProfile()`, `getProfileRootDir()`, and `normalizeProfileName()` — plus `resolveProfileEnv()`, which selects the active profile from `OMP_PROFILE` (canonical; takes precedence) then `PI_PROFILE` (legacy fallback, consulted only when `OMP_PROFILE` is unset).
|
||||
- Added support for a runtime `overrides` map in `RuntimeInstallSpec`, which is now written into generated runtime `package.json` manifests to force dependency pins (including transitive ones) across the runtime tree
|
||||
- Added a lightweight loop-phase breadcrumb stack (`pushLoopPhase`/`popLoopPhase`/`currentLoopPhase`, plus `takeRecentLoopPhase` which returns the live phase or the most recently popped one and clears it) so the TUI event-loop watchdog can attribute a main-thread block to the phase that caused it — including a synchronous phase already popped before the watchdog's delayed tick runs ([#2485](https://github.com/can1357/oh-my-pi/issues/2485))
|
||||
- Added `FetchWithRetryOptions.timeout` (forwarded to the underlying `fetch` call). `false` disables Bun's native ~300s pre-response timeout; a positive number overrides the ceiling. Bare browser/Node fetch ignores it ([#2422](https://github.com/can1357/oh-my-pi/issues/2422))
|
||||
- Added the side-effect-free `@oh-my-pi/pi-utils/worker-host` module (`declareWorkerHostEntry()` / `workerHostEntry()`), extracted from `env` (still re-exported there) so worker spawn sites can resolve the self-dispatching CLI host entry without importing `env`'s side-effecting module graph.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed profile directory isolation when a profile's agent `.env` customizes directory roots: directory-affecting keys (`XDG_DATA_HOME`/`XDG_STATE_HOME`/`XDG_CACHE_HOME`, and a default-mode `PI_CODING_AGENT_DIR`) are now honored. The `env` loader rebuilds the `dirs` resolver after applying `.env` files (`refreshDirsFromEnv()`), so a profile `.env` that points XDG roots elsewhere no longer leaks state into the home-based config dir.
|
||||
- Fixed `installRuntimeModuleResolver()` to keep bare requests from runtime-cache modules inside that registered runtime before falling back to host/workspace packages.
|
||||
|
||||
## [15.13.1] - 2026-06-15
|
||||
|
||||
### Added
|
||||
|
||||
- Added support for a runtime `overrides` map in `RuntimeInstallSpec`, which is now written into generated runtime `package.json` manifests to force dependency pins (including transitive ones) across the runtime tree
|
||||
- Added a lightweight loop-phase breadcrumb stack (`pushLoopPhase`/`popLoopPhase`/`currentLoopPhase`, plus `takeRecentLoopPhase` which returns the live phase or the most recently popped one and clears it) so the TUI event-loop watchdog can attribute a main-thread block to the phase that caused it — including a synchronous phase already popped before the watchdog's delayed tick runs ([#2485](https://github.com/can1357/oh-my-pi/issues/2485))
|
||||
- Added `FetchWithRetryOptions.timeout` (forwarded to the underlying `fetch` call). `false` disables Bun's native ~300s pre-response timeout; a positive number overrides the ceiling. Bare browser/Node fetch ignores it ([#2422](https://github.com/can1357/oh-my-pi/issues/2422))
|
||||
|
||||
### Fixed
|
||||
|
||||
- Made `TempDir` cleanup retry transient Windows `EBUSY`/`EPERM`/`ENOTEMPTY` removal failures so tests are less likely to fail when deleting just-used temp directories.
|
||||
- Fixed `installRuntimeModuleResolver()` to keep bare requests from runtime-cache modules inside that registered runtime before falling back to host/workspace packages.
|
||||
|
||||
## [15.12.4] - 2026-06-13
|
||||
|
||||
|
||||
@@ -4,11 +4,11 @@
|
||||
|
||||
### Breaking Changes
|
||||
|
||||
- Bumped `COLLAB_PROTO` to `3`: the `ui-request`/`ui-request-end` host frames and `ui-response` guest frame are part of the handshake contract. Proto-2 guests, which silently dropped host ask requests, are now rejected at hello with the protocol-mismatch error.
|
||||
- Upgraded the collaboration protocol (COLLAB_PROTO) to version 3. Guests using version 2 are now rejected during the handshake with a protocol-mismatch error due to new interactive UI request/response requirements.
|
||||
|
||||
### Added
|
||||
|
||||
- Added collaboration UI request/response wireframes, enabling browser guests to respond to host-side interactive prompts.
|
||||
- Added collaboration UI request and response frames, enabling browser guests to respond to interactive prompts initiated by the host.
|
||||
|
||||
## [16.1.8] - 2026-06-20
|
||||
|
||||
|
||||
Reference in New Issue
Block a user