The #3063 fix introduced a second mutating step — `bun update <name>` —
that rewrites bun.lock before extension validation runs. Three failure
paths could still leave the rejected commit pinned in the lockfile or
active tree:
- Extension validation throwing after `bun update` had refreshed
bun.lock — rollback restored package.json and node_modules/<name>
but never touched bun.lock.
- Feature validation (`omp plugin install pkg[ghost]`) throwing
outside the rollback block entirely.
- Runtime-config save failing after a successful install with no
rollback path.
Snapshot bun.lock alongside package.json before `bun install` runs and
route every post-install step (resolution, update, package.json read,
feature validation, extension validation, runtime-config save) through
one outer catch that restores all three (package.json + bun.lock +
node_modules/<name> from snapshot). `#rollbackFailedInstall` now
tolerates an unresolved `actualName` for failures that throw before
the dep key is known.
Three regression tests in plugin-install-validation.test.ts pin the
new contract: bun.lock restoration after a git reinstall fails
validation, bun.lock removal when it didn't exist pre-install, and
rollback on an unknown feature request.
Addresses review feedback on #3069.
bun install <spec> respects the existing bun.lock pin when the spec is
unchanged and never re-resolves the remote ref, so re-running
`omp plugin install github:owner/repo` on an already-installed plugin
reported success while silently keeping the user on the original
resolved commit (1ms no-op, no network).
PluginManager.install now follows a git re-install with
`bun update <name>` to force re-resolution of the ref against the
upstream. First-time installs (no prior dep entry) skip the update —
the initial bun install already fetches HEAD. bun update failures
trigger the same rollback path as validation failures.
Fixes#3063
Forwarded peekPendingInvoker and clearPendingInvokers from the production toolSession literal so the resolve tool can dispatch staged previews and drain stale gates in real CLI sessions, not only in unit tests.
Added a regression test exercising the production wiring through AgentSession and a phantom-gate drain via a no-invoker facade.
Fixes#3061
Cleared pending-preview markers when resolve has no runnable handler so a stale gate cannot keep forcing resolve after the invoker is gone.
Added regression coverage for apply and discard draining stale pending markers.
Fixes#3061
- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
Kept observing a status-line usage fetch after the startup timeout so late successful reports still refresh the quota segment instead of being hidden behind the timeout backoff.\n\nFixes #3057
Deferred the status-line quota refresh off the render path and raced it against a short startup timeout so slow Anthropic usage lookups cannot pin interactive startup. Added regression coverage for non-synchronous refresh startup and timeout backoff.\n\nFixes #3057
- Tighten invalidation logic to only trigger when a demonstrably warm cache (previously read) goes cold.
- Ignore cold transitions following write-only turns to prevent spurious markers during initial cache warming or after natural TTL expiry.
- Update test suite to verify that consecutive cold turns following an initial write do not trigger invalidation alerts.
- Improved thinking loop detection logic by refining text normalization and treating stalls as retryable errors.
- Instrumented the agent session to recognize thinking loop markers within retryable error conditions.
- Automated the clearing of stale error banners upon successful auto-retry execution.
- Added comprehensive test coverage for chunked thinking loop errors and banner management.
- Moved authentication logic to `perplexity-auth.ts` to share logic between search providers and CLI commands.
- Updated authentication priority to prefer browser cookies over OAuth tokens during search operations.
- Modified the `token` CLI command to display active OAuth tokens when both an OAuth token and an API key are configured.
- Added comprehensive unit tests in `perplexity.test.ts` to verify authentication priority and precedence.
- Replaced usage of `ReturnType<typeof setTimeout>` and `ReturnType<typeof setInterval>` with the explicit `Timer` type across the codebase.
- Updated several type definitions and function signatures to use concrete types instead of inferred return types for improved clarity and maintainability.
The PR fixed the -32601 hang for the defined server->client refresh
requests but missed two real spec methods of the identical class:
workspace/inlineValue/refresh (LSP 3.17) and workspace/foldingRange/refresh.
A server emitting either still received Method not found and could stall.
Add both to the void-ack chain and extend the regression test.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Renamed the global `INTENT_FIELD` constant from `_i` to `i`.
- Updated documentation strings, type annotations, and test expectations across packages to reflect the new field name.
- Ensured consistent usage of the constant in tool schema construction and intent serialization.
- Replaced regex-based markdown detection with a stateful parser in `detectLiveReflowingMarkdown`.
- Correctly ignore table delimiters and mermaid markers when they appear inside fenced code blocks.
- Added tests to verify that code blocks containing markdown-like syntax do not prevent commit stability.
- Prevent object reference sharing between agent snapshots and stream events by deep-cloning tool-call arguments.
- Stabilize GFM tables and Mermaid diagrams during streaming by delaying transcript block commits until content finalization.
- Implement session resume safety to prevent crashes when working directories are missing.
- Add comprehensive test suites to verify streaming commit stability and immutable snapshot isolation.
- Shortened language in `browser.md`, `eval.md`, and `prompt.md` to improve clarity and reduce token consumption.
- Refined instructional phrasing throughout the tool documentation for better readability.
- Replaced O(n) `Array.shift()` calls with O(1) head-index tracking in `RawSseDebugBuffer`.
- Implemented lazy compaction to reclaim the dead-prefix memory only when the leading window grows sufficiently large.
- Added comprehensive tests to verify ordering, dropped count accuracy, and buffer limits under heavy load.
- Add `onFirstChatDispatch` hook to `CreateAgentSessionOptions` to track the boundary between session creation and the initial model request.
- Update `runSubprocess` and `TaskTool` to measure and log detailed latency metrics across the subagent lifecycle, including semaphore queue wait, setup time, and dispatch latency.
- Added weighted selection logic to bias discovery toward "[NEW]" tips.
- Updated the tip marker regex to allow trailing whitespace.
- Migrated tip selection to use the new biased distribution algorithm.
- Delay the creation of the extension context until a relevant handler is identified.
- Avoid unnecessary performance overhead for events that have no registered handlers, such as frequent streaming updates.
- Added an early return to stop event propagation if no handlers are registered for the event type.
- Updated the `agent_start` logic to return early, ensuring consistent flow control.
- Introduced a "[NEW]" marker in tip text that triggers an animated rainbow-colored tag in the welcome component.
- Implemented `renderNewTag` to calculate cyclic HSL hue offsets based on render timing.
- Updated `renderWelcomeTip` to intelligently wrap the tag at the end of the last line or on a new line to prevent layout overflow.
- Simplified system and personality prompts for improved conciseness and clarity.
- Streamlined tool instruction sets and parameter descriptions across all agent modules.
- Refactored prompt documentation in `hashline` to clarify terminology and task-specific constraints.
- Updated tool metadata in TypeScript service definitions to align with reduced documentation verbosity.
- Removed the customizable `timeout` parameter from the tool schema and implementation.
- Fixed the timeout duration to 5 seconds to simplify tool usage.
- Updated documentation and error messages to advise narrowing search patterns instead of adjusting timeouts.
- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.
- Updated the system schema and prompt options to disable inline tool descriptors by default.
- Modified the build system prompt logic to reflect the new default fallback value.
- Export `directoryExists` in `utils` to safely validate working directories before traversal.
- Update `SessionManager` and startup logic to fallback to the launch directory if a session's recorded working directory no longer exists.
- Add regression tests to ensure sessions now correctly adopt the launch directory instead of crashing on missing paths.
- Add `directoryExists` check to validate session working directories.
- Skip updating the session cwd if the recorded directory no longer exists on disk to prevent runtime errors.