Passed the per-request model through SDK payload callbacks and ExtensionRunner context creation.
Added regression coverage for Codex-primary and Anthropic-request contexts.
Fixes#6006
- Parsed live named efforts, mandatory-thinking state, and model protocol metadata.
- Sent native Kimi named efforts and adaptive Anthropic override efforts without generic token budgets.
Fixes#5893
Caller-provided output schemas are free-form JSON and cannot be represented by OpenAI strict tool schemas. Keep todo strict while explicitly sending task as non-strict.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
Resolved plan-mode exit overlap with #5662 (kept restore/rollback
structure, routed pending-switch clearing through
clearPendingPlanModelSwitch) and unioned additive test blocks with
#5672/#5662.
Resolved sdk.ts overlap with #5651 (kept getCursorTools alongside the
extracted transformToolCallArguments) and beginDispose overlap with
#5668 (kept both title-generation and autolearn-capture aborts).
Extension/SDK/RPC registerTool defaulted an omitted loadMode to
"discoverable". A UI-only re-register of an essential built-in
(read/write/bash/edit/glob) then became discoverable and, with tools.xdev
on, was unmounted from the top-level schema. read/write dropping also
broke the xd:// transport (read xd://, write xd://<tool>), leaving the
model with no callable coding essentials.
- Add defaultLoadModeForToolName: omitted loadMode resolves to "essential"
for known essential built-in names, "discoverable" otherwise.
- Apply it at all four adapter boundaries (extension wrapper, custom-tools
wrapper, sdk customToolToDefinition, rpc normalizeHostToolDefinitions).
- Transport invariant: read/write never mount under xdev regardless of
loadMode (they carry the transport).
- Regression test covering the demotion, transport invariant, and a drift
guard tying the essential-name set to the tool classes.
Fixes#5764
- Prevented compaction from reopening a settled terminal answer unless queued work or an active goal remains.
- Ran auto-learn capture in an abortable detached agent with constrained tools and isolated provider state.
- Replaced primary-turn capture coverage with private-capture regression tests.
Fixes#5715
Built-in xd:// devices are mounted before the SDK wraps registry tools in ExtensionToolWrapper, so Cursor executed them via tool.execute() without the deny/prompt approval gate that write xd:// enforces. Wrap unwrapped devices in the Cursor resolver, skipping already-wrapped dynamic mounts.
Fixes#5650
Forwarded the session xd registry into Cursor provider tool contexts.
Routed Cursor MCP execution through the mounted registry fallback and added regression coverage for built-in devices and external MCP tools.
Fixes#5650
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated system and todo prompt templates to require batching todo tool calls with real action calls instead of sending them alone.
- Passed a new `prewalkArmed` session flag in `createAgentSession`, set from whether `prewalk` was supplied.
Switching from a vision model to a text-only model kept replaying
historical image content blocks to the new provider, which rejected them
with invalid_argument. The convertToLlm wrapper in sdk.ts only filtered
images when images.blockImages was set, never by model capability, so
the outbound request carried image blocks the active model could not
accept.
Add replaceLlmImagesWithText() to scrub image blocks out of the
already-converted LLM message view, and call it from the sdk wrapper
when the active model's input lacks "image". History on disk keeps its
images; only the provider request is scrubbed, and the check reads the
active model dynamically so a /model switch takes effect next turn.
Fixes#5400
generate_image was registered as a custom tool and force-activated via the alwaysInclude list in createAgentSession, so it survived --no-tools (empty toolNames) and any explicit whitelist that omitted it. There was also no generate_image.enabled setting, so /settings had no toggle.
Add a generate_image.enabled setting and only register the tool when enabled and either no whitelist is given or it names generate_image.
Fixes#5305
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Removed the Google Interactions transport and deleted interaction-specific request options from the shared AI stream typing/API surface.
- Simplified Google provider routing to eliminate interactions auto-selection logic and keep `streamGoogle` on the `:streamGenerateContent` path.
- Updated Vertex request handling to use resolved stream hosts without `/interactions`/`Api-Revision` and removed related interaction constants.
- Deleted obsolete Interactions tests and updated remaining Google stream tests to no longer reference `useInteractionsApi`/`storeInteraction`/`previousInteractionId`.
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.
- Replaced legacy reasoning-slide functionality with new downshift and plan-yolo capabilities.
- Updated CLI arguments, slash commands, and configuration schemas to support the new model-switching and execution behaviors.
- Refactored agent session logic to handle downshift arming, plan-yolo headless execution, and context scrubbing.
- Renamed and added system prompts to align with the updated downshift and plan-yolo workflows.
Detect pending tool calls from assistant content and use the selected model metadata to terminate first-turn user tails after abnormal process exits.
Signed-off-by: Christian Stewart <christian@aperture.us>
Appended one terminal aborted assistant record when a persisted abnormal-exit diagnostic follows a non-terminal transcript tail, preserving partial history while making resumed context valid.
Signed-off-by: Christian Stewart <christian@aperture.us>
- Transitioned vibe tools to an ephemeral registration model where they are installed only when entering `/vibe` mode and removed upon exit.
- Added `activateVibeTools` and `deactivateVibeTools` methods to `AgentSession` to manage these transient tool registrations.
- Removed vibe tools from the default global tool registry, preventing unnecessary background exposure.
- Implemented Vibe mode to enable worker session management and director-role context injection.
- Added a `/vibe` slash command and integrated status line UI to display mode activity.
- Configured restricted toolsets and guards to prevent concurrent conflicts with existing Goal or Plan modes.
- Provided system prompts and tool templates to support specialized agent communication and task orchestration.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.
- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.
- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.
Fixes#5035
Interactive sessions defer MCP discovery, so CLI --tools produced an initial built-in-only active set and later MCP refreshes respected that filtered set.
Force-activate deferred MCP tools when MCP discovery mode is disabled, matching the blocking startup path while leaving discovery-mode selection intact.
Fixes#5013
- Copied localProtocolOptions through SDK-created custom tool contexts so startup MCP tools resolve '/data/workspaces/can1357__oh-my-pi__4946/.omp-session/2026-07-09T16-06-43-993Z_019f47a1-a619-7000-9062-5f5d863afa45/local' against the active session.
- Exposed localProtocolOptions on extension contexts and covered the runner propagation path.
Fixes#4946