A budget-aborted keep-alive subagent's job row (job id == agent id) settles
failed and is retained ~5 min; executeCancel short-circuited to
already_completed for that window, leaving the zombie registration
unkillable exactly when the user wants it dead. Fall through to
cancelAgentRegistration for settled rows; keep already_completed when no
lingering registration exists.
Also stop wiring AgentLifecycleManager.global() onto SDK sessions created
with a caller-supplied agentRegistry: the global lifecycle releases through
AgentRegistry.global(), so it would report a cancel while releasing an
unrelated global ref. Without a lifecycle, cancel falls back to
dispose + unregister on the session's own registry.
Addresses both Codex P2 review findings on #6319.
The getDiscoverableProviders() guard skipped awaiting runtimeDiscoveryPromise
when no config-discovery providers exist, so a cold deferred selector backed
only by runtime model managers (extension fetchDynamicModels) with implicit
local discovery disabled still resolved against the offline cache. Awaiting
unconditionally is free when no runtime managers are registered
(refreshRuntimeProviders early-returns); the full refresh fallback stays
gated on discoverable providers.
Three read paths raced background model discovery on cold start:
1. `get_available_models` RPC (rpc-mode.ts) read the registry
synchronously and returned a partial catalog containing only
statically-bundled models.
2. `set_model` RPC (rpc-mode.ts) read the registry synchronously and
rejected discovery-backed selectors with "Model not found".
3. `--model <provider>/<pattern>` CLI flag deferred retry (sdk.ts:2078)
resolved synchronously after extension registration, before
discovery-backed providers had populated `#models`.
Paths 1 and 2 are fixed by exposing the existing in-flight background
refresh promise (`#backgroundRefresh`, already tracked and cleared by
`refreshInBackground`) via a new public
`ModelRegistry.awaitBackgroundRefresh()` method, and awaiting it at each
RPC read site. No-op when no refresh is in flight (warm sessions
unaffected).
Path 3 mirrors the cold-cache race fix already applied to the
default-role fallback on this branch (issues #6114, #6162, sdk.ts:2343):
when a deferred pattern is unresolved and any discoverable provider is
registered, run a cache-aware `refresh("online-if-uncached")` pass
before the retry. Reuses the existing discovery machinery rather than
introducing a new ordering dependency.
The `omp models` CLI never had this bug because it awaits
`modelRegistry.refresh()` directly before listing.
Behavioral characteristics:
- **Warm-session fast path preserved**: when no refresh is in flight
(`#backgroundRefresh === undefined`), `await undefined` resolves in a
microtask. No regression for sessions that don't need discovery or
have already settled.
- **Failure isolation preserved**: `refreshInBackground()` already
swallows discovery errors via `.catch(...)`, so `awaitBackgroundRefresh()`
resolves even when discovery fails — callers then read whatever models
made it into `#models` (built-in + cached). No new failure modes.
- **Scoped**: doesn't change `refreshInBackground()` semantics. Adds a
new read-only awaiter with minimal API surface. Reuses the
well-established `refresh("online-if-uncached")` pattern for the
deferred retry path.
Reproduction (get_available_models RPC, with any discovery-backed
provider configured in `~/.omp/agent/models.yaml`):
cd ~
{
sleep 1
printf '%s\n' '{"id":"m1","type":"get_available_models"}'
sleep 5
} | timeout 15 omp --mode rpc-ui --approval-mode yolo 2>/dev/null \
| grep '"id":"m1"' | jq '.data.models | {count: length, providers: ([.[].provider]|unique)}'
Before: discovery-backed provider absent from the response on cold start.
After: discovery-backed provider present.
Reproduction (--model CLI flag, same config):
omp --mode rpc-ui --model <discovery-provider>/<model-id> --approval-mode yolo
Before: exits 1 with "Model \"<discovery-provider>/<model-id>\" not found".
After: starts rpc-ui session with the requested model selected.
A keep-alive subagent force-stopped for exceeding its soft request budget
is kept resumable (status idle, adopted by AgentLifecycleManager) so its
context can be salvaged, but its async job row settles and is reaped after
~5 min. After that, hub cancel <id> only reported "Background job not
found" because executeCancel consulted AsyncJobManager alone, leaving the
registration unkillable short of a broker restart.
hub cancel now falls through to the agent registration when no live job
matches: for a sub the caller spawned, it aborts any in-flight turn,
disposes the session, and releases it from the lifecycle. Cross-agent
kills stay impossible and Main/advisor refs are never targeted.
Fixes#6315
customToolToDefinition rebuilt ToolDefinition without strict, so the
Task proxies' (and startup MCP tools') explicit strict:false was dropped
before RegisteredToolAdapter and the OpenAI-family serializers saw it.
Declare strict on ToolDefinition, copy it in the bridge, and cover the
proxy -> definition -> registered adapter path in the parity test.
A retry-chain entry without its own :level suffix now inherits the
unavailable primary's configured thinking level, matching runtime
fallback-chain semantics. Regression test asserts a level that differs
from the fallback model's default.
- Carried startup-selected fallback role and primary selector into AgentSession.
- Continued remaining role fallback entries after the startup fallback fails.
- Added regression coverage for chained startup failover.
Fixes#6283
- Carried configured role identity through deferred CLI model resolution.
- Consulted ordered authenticated role fallbacks after unavailable primaries.
- Added startup regression coverage for missing primary and fallback entries.
Fixes#6283
On a cache-cold interactive launch, a modelRoles.default pointing at a
models.yml discovery provider (openai-models-list) was silently replaced
by an unrelated authenticated provider's default. createAgentSession
resolves the default role before background discovery populates the
catalog, so the configured provider had no models yet and the fallback
fell through to pickDefaultAvailableModel.
Await one cache-aware discovery pass and re-resolve the configured
default whenever it is still unresolved (not only when nothing resolved
at all) before accepting a bundled-provider fallback.
Fixes#6162
Passed the per-request model through SDK payload callbacks and ExtensionRunner context creation.
Added regression coverage for Codex-primary and Anthropic-request contexts.
Fixes#6006
- Parsed live named efforts, mandatory-thinking state, and model protocol metadata.
- Sent native Kimi named efforts and adaptive Anthropic override efforts without generic token budgets.
Fixes#5893
Caller-provided output schemas are free-form JSON and cannot be represented by OpenAI strict tool schemas. Keep todo strict while explicitly sending task as non-strict.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
Resolved plan-mode exit overlap with #5662 (kept restore/rollback
structure, routed pending-switch clearing through
clearPendingPlanModelSwitch) and unioned additive test blocks with
#5672/#5662.
Resolved sdk.ts overlap with #5651 (kept getCursorTools alongside the
extracted transformToolCallArguments) and beginDispose overlap with
#5668 (kept both title-generation and autolearn-capture aborts).
Extension/SDK/RPC registerTool defaulted an omitted loadMode to
"discoverable". A UI-only re-register of an essential built-in
(read/write/bash/edit/glob) then became discoverable and, with tools.xdev
on, was unmounted from the top-level schema. read/write dropping also
broke the xd:// transport (read xd://, write xd://<tool>), leaving the
model with no callable coding essentials.
- Add defaultLoadModeForToolName: omitted loadMode resolves to "essential"
for known essential built-in names, "discoverable" otherwise.
- Apply it at all four adapter boundaries (extension wrapper, custom-tools
wrapper, sdk customToolToDefinition, rpc normalizeHostToolDefinitions).
- Transport invariant: read/write never mount under xdev regardless of
loadMode (they carry the transport).
- Regression test covering the demotion, transport invariant, and a drift
guard tying the essential-name set to the tool classes.
Fixes#5764
- Prevented compaction from reopening a settled terminal answer unless queued work or an active goal remains.
- Ran auto-learn capture in an abortable detached agent with constrained tools and isolated provider state.
- Replaced primary-turn capture coverage with private-capture regression tests.
Fixes#5715
Built-in xd:// devices are mounted before the SDK wraps registry tools in ExtensionToolWrapper, so Cursor executed them via tool.execute() without the deny/prompt approval gate that write xd:// enforces. Wrap unwrapped devices in the Cursor resolver, skipping already-wrapped dynamic mounts.
Fixes#5650
Forwarded the session xd registry into Cursor provider tool contexts.
Routed Cursor MCP execution through the mounted registry fallback and added regression coverage for built-in devices and external MCP tools.
Fixes#5650
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated system and todo prompt templates to require batching todo tool calls with real action calls instead of sending them alone.
- Passed a new `prewalkArmed` session flag in `createAgentSession`, set from whether `prewalk` was supplied.
Switching from a vision model to a text-only model kept replaying
historical image content blocks to the new provider, which rejected them
with invalid_argument. The convertToLlm wrapper in sdk.ts only filtered
images when images.blockImages was set, never by model capability, so
the outbound request carried image blocks the active model could not
accept.
Add replaceLlmImagesWithText() to scrub image blocks out of the
already-converted LLM message view, and call it from the sdk wrapper
when the active model's input lacks "image". History on disk keeps its
images; only the provider request is scrubbed, and the check reads the
active model dynamically so a /model switch takes effect next turn.
Fixes#5400
generate_image was registered as a custom tool and force-activated via the alwaysInclude list in createAgentSession, so it survived --no-tools (empty toolNames) and any explicit whitelist that omitted it. There was also no generate_image.enabled setting, so /settings had no toggle.
Add a generate_image.enabled setting and only register the tool when enabled and either no whitelist is given or it names generate_image.
Fixes#5305
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Removed the Google Interactions transport and deleted interaction-specific request options from the shared AI stream typing/API surface.
- Simplified Google provider routing to eliminate interactions auto-selection logic and keep `streamGoogle` on the `:streamGenerateContent` path.
- Updated Vertex request handling to use resolved stream hosts without `/interactions`/`Api-Revision` and removed related interaction constants.
- Deleted obsolete Interactions tests and updated remaining Google stream tests to no longer reference `useInteractionsApi`/`storeInteraction`/`previousInteractionId`.
- Modify downshift logic to ignore `todo` tool calls as triggers, requiring them instead to open a "gate" that permits switching only on subsequent `edit`/`write` actions.
- Ensure the starting model consistently handles implementation until the todo list is established, preventing premature hand-off to the fast/cheap model.
- Update system prompts to emphasize strict validation and multi-test execution requirements for the boomerang model upon returning to the primary context.
- Included the todo tool in downshift action triggers, gated on the plan nudge being in context — a post-nudge todo init is the planning-complete signal, while a turn-one todo remains bookkeeping.
- Updated flag help, settings schema, SDK docs, slash-command text, and plan-nudge instructions.
- Added a regression test covering the nudge-gated todo trigger.