- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
- Added `qwenTemplateReasoningEffort` model compatibility option to disable Qwen chat template kwargs for strict local servers.
- Implemented fallback handling to strip rejected `chat_template_kwargs.reasoning_effort` and hoist values to top-level fields.
- Added comprehensive test coverage for Qwen reasoning effort fallback and keyword rejections.
bash.patterns only feeds the bash tool's approval decision. The eval
tool declares the exec tier and can spawn a shell via subprocess, so a
deny rule there does nothing for the same command run through eval;
under yolo the exec call resolves to allow. Note the scope and point at
tools.approval.eval as the lever that closes the path in
bash-tool-runtime.md, approval-mode.md, and settings.md.
Fixes#8838
- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
The idle watchdog aborts the request signal and cursor.ts closes
the Connect stream, so there is no in-flight server exec to race.
Unmarked MCP/todo blocks can continue once every emitted call has
a matching result, same as HTTP/2 RST.
Cursor hosted web search / Exa / unnamed field-9 WebFetch send
interaction_query and block the Run RPC until the client writes
interaction_response. Dropping the frame leaves the HTTP/2 stream
alive on heartbeats that are not semantic progress, so the 300s
idle watchdog aborts with "Provider stream stalled while waiting
for the next event".
Approve network permission gates and reject interactive
ask / switch-mode / create-plan. Leave VM setup unanswered
rather than inventing a success result.
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
The onPayload hook contract (README, docs/extensions.md) is that a non-undefined
return replaces the provider request payload, and every provider except these
three implements it (anthropic, openai-responses family, google, ollama — see
the earlier fix for the responses providers). openai-completions, amazon-bedrock
and cursor invoked the hook fire-and-forget and sent the original payload, so
extensions hooking before_provider_request could never transform the wire body
on these providers.
- openai-completions: await the hook and apply a non-undefined replacement to
the params used for the request body, raw request dump and error-path
fallback state
- amazon-bedrock: same for the ConverseStream command input
- cursor: await the hook for the AgentRunRequest; buildGrpcRequest becomes
async and is exported for direct testing (transport is HTTP/2)
- devin-agent intentionally unchanged: it does not fire the hook at all (its
payload is a protobuf object), which is a feature gap rather than a dropped
replacement; documented in README/docs instead
- regression tests: captured wire body reflects async/sync replacement, and an
undefined return keeps the original payload (completions + bedrock over a
mocked fetch; cursor by decoding the serialized run request)
On macOS the shared headless browser daemon launched from the system Google Chrome app bundle, running as a com.google.Chrome instance. macOS LaunchServices could then deliver the user's open-URL Apple Events to the daemon instead of their own Chrome, silently swallowing link clicks.
ensureChromiumExecutable now prefers the isolated Chrome for Testing binary (com.google.chrome.for.testing) on macOS, falling back to system Chrome only when Chrome for Testing cannot be obtained. Other platforms keep the download-avoiding system Chrome preference.
Fixes#8673
When a payload reports weighted counters but omits hard_cap, the soft/hard
split collapses to a single umans:requests row keyed off the authoritative
weighted effective-request counter, so a spent account can report exhausted
instead of being stuck at warning forever. Raw burst traffic above the limit
still never fabricates exhaustion; weighted headroom stays decisive (#7858).
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
reloadServer() sent the rust-analyzer-specific rust-analyzer/reloadWorkspace
request to every server before falling back to workspace/didChangeConfiguration.
Servers that crash on an unknown method instead of replying -32601 (Roslyn,
dotnet/roslyn#84890) were killed by `lsp reload` rather than reloaded.
Gate the request on isRustAnalyzerClient (exported from client.ts) or a
"rust-analyzer" server name; every other server reloads via
workspace/didChangeConfiguration directly. Consolidate tool.ts's inline
rust-analyzer detection onto the same helper.
Fixes#8571
Exa MCP servers were always filtered out because the native Exa integration covers web_search_exa. But configs that explicitly request web_fetch_exa or web_search_advanced_exa have no native equivalent, so filtering them made /mcp reconnect exa fail and hid those tools. Keep the MCP server mounted when its tools restriction includes anything beyond web_search_exa, while still extracting the API key for native search.
Review follow-up on #8052. The fallback trampolines captured one
`ExtensionContext` at `ExtensionRunner.initialize` and reused it for the
life of the session, unlike every other dispatch site, which builds one
per call. `createContext()` materializes `cwd` and `hasUI` as values, so
a handler kept seeing the workspace the runner initialized in — wrong
the moment `SessionManager.moveTo()` relocates the session (`/move`),
where a handler that scopes or prompts against `ctx.cwd` would allow the
old workspace and deny the new one.
Both trampolines now build the context inside the invocation. A denied
mutation is a rare path, so the extra object costs nothing that matters,
and anything else `createContext()` snapshots is refreshed with it.
Covered by an integration test that moves the session's cwd after
initialize and asserts the handler sees the new one; hoisting the
context back out of the invocation fails it.
Review on #8052 found three problems in the write/delete fallback seam.
A symlink guard that only `lstat`'d the final component let the same
escape through a symlinked ancestor: `ws/link/file` under a
`ws/link -> /outside` link reached a handler as a lexically innocent
path, so a helper's prefix allowlist passed while the bytes landed
outside. Refusing every symlinked component is not available, since
`/var` and `/tmp` are links on macOS and every path under `os.tmpdir()`
traverses one. So `req.dst` is now the path the failed syscall itself
acted on, via `resolveSyscallTarget` beside `confineToWorkspace`: fully
resolved for a write, resolved up to the last component for a delete,
because `unlink` removes a link rather than following it. Resolving also
closes the TOCTOU window a refusal left open. A path that cannot be
canonicalized — a dangling final link, or an ancestor whose own
resolution is denied — is not brokered at all.
Per-handler throw isolation lived outside the per-extension trampoline,
so a throw from one extension's first handler advanced the registry to
the next extension and skipped every later handler that one had
registered. Each handler call is now wrapped individually.
The registry stays process-wide. A subagent spawned with restricted
tools gets `preloadedExtensionPaths: []` and loads no extensions of its
own, so scoping resolution to the originating session would turn its
brokered writes into hard failures, and a host that registers once in
its top-level session expects its subagents covered. The request names
its origin instead: `req.sessionId` against the handler's own
`ctx.sessionManager.getSessionId()`, entered by `ExtensionToolWrapper`,
which `sdk.ts` already puts around the whole tool registry whenever a
runner exists. Both registries are also walked over a snapshot, so a
concurrent session shutdown cannot make another session's walk skip
whichever handler shifted into the hole.
Unit tests go 30 -> 35 and integration 6 -> 7, covering the resolved
target, a symlinked ancestor on both seams, a dangling link, a target
whose own metadata is behind the boundary, same-extension handler
ordering after a throw, and `req.sessionId` matching the handler's own
session end to end.
A host that runs omp inside an OS sandbox can grant a path mid-session but cannot
apply that grant to an in-process write: `write` and `edit` do their I/O in the
agent process, so an out-of-workspace write fails and stays failed until the
process restarts under a wider profile.
Nothing available today closes that. A `tool_call` handler can block and a
`tool_result` handler can rewrite content, but neither can re-run a tool.
`ctx.invokeTool` delegates execution, but the delegated native tool runs in the
same process under the same restrictions. And the failure lands AT the write
syscall - after the tool computed the final content, before it returned - so the
bytes are gone with the throw, and reconstructing them means reimplementing
`edit`'s hashline protocol and the snapshot bookkeeping.
The byte-write that `write`, `edit` and `apply_patch` perform on an ordinary file
path already funnels through one two-line primitive
(`file ? file.write(content) : Bun.write(dst, content)`) at four call sites.
Routing that primitive through `writeFileWithFallback` gives an embedder a single
seam to intercept a permission-denied write: the native tool still records its own
snapshot under the real destination path once a handler reports success, so a
follow-up hashline `edit` on that path keeps working.
Only a permission boundary diverts - `EPERM`/`EACCES`/`EROFS`. Two cases needed
more than that:
- `Bun.write` creates missing parents itself, and when that `mkdir` is the denied
operation it reports the subsequent `open()`'s `ENOENT` instead of the denial -
making a sandboxed write into a new out-of-tree directory indistinguishable from
an ordinary bad path. Redoing the `mkdir` explicitly recovers the real errno, and
because it runs through the same enforcement path as the write it also sees
kernel-level denials (Seatbelt, LSM) that a `stat`/`access` probe reports as
writable. If no handler takes the write, the original `ENOENT` is still what
propagates, with the recovered denial attached as its `cause`.
- `apply_patch` creates the parent as a separate step before writing, so a denial
there threw before the seam was ever reached. That `mkdir` now tolerates a
permission denial when a fallback is registered, letting the write report it.
A denial reached through a SYMLINK is never brokered. The in-process write follows
the link, so the kernel denied the link's TARGET, but a handler receives `dst` and
a privileged helper opening it with ordinary follow semantics would land the bytes
wherever the link points. That also defeats the obvious helper-side defence, since
a prefix allowlist passes when the link sits inside the allowed root while its
target does not. omp cannot vouch for the destination, so it refuses rather than
hand the ambiguity to a privileged writer - the same answer `confineToWorkspace`
already gives an unresolvable link.
Removing a file is a different primitive, so it gets its own seam
(`deleteFileWithFallback`, `registerFileDeleteFallback`) covering `edit`'s `REM`,
a hashline `MV`'s source unlink, and `apply_patch`'s delete op. Two differences
from the write path: `ENOENT` is never diverted, since nothing is created on the
way to an unlink and `REM` needs it to become a not-found error; and the seam
refuses a target it can confirm is a directory, because `unlink` on a directory
reports `EPERM` on Darwin and is otherwise indistinguishable from a sandbox
denial. That check cannot always run - a sandbox denying the unlink usually denies
the target's metadata too - so the request carries `confirmedFile`, and a handler
is required to use a plain unlink rather than resolving or recursing.
The two registries are deliberately separate. A write handler brokers `content` to
`dst`, so a delete request reaching it with no content invites brokering an empty
write and truncating the file it was asked to remove.
With nothing registered both seams are inert: the primitives run exactly as
before, a failure rethrows from the same place, and no extra syscalls are
performed.
Scope is deliberately narrow. Archive-member and SQLite writes are unchanged -
neither is a byte-write to a path, so brokering them needs a different request
shape - along with the ACP bridge's `writeTextFile`, the `lsp` tool's own
workspace-edit and formatter writes, and directory removal.
Move the per-request date/cwd line out of the system prompt into a
first-turn system-reminder so open-weight providers keep their tool-schema
prefix cache; the reminder refreshes itself at midnight. Closes#7404.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <noreply@codebuff.com>
Split the request limit into a weighted soft-cap row and a raw burst-ceiling
row so healthy accounts no longer read as exhausted; surface the rolling
window's resets_at as a countdown. Closes#7858.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <noreply@codebuff.com>
- Switched Docker images to build native addons via cargo/napi-rs (`OMP_NATIVE_BUILD_BACKEND=cargo`) instead of Bazel.
- Updated Cargo.toml workspace members explicitly to prevent loading errors from stale directories under crates/.
- Added depth-agnostic patterns to .dockerignore files to exclude nested build outputs from Docker build contexts.
- Added OMP_NATIVE_CARGO_PROFILE environment variable to support configuring cargo profiles for addon builds.
The win32 addon static-CRT switch enabled the static_link_msvcrt cc
feature, but audiopus_sys's bundled opus is built through the generated
CMake toolchain, which pinned CMAKE_MSVC_RUNTIME_LIBRARY=MultiThreadedDLL
(/MD) as authoritative under CMP0091 NEW. The opus objects could then
still emit /MD and pull VCRUNTIME140.dll / conflict with the static CRT
the rest of the addon links, leaving the outcome dependent on compile-
flag ordering.
Pin CMAKE_MSVC_RUNTIME_LIBRARY to MultiThreaded (static release /MT) in
the msvc toolchain.cmake so opus deterministically matches rustc's
+crt-static and the static_link_msvcrt feature.
Verified with a fully cold `bazel build //:natives-win32-x64-baseline`
(opus recompiled): the produced .node imports no VCRUNTIME140.dll and no
api-ms-win-crt-* — only core Windows system DLLs.
Fixes#8439
The shipped win32-x64 pi_natives addon linked the dynamic MSVC CRT (/MD)
and imported VCRUNTIME140.dll from the Visual C++ Redistributable, which
is absent on a clean Windows install. LoadLibrary of the extracted .node
then failed with error 126 ("The specified module could not be found"),
so omp could not start after a fresh `irm install.ps1 | iex`.
Static-link the CRT for the win32 addon: +crt-static for rustc (crate
BUILD select) plus the static_link_msvcrt cc feature enabled for win32 in
the native_addon transition, so its C deps (opus/cmake, tree-sitter,
blake3, ring) compile /MT in lock-step. The rebuilt .node imports only
core Windows system DLLs -- no VCRUNTIME140.dll, no api-ms-win-crt-*.
Fixes#8439
Kept direct Herdr panes on the host-safe in-place resize path so streaming redraws no longer clear and replay pane scrollback.
Updated the resize regression, stress scenario, renderer docs, and changelog.
Fixes#8431
- Generalized thinking loop guard and helper functions to support Gemini, DeepSeek, and Grok model families.
- Replaced `withGeminiThinkingLoopGuard` and related Gemini-specific symbols with generalized counterparts.
- Removed deprecated `enableGeminiThinkingLoopGuard` options and associated tests.
- Updated test suites and agent session logic to use the generalized thinking loop guard and model family tokens.
- Replaced blanket subagent advisor global settings with fine-grained per-agent configuration and frontmatter support.
- Added dashboard keybindings and inline override editors for managing agent advisor patterns.
- Implemented settings migration logic to convert legacy global options into per-agent settings.
- Updated session persistence and execution layers to restore and enforce per-agent advisor behaviors.