- Added a shallow string-record sanitizer to normalize unknown model role values before applying updates.
- Updated setModelRole and overrideModelRoles to base persistence on global roles while retaining matching runtime overrides.
- Added model role override tests covering non-persistent temporary overrides, override clearing, and consistency after role updates over overrides.
- Added the active session model to compaction candidate selection before role-based candidates.
- Updated compaction routing so role-based models are only considered after the current chat model.
- Added a regression test proving an Anthropic session prefers its active model over `modelRoles.default` on OpenAI.
- Added a new shimmerText helper that computes a moving accent shimmer band across characters.
- Updated the interactive mode loader and slash-command ASCII bar renderer to use shimmer styling, with interrupt hints kept dim.
- Added tests for shimmer-enabled progress rendering and verified the visible bar output remains correct.
A per-cell timeout used to kill the persistent Python kernel, losing
all session state. The kernel.ts timeout path now sends SIGINT first
(letting the cell raise KeyboardInterrupt) and only escalates to
shutdown after a 5s grace window if the interrupt is ignored.
KernelExecuteResult gains an optional kernelKilled flag (defaulting to
false, propagated by the escalation timer and the unexpected-exit
handler). executor.ts formats two distinct timeout annotations: one
that says the kernel is still alive and reset:true would clear state,
one that says the kernel was killed and will be recreated.
Adds a shared classifyProviderHttpError helper that maps the well-known
failure shapes (status 401/402/403 and bodies matching credits / quota
/ insufficient) into compact SearchProviderError messages so the
orchestrator advances to the next provider instead of bailing.
Wired into every HTTP-talking provider (codex, exa, gemini, anthropic,
brave, jina, kimi, perplexity, searxng, synthetic, tavily, zai) and the
two wrapped-error providers (kagi via KagiApiError, parallel via
ParallelApiError). The orchestrator now collects per-provider failures
and emits a joined summary ("exa: 403 forbidden; codex: credits
exhausted; ...") when the whole chain fails.
Test updates: web-search-{exa,tavily,kagi} message expectations now
assert the compact "<id>: <status> <reason>" form.
- get op now returns paused goals (was returning null when enabled=false)
- complete op now works on paused goals; previously required enabled=true
which always failed after an interrupt set enabled=false
- create op now allowed after previous goal status is 'complete'; was
incorrectly blocked by the same guard as 'dropped' check
- goal tool is re-added to the active tool set on session reload when a
paused/active goal is persisted to disk; sdk.ts:1599 excludes 'goal'
from initial active tools unconditionally, so restoreModeFromSession
now re-adds it and saves #goalModePreviousTools for later cleanup
- goal_updated event for 'dropped' status now triggers #exitGoalMode
before clearing goalModeEnabled, ensuring the previous tool set is
restored when the agent drops a goal via the tool
- added 'resume' and 'drop' ops to goal tool schema and execute path
- updated goal.md prompt to document new ops and the paused-goal workflow
Fixes#1249
- Extended hashline parsing APIs to accept an optional target path so padding checks can be path-sensitive.
- Refined separator-padding heuristics to warn only on uniform single-space-before-content payloads while skipping those checks for indent-sensitive extensions.
- Threaded section paths through execute/diff callers and added tests covering warning behavior for both exempt and non-exempt file types.
Rolled back binary updater replacements when post-install version verification fails instead of deleting the previous working binary first.
Added a release workflow gate that downloads the published macOS arm64 asset and verifies codesign plus --version before npm publishing.
Fixes#1240
Route /force through a named Ollama tool choice and scope the Ollama request tools to that selected name so local models cannot pick a different tool.\n\nFixes #1236
ACP clients own MCP server configuration via session/new.mcpServers and AcpAgent#configureMcpServers. The ACP session factory previously left enableMCP at its default (true), so createAgentSession ran discoverAndLoadMCPTools on every session/new and the resulting host MCP tools landed in the session tool registry alongside the client-supplied ones. search_tool_bm25 then surfaced only the host tools.
Force enableMCP: false on every session created through createAcpSessionFactory so on-disk discovery is bypassed in ACP mode. Non-ACP modes (omp interactive, print, RPC) keep auto-discovery.
Fixes#1234
Prevent disabled providers from being registered for implicit local discovery and from creating built-in model discovery managers. Added regression coverage for disabled local providers during model registry refresh.
Fixes#1232
Bun's WinHTTP backend can ignore AbortSignal once a TCP/TLS connection
stalls (oven-sh/bun#15275, oven-sh/bun#18536), so Esc never reached the
in-flight `web_search` fetch on Windows and the session froze until
Ctrl+C. Only kimi shipped any timeout at all (server-side); every other
provider passed `signal` to `fetch` with no client-side bound.
Introduced `withHardTimeout(signal, ms=60_000)` in providers/utils and
wired it into every web-search provider's outbound fetch — anthropic,
brave, codex, exa, gemini, jina, kagi, kimi, parallel, perplexity
(api-key and oauth), searxng, synthetic, tavily, z.ai. 60s tolerates
legitimate slow LLM-mediated responses while still guaranteeing the
request settles within a minute when Bun's abort fails to propagate.
Independently, `executeSearch`'s provider-fallback loop swallowed every
`AbortError` as a regular provider error and returned
"All web search providers failed", masking cancellation on every
platform. The catch block now calls `throwIfAborted(signal)` first so a
caller-initiated cancel propagates as `ToolAbortError`.
Fixes#1221
- Added isStreaming-aware diff options and trimmed-input handling for streaming/non-streaming previews.
- Added helpers to trim trailing partial lines and strip unmatched trailing `-`/`@@` blocks during streaming.
- Reworked apply_patch and hashline preview builders to keep files in input order with per-line added mapping.
- Added streaming preview regression tests and Unreleased Fixed changelog notes for partial-line and ordering fixes.
- Added reusable loop auto-submit deferral and readiness helpers for loop mode.
- Updated loop iteration flow to defer next prompts while the session is streaming or compacting.
- Added tests verifying loop submissions wait until compaction/streaming completes before resolving.
- Preserved launch and attach request failures when configurationDone also fails.
- Handled initial stop-outcome watcher rejections for failed launch and attach attempts.
- Rejected directory-valued debug launch programs before adapter selection and documented the debugpy launch shape.
Fixes#1187
- Added splitInternalUrlSel to iteratively peel internal-URL selector chunks while preserving unsupported schemes like mcp://.
- Updated ReadTool to use the internal splitter before routing so selector parsing is handled via parseSel.
- Added unit tests covering malformed selectors, namespaced skill hosts, and unchanged behavior for non-URLs or unsupported schemes.
Grammar-constrained models (e.g. Qwen3.6-35B-MTP via llama.cpp) emit
`extra: { title: {} }` instead of `extra: { title: "<string>" }` because
the resolve schema declares `extra` as Record<string, unknown> with an
open value schema, leaving the model free to drop in an empty object.
The apply guard then threw 'Plan approval requires extra: { title: ... }'
on every retry, looping the model indefinitely (issue #1179).
Plan approval now uses a layered title resolution:
1. `extra.title` if it is a non-empty string (and sanitizes to non-empty)
2. First `# Heading` in the plan content
3. Filename stem of `planFilePath` (`'/data/workspaces/can1357__oh-my-pi__1179/.omp-session/2026-05-19T03-59-21-254Z_019e3e63-62a6-7000-be63-371f2cd6d67d/local/PLAN.md'` → `PLAN`)
4. Literal `plan` as a final safety net
Each candidate is run through `normalizePlanTitle`; rejected ones fall
through. Extracted as `resolvePlanTitle` in plan-mode/approved-plan.ts
so it's unit-testable.
Prompt language relaxed from MUST to SHOULD for `extra.title` in
plan-mode-active.md and plan-mode-tool-decision-reminder.md, noting the
fallback so models don't waste turns on a now-optional field.
Fixes#1179
- Added capParseErrors in shared render utilities and updated ast_grep and ast_edit to return capped parseErrors plus parseErrorsTotal.
- Threaded the preserved totals into parse-error formatting and renderer output so labels and overflow counts report the full number of issues.
- Added an ast_grep test asserting parse errors are capped at PARSE_ERRORS_LIMIT while parseErrorsTotal retains the original count.
- Implemented strict base64 validation and normalization for image `displayValue` payloads in the JS runtime, supporting strict strings, `Uint8Array`, `Buffer`, `ArrayBuffer`, typed-array views, and JSON `Buffer` objects.
- Dropped unrecognized image payloads while emitting a warning and fallback text instead of forwarding malformed data.
- Added tests covering successful coercions and invalid image data rejection paths.
- Guard renderInlineMarkdown against non-string input: partial JSON during
streaming can leave option label fields as undefined, causing marked.lexer
to throw 'undefined is not an object (evaluating e.replace)'. The ask tool
renderer now silently falls back to an empty string or baseColor output.
- Sanitize normalizePlanTitle instead of hard-rejecting: models that produce
natural-language plan titles like 'My Improvement Plan' were getting a
ToolError on every resolve call, causing an infinite retry loop. Spaces are
now converted to hyphens, remaining invalid chars are dropped, and only
truly unresolvable titles (empty after sanitization, path separators) throw.
- Fix ask.md prompt example: the example showed the legacy single-question
format (question/options/recommended at the top level) while the schema
requires questions: [{id, question, options}]. Models that follow examples
closely (Qwen3) generated calls that always failed schema validation.
Fixes#1176
Rewrites JSON-roundtripped Zod 4 schema objects that leak Zod internals as
JSON Schema keywords (e.g., `type:"enum"`, `enum:{...}`) into valid
JSON Schema 2020-12. This prevents validation failures when such schemas
are used as tool input schemas (e.g., from MCP servers).
Updates `isZodSchema` to reject deserialized Zod impostors that retain
`_zod` but lose their prototype.
Wires `toJSON` methods onto TypeBox shim schemas to ensure `JSON.stringify`
produces clean JSON Schema, preventing future leaks.
Fixes#1101
- Added stats sync, browser tab, and JS eval workers as explicit --compile
entrypoints in scripts/ci-release-build-binaries.ts so Bun emits them
into bunfs in published binaries, matching the dev build script and the
AGENTS.md worker spawn contract.
- Switched the release-binary smoke step in .github/workflows/ci.yml to
invoke --smoke-test (in addition to --version) so this regression
cannot ship again.
- Added packages/coding-agent/test/issue-1150-repro.test.ts pinning the
symmetric contract: both build scripts must list every worker entry.
Fixes#1150
- Set GIT_CONFIG_* and GIT_TERMINAL_PROMPT environment variables to prevent test interference from user gitconfig, LFS filters, signing, and credential helpers.
- Enhanced git error messages to include stdout when stderr is empty, providing better diagnostics for test failures.
- Replaced symlink-resolving pathIsWithin with lexical path containment check to prevent test isolation bypass via symlinked extensions.
- Tracked ACP tool-call inputs per session and replayed them via `toolArgsById`/`getToolArgs` plumbing.
- Merged ACP tool execution end content from start and result events so command output replay preserves original args.
- Scoped ACP async-job draining by session `ownerId` and `agentId` with in-flight tracking and permission-gated deferred turns.
- Refactored compaction telemetry and async tests with per-test telemetry setup and asynchronous teardown resets.