- Fixed bash interceptor to check both raw and cwd-normalized commands, catching commands hidden behind leading `cd ... &&` wrappers.
- Fixed LSP client shutdown to await graceful shutdown with a 5s timeout before killing the process, and parallelized `shutdownAll` via `Promise.allSettled`.
- Fixed concurrent bash command tracking by replacing a single abort controller with a Set, preventing premature cancellation of parallel commands.
- Removed `./hooks` and `./hooks/*` export entries from the coding-agent package exports map.
- Updated pinned Rust nightly toolchain from `nightly-2026-03-27` to `nightly-2026-04-29` in `rust-toolchain.toml` and CI workflow.
- Replaced custom already-published detection in `ci-release-publish.ts` with `bun publish --tolerate-republish` flag.
The helper guarded `model.api === "openai-completions"` and returned undefined
for openai-responses models. The discoverable-custom-compat test sets
`api: "openai-responses"` on a custom model with `compat.extraBody`, so the
post-refresh assertion saw `undefined` instead of the configured proxy hint.
The OpenAICompatSchema gates user-facing custom-model compat regardless of the
underlying api wire format, so reading the field as OpenAICompat for any api
matches what the registry actually stores.
Commit a190397d8 made `Model.compat` resolve to `OpenAICompat | AnthropicCompat`
under the default `TApi = any`. The widened union broke every site that treated
`compat` as openai-shaped: model-registry deep-merge, openai-completions resolved
compat, and ~20 test fixtures. This restores the assumption locally instead of
papering over it with casts.
- getBundledModel is now generic on TApi so test fixtures that spread it into
`Model<"openai-completions">` get the narrow compat back.
- mergeCompat is generic over TBase/TOverride; the schema-driven model-registry
override path keeps its OpenAICompat-shaped merge fields, anthropic overrides
pass through untouched.
- OpenAICompatSchema gains the openai-only fields it was missing
(requiresMistralToolIds, reasoningContentField, requiresReasoningContent*,
thinkingFormat, requiresThinkingAsText, disableReasoningOnForcedToolChoice).
- resolveOpenAICompat fills in disableReasoningOnForcedToolChoice so the
Required<OpenAICompat> shape stays satisfied.
- Anthropic tool-result block id assignment uses the proper unknown double-cast.
- isForcedToolChoice accepts unknown so it can read `params.tool_choice` whose
type comes from the OpenAI SDK ChatCompletionToolChoiceOption (now wider than
our local OpenAICompletionsToolChoice).
- Test fixtures and Required<OpenAICompat> literals updated for the field set.
Fixes CI red on main.
- Changed the parameter type from OpenAICompletionsToolChoice to unknown to accommodate widening type changes in the OpenAI SDK's ChatCompletionToolChoiceOption.
- The function only checks for specific open/forced values, so it safely handles any input shape.
Mirror anthropic.ts:disableThinkingIfToolChoiceForced for backends that 400
on combined reasoning + forced tool_choice. Kimi explicitly rejects this
combination ('tool_choice specified is incompatible with thinking enabled')
on its native API, OpenCode-Go, OpenRouter, etc. Anthropic itself enforces
the same constraint, so Claude reached through OpenAI-compat proxies
(LiteLLM, Vertex chat-completions, OpenRouter) needs the same handling.
Adds disableReasoningOnForcedToolChoice compat flag, defaulted on for any
Kimi (moonshotai/kimi*, kimi-* ids) or Anthropic (provider/baseUrl/claude*
ids) model. When tool_choice resolves to anything that forces a tool call
(required, named function), reasoning_effort and the OpenRouter-shaped
nested reasoning object are dropped for that turn. The forced tool_choice
itself stays so the agent still gets the tool call.
Replaces the previous (incorrect) approach of unconditionally dropping
tool_choice for kimi reasoning models, which broke explicit tool routing.
Fixes#827
Standalone Bun binaries on WSL (and any host where the user moves the
binary away from the build-host's checkout) failed to load
pi_natives.<platform>-<arch>*.node. The loader's isCompiledBinary
detection relied on two signals that are both false in shipped binaries:
process.env.PI_COMPILED (bun --define PI_COMPILED=true substitutes the
bare identifier, not property accesses on process.env) and
__filename.includes("$bunfs") (Bun retains the build-host absolute path
in __filename for required CJS modules — only import.meta.url is
rewritten). Detection therefore returned false, embedded-addon
extraction was skipped, and the only candidates probed were the
build-host nativeDir and execDir.
Make embedded-addon presence the authoritative compiled-mode signal
(it is null in the post-build --reset stub, populated when embed:native
ran for the standalone build), eagerly require the manifest, and
extract candidate-path computation into a pure helper covered by a
host-platform-agnostic unit test. Also fix the build-time --define so
process.env.PI_COMPILED is genuinely set at runtime as a defensive
fallback.
Fixes#823
renderContentPreview crashed on the Windows prebuilt with "Failed to
convert napi value Null into rust type `u8`" (and the related `bool`
variant) because the wrapper forwarded `null` to napi-rs 3's
`Option<Ellipsis>` / `Option<bool>` parameters; the Windows binding
rejects null where Linux/macOS tolerates it. `maxWidth` was also passed
through unchecked, so a stray null/undefined upstream blew up the
required `u32` everywhere. Pass concrete defaults that mirror the Rust
`unwrap_or`s and clamp `maxWidth` to a sane non-negative integer.
Fixes#848
flushCompactionQueue() fires session.prompt(text) on the first non-slash
queued message. If the session is still streaming when compaction-end
lands (race between isStreaming flipping false and the event arriving),
prompt() throws AgentBusyError, restoreQueue() dumps the message back
into compactionQueuedMessages, and it stalls there: nothing drains that
array except the next compaction-end. The user sees the steer preview
but cannot deliver the message (Alt+Up consults session.clearQueue, not
compactionQueuedMessages).
Pass streamingBehavior derived from the queued message's mode so
prompt() routes into the steer/follow-up queue when busy and runs as a
fresh prompt when idle.
Fixes#825
buildSessionContext walked the entry path and unconditionally overwrote
models.default from every assistant message's reported model. Temporary
fallbacks (retry fallback, context promotion) and codex-side model
downgrades both produce assistant messages tagged with a different model
id, which clobbered the user's explicit /model pick on resume and made
the session silently revert to the older model.
Treat assistant-message inference as a legacy fallback that only fills
in models.default when no explicit `model_change` with role="default"
has been seen on the path.
Fixes#849
Windows Terminal lacks Kitty keyboard protocol, so Alt+Shift+P is emitted
as the legacy two-byte sequence ESC + 'P'. The native key parser only
handled lowercase ESC pairs (alt+letter), so matchesKey("\x1bP",
"alt+shift+p") and parseKey("\x1bP") both failed, breaking the
app.plan.toggle keybinding (and every other Alt+Shift+letter shortcut)
on Windows.
Add ALT_SHIFT_LETTERS table and route ESC+A..Z through it in both
parse_esc_pair (parsing) and matches_key_inner (matching) when Kitty
protocol is inactive.
Fixes#879
Phase1 caught every per-claim failure, recorded the reason in
jobs.last_error, and surfaced only an aggregate failed count via the
phase1 completion debug line. Users hitting setup-time failures (e.g.
WSL2 stale rollout paths producing ENOENT before any LLM call) had no
diagnostic. Emit logger.error per failed claim with threadId,
rolloutPath, and reason so the actual error is visible in omp.log.
Fixes#846
Add optional defaultLevel to ThinkingConfig schema/type so models.yml can
declare a preferred starting thinking level per model. On model switch
the agent session adopts model.thinking.defaultLevel when present (with
explicit caller-supplied level still winning); otherwise current behavior
is preserved. SDK initial selection prefers the model's defaultLevel
before falling back to the global defaultThinkingLevel setting.
Fixes#775
Adds a DeepSeek provider descriptor (api.deepseek.com, openai-completions,
DEEPSEEK_API_KEY) with a models.dev mapping for the deepseek-v4 family. The
v4 reasoning models map xhigh -> reasoning_effort=max, disable tool_choice
(returns 400 against DeepSeek when reasoning_effort is set per #830 thread),
and round-trip reasoning_content on tool calls to keep chain-of-thought
attached across tool steps.
Fixes#830
Vertex AI's Anthropic-compatible endpoint rejects strict tool defs
(`tools.<n>.custom.strict: Extra inputs are not permitted`) and the
adaptive thinking tag (`Input tag 'adaptive' ... does not match`).
Add an `AnthropicCompat` shape on `Model.compat` with two opt-in
flags: `disableStrictTools` drops `strict: true` from tool
definitions, `disableAdaptiveThinking` falls back to budget-based
`thinking: { type: "enabled" }`. Default behavior unchanged.
Fixes#826
When entering plan mode while the session is streaming, #applyPlanModeModel
defers the switch into #pendingModelSwitch and snapshots the previous model.
On exit, the snapshot was restored but the deferred switch was left queued,
so the next agent_end flush landed the session on the plan-role model after
the user had already left plan mode.
Drop the pending switch in #exitPlanMode when its target matches the
plan-role resolution; leave any other queued switch alone.
Fixes#816
Xiaomi splits API keys across two backends: standard sk- keys use
api.xiaomimimo.com while token-plan tp- keys use the EU host
token-plan-ams.xiaomimimo.com. Previously the xiaomi provider hardcoded
the standard host for both /login validation and model discovery, so
tp- keys were rejected at the validation step.
Detect the tp- prefix in loginXiaomi and xiaomiModelManagerOptions and
switch the Anthropic-compat base URL accordingly. Single provider with
prefix detection per AGENTS.md design integrity (no parallel mimo-code
provider needed).
Fixes#772
The local Ollama provider relied on /v1/models, which omits the
per-model context length and so reported a hardcoded 128k for any
model lacking a bundled reference (e.g. deepseek-v4-flash:cloud at
1M tokens). Now query /api/show per discovered model, parse
model_info.<arch>.context_length, and fall back to 128k only when
the call fails or the field is absent. Results are cached per model
id so repeated fetchDynamicModels calls do not refetch.
Fixes#847
Z.AI's Anthropic-compatible proxy at api.z.ai/api/anthropic deserializes
tool_result blocks into a Python class whose code path accesses .id, even
though Anthropic's schema only carries tool_use_id. Every request with a
tool_result returned 500 'ClaudeContentBlockToolResult object has no
attribute id'.
Detect the z.ai endpoint (provider === 'zai' or hostname api.z.ai) and
emit a non-standard 'id' field aliased to tool_use_id on tool_result
blocks. Other Anthropic-compatible endpoints (api.anthropic.com, xiaomi,
minimax, cloudflare gateway, etc.) keep the canonical schema.
Fixes#814
Claude marketplace plugins (e.g. context7@claude-plugins-official) ship
.mcp.json with the server map at the top level rather than under the
mcpServers key. The loader only accepted the nested shape, so install
appeared to succeed but no MCP tools were registered. Detect both shapes
and validate that each entry declares command (stdio) or url (HTTP/SSE)
before registering.
Fixes#851
DeepSeek V4 (v4-pro / v4-flash and reasoning-capable v3.x variants) reject follow-up
requests with 'The reasoning content in the thinking mode must be passed back to the
API' whenever a prior assistant tool-call turn lacks reasoning_content. The compat
flag that triggers placeholder injection was scoped to Kimi and OpenRouter-routed
reasoning models, so DeepSeek reached through api.deepseek.com, Deepinfra, Kilo,
NVIDIA NIM or Zenmux silently failed.
- Match DeepSeek family by provider, baseUrl, model id, or model name in
detectOpenAICompat (gated by model.reasoning).
- Recompute hasReasoningField after injecting the placeholder so the existing
'content === null && hasReasoningField -> ""' DeepSeek normalization actually
runs on tool-only assistant turns.
Fixes#883Fixes#810
The autoContinue post-compaction prompt echoed the summary's '## Next
Steps' heading, but that section is generated only from the compacted
tail; the kept ~20k recent tokens are not fed to the summarizer. When
the user pivoted within the kept window, 'Continue if you have next
steps.' anchored the model on the now-outdated plan instead of the
latest intent.
Move the prompt to prompts/system/auto-continue.md (per AGENTS.md
no-inline-prompts rule) and rewrite it to direct the model to re-read
the kept recent messages and follow the user's most recent request,
explicitly allowing it to stop when nothing remains.
Fixes#840
isPathInDirectory only normalized strings via path.resolve, so on Windows
when Bun is installed via Scoop (~/.bun is a junction to scoop\persist\
Oven-sh.Bun\.bun) the omp path from $which and the bunBinDir from
'bun pm bin -g' compared as different directories, causing 'omp update'
to take the binary-swap path instead of 'bun install -g' and fail with
EPERM unlinking omp.exe.bak (Bun has the running exe open). Layer
fs.realpathSync.native on top of the existing lexical guard, resolving
the file's parent dir so non-existent target paths still fall through.
Fixes#845
- Updated qwen3.5-plus and qwen3.6-plus to use openai-completions API with /v1 base URL path.
- Added idOverrides parameter to createOpenCodeApiResolution to allow per-model API routing corrections.
- Configured OPENCODE_GO_API_RESOLUTION to override Qwen models to openai-completions, matching the actual OpenCode Go endpoint table.
- Added a shared session path resolver that maps local:// URLs through local-protocol options, skips other internal schemes, and returns an absolute filesystem path for real files.
- Updated streaming-edit pre-cache and post-edit cache invalidation to use the shared resolver, preventing internal-scheme assertions while keeping filesystem-based flow for local plan files.
- Extended streaming-edit tests to confirm local:// plan edits complete without panicking and that auto-generated checks receive resolved absolute paths.
- Added support for batch PR operations by accepting `pr` as string or array and dropping `worktree` input.
- Updated `pr_view` and `pr_diff` to normalize PR IDs, process multiple PRs in parallel, and emit combined summaries.
- Refactored checkout into `checkoutPullRequest`, added repo-locking, fixed worktree paths, and summary metadata outputs.
- Updated `remote.add` handling with URL-aware idempotency and per-repo queueing for serialized git mutations.
- Added temp-home test scaffolding and expanded tests for batched PR flows and remote add conflict/no-op cases.
The task description prompt was compressed in c6cab807d — the explicit 'Current input mode' and 'Every assignment must stand on its own.' phrasings are gone. Replace those literal-string assertions with checks against the rendered mode-dependent content that actually differentiates the modes ('context` or `assignment`' for schema-free, 'each `assignment`' for independent). Mode-driven schema differences remain covered by the surrounding parameter-bullet assertions.
- Removed example and verbose description metadata from many debug tool schema properties while preserving field types and names.
- Simplified Enum declarations for `action` and `access_type` by using direct enum lists.
- Normalized several numeric and string field definitions in the schema without changing parameter names.
- Added a PROMPT_TASK_LIMIT constant set to 20 in the recipe runner model.
- Updated buildPromptModel to include only the first 20 tasks from each runner when generating prompt tasks.
- Added an optional hiddenTaskCount field to PromptRunnerModel to represent truncated tasks.
- Lowered the default session scan limit for edits and tools from 100000 to 1000 and updated usage text to match.
- Reworked grand-total and per-tool table rendering in session-stats to compute aligned column widths and include an aggregated "(N others)" row.
- Tuned the session-stats release profile by replacing full stripping with line-tables-only debug symbols and disabled split debuginfo stripping.
- Added a `/context` slash command flow from registry to interactive-mode command dispatch.
- Added `handleContextCommand()` to the mode context interface and command-controller wiring.
- Added context usage breakdown utilities, cell allocation, and 20x10 usage rendering for token categories.
- Reworked compaction token estimation to use tokenizer counts, role aggregation, image token estimates, and fallback handling.
- Exported `resolveThresholdTokens()` as a public compaction helper.
- Updated parseChunkUsage to subtract prompt_tokens_details.cache_write_tokens from prompt token input so OpenRouter write tokens are not misclassified as billable input.
- Set cacheWrite and total token counts to include cache-write usage, while preserving cache-read behavior from existing cached_tokens handling.
- Added OpenRouter attribution tests verifying cacheWrite and cacheRead totals for write-heavy and cache-warm prompts.
- Added a new `tokens` module in `crates/pi-natives` using `tiktoken-rs`, exposing `count_tokens` and `count_tokens_batch` through N-API with `Encoding::O200kBase` as the default.
- Exported the `Encoding` enum plus token counting functions in `packages/natives/native` TypeScript declarations and JS runtime bindings.
- Registered `tiktoken-rs` in `crates/pi-natives/Cargo.toml` and updated `Cargo.lock` with its resolved dependency entries.
- Extended `Usage` typing with `reasoningTokens`, `cttl`, and `server` fields for richer token accounting.
- Added conditional `reasoningTokens` output across OpenAI and Google providers when token counts are positive.
- Updated `parseChunkUsage` and OpenAI usage parsing to prevent `reasoning_tokens` double-counting.
- Added Anthropic usage extras mapping to emit TTL and server tool counters only when non-zero.
- Preserved Anthropic cache TTL and server counters when later usage events omit those fields.
- Added usage-attribution tests for parse helpers, missing fields, and zero-value merge scenarios.
- Allowed bare `LidA..LidB` to recover a missing-range-delete typo and accepted `|` as a legacy range replacement separator while validating ranges and replacement text.
- Enabled indented hashline statements to parse as replacement edits and added a preflight pass that validates all atom sections before any file write occurs.
- Updated hash-mismatch messaging, prompt wording, and tests to reflect hash-only rebase candidates and the new range/section behaviors.
- Updated the Atom grammar to parse replacement blocks via a set block rule that no longer targets range replacements only.
- Extended continuation preprocessing to allow backslash lines after single-line replace operations (including legacy `@` and `|` forms) while preserving the active-replacement check.
- Added tests and prompt documentation for single-line continuation cases and for rejection of backslash-like text outside an active replacement.