Stripped advisory wrapper tags in the collab web Markdown renderer while preserving escaped advisory content and existing raw HTML escaping.
Added a regression test for assistant advisory blocks in the collab transcript renderer and documented the fix in the collab-web changelog.
Fixes#3559
Hidden slider means the operator made no choice; a singleton cycle built around the active plan model must not be pinned as executionModel, otherwise approval re-applies the plan model after #exitPlanMode restored the pre-plan one.
Added regression coverage for the plan-only role configuration.
Refs #3554
- Ensure `-exec` commands run in the shell current working directory rather than the host process CWD.
- Update operand path resolution in `find` matchers to use the display path, matching behavior expected by shell-integrated utilities.
- Add an integration test to verify path substitution and execution context for shell-integrated `find`.
- Introduced `pi-uutils-ctx` to manage thread-local I/O redirection, environment context, and path resolution for in-process utilities.
- Integrated a suite of vendored coreutils (cat, find, grep, head, ls, mkdir, mv, rm, sort, tail, uniq, wc) as shell builtins.
- Added infrastructure for command cancellation, streaming I/O, and locale-aware error reporting within the shell environment.
- Configured shell-side dispatch logic to route commands through the new utility context, enhancing performance and binary integration.
Same-model role with an explicit thinking suffix that differs from the pre-plan thinking now passes through applyRoleModel instead of being treated as an implicit match.
Added regression coverage for the sonnet:off vs pre-plan thinking-high case.
Refs #3554
Compared the selected approval tier against the model restored after plan mode instead of the active plan-mode tier.
Added regression coverage for keeping the active planning model selected on approval.
Fixes#3554
- Implement `drainMicrotasksUntil` helper to manage async race conditions during tests.
- Add fake timer validation to ensure the first-item watchdog is cleared when the upstream source rejects.
- Restore real timers in `afterEach` to prevent test contamination.
Review catch: the hoisted twin emit set `params.preserve_thinking = true`
unconditionally for any Qwen + local target, but NVIDIA NIM's request
schema is `additionalProperties: false` and rejects unknown top-level
fields with HTTP 400 — the same #2299 quirk the catalog already routes
around by sending `enable_thinking` only under `chat_template_kwargs` for
the `qwen-chat-template` dialect. The NIM Qwen Q on a loopback baseUrl
(qwenPreserveThinking auto-enables) would have been broken before it
reached the kwargs copy.
Mirror the dialect split that already gates `enable_thinking`:
- thinkingFormat === 'qwen' (llama.cpp / Alibaba compatible-mode): set
BOTH top-level `preserve_thinking` and the `chat_template_kwargs`
mirror, like before.
- thinkingFormat === 'qwen-chat-template' (NVIDIA NIM, vLLM/SGLang's
chat-template-kwargs path): set ONLY
`chat_template_kwargs.preserve_thinking`, leave the top-level field
unset so NIM doesn't 400.
Flipped the existing NIM wire pin to assert top-level `preserve_thinking`
is undefined (and `enable_thinking` is undefined too — it was already
routed to kwargs by the catalog). Changelog reworded to document the
dialect split alongside the NIM rationale.
Verification:
bun --cwd packages/ai test ./test/issue-3528-repro.test.ts → 26/26 pass
bun --cwd packages/ai test ./test/issue-3434-repro.test.ts ./test/deepseek-reasoning-content.test.ts ./test/ollama-thinking-disable.test.ts ./test/openai-compat-policy.test.ts ./test/openai-completions-compat.test.ts ./test/openai-completions-tool-result-images.test.ts ./test/issue-967-vision-guard.test.ts → 111/111 pass
bun --cwd packages/catalog test → 325/325 pass
bun --cwd packages/ai check:types → clean
Allowed resource-server fallback OAuth discovery to accept authorization-server metadata whose issuer differs from the resource URL while keeping issuer matching for advertised auth-server candidates.
Added an Atlassian-shaped regression test so the fallback path no longer returns null.
Fixes#3551
Review catch: `applyChatCompletionsCompatPolicy` emitted preserve_thinking
only inside the `if (reasoning.enabled)` branch, so it never reached the
wire for the most common local-llama.cpp case: models built by
`discoverOpenAICompatibleModels` carry `reasoning: false` (the generic
/v1/models endpoint can't advertise the capability), `model.reasoning ===
false` short-circuits the reasoning encoder, and the request still shipped
`reasoning_content` (via replayReasoningContent, correctly ungated since
#3532's b6d81f22) for the template to strip <think> from older turns
anyway.
Hoist the emission above the early-return + reasoning-state branches so it
fires whenever `compat.qwenPreserveThinking` is set, regardless of
`reasoning.enabled` / `reasoning.disabled`. Three cases now covered that
were missed before:
1. Discovered local Qwen (reasoning: false in spec) — the case the
reviewer flagged.
2. Caller-disabled reasoning (/think off) — the kwarg is a HISTORY
rendering knob, not a per-turn switch, and the slot still holds
<think> tokens from earlier turns. Stripping them on re-render
invalidates the cache at the first historic <think>.
3. Forced-tool-choice / DeepSeek-style auto-disable — same history
reasoning as (2).
The qwen-template-false branch in the same function now spreads existing
chat_template_kwargs when setting enable_thinking: true so the
preserve_thinking entry hoisted above survives the merge (the old bare
assignment would have clobbered it).
Three new wire pins added on top of the existing five:
- Discovered local Qwen (reasoning: false, no reasoning option):
preserve_thinking: true rides; enable_thinking stays unset (server
falls back to template default).
- Discovered local Qwen + disableReasoning: preserve_thinking still
rides — the model's reasoning gate doesn't suppress history rendering.
- NVIDIA NIM Qwen on loopback (qwen-chat-template dialect):
chat_template_kwargs carries both enable_thinking and
preserve_thinking; the spread merge keeps both.
Existing "does NOT emit when disabled" pin flipped to "emits even when
disabled" to match the corrected history-knob semantics.
Verification:
bun --cwd packages/ai test ./test/issue-3528-repro.test.ts → 26/26 pass
bun --cwd packages/ai test ./test/issue-3434-repro.test.ts ./test/deepseek-reasoning-content.test.ts ./test/ollama-thinking-disable.test.ts ./test/openai-compat-policy.test.ts ./test/openai-completions-compat.test.ts ./test/openai-completions-tool-result-images.test.ts ./test/issue-967-vision-guard.test.ts → 111/111 pass
bun --cwd packages/catalog test → 325/325 pass
The reload-cache regression fixture used ReturnType<typeof loadConfig>,
which violates the repository style rule banning ReturnType<>. Import and
use the explicit LspConfig type instead.
Fixes#3546
getConfig() in packages/coding-agent/src/lsp/index.ts cached the first
loadConfig() result per cwd permanently. If .omp/lsp.json, root markers,
or plugin LSP configs were added after the first LSP call, they stayed
invisible for the remainder of the process lifetime — even after the
user explicitly requested 'reload *' — because the reload handler
operated on the same stale config object retrieved at the top of
execute().
The reload-workspace branch now deletes the per-cwd cache entry and
re-runs getConfig() before iterating servers, so the refresh behaves as
the prompt documents. The cache is repopulated by the fresh read, so
subsequent calls still avoid the disk hit until the next 'reload *'.
Fixes#3546
The Qwen3 / Qwen3.6 chat template strips <think>...</think> from every
assistant turn whose loop.index0 <= ns.last_query_index, so the moment a
new user message (the user's real next prompt OR the auto-learn
capture-at-stop nudge) lands, every prior assistant turn becomes 'older'
and is re-rendered without its <think> block — diverging from the
generation tokens still in the local slot's KV cache and forcing full
prompt re-processing on SWA models.
Sending reasoning_content alone (the #3528 fix) does not help: the
template's older branch renders only `content`, never the
reasoning_content field. The official Qwen3.6 fix is
`preserve_thinking: true`, which makes the template render
<think>\n{reasoning_content}\n</think>\n\n{content} for every assistant
turn regardless of position.
- packages/catalog/src/compat/openai.ts: new qwenPreserveThinking shared
compat flag, auto-enabled when the resolved thinkingFormat is `qwen`
or `qwen-chat-template` AND replayReasoningContent is on (the four
local provider ids plus loopback / RFC1918 / *.local baseUrls).
Responses-API builder pins it false — it's a chat-template knob,
irrelevant on the Responses surface.
- packages/ai/src/providers/openai-shared.ts: chat-completions encoder
emits preserve_thinking: true alongside enable_thinking: true in both
Qwen disable-mode branches (twin top-level + chat_template_kwargs
emission so llama.cpp / vLLM / SGLang and Alibaba's compatible-mode
wire shapes all pick it up). Stays off when thinking is disabled.
- packages/ai/test/issue-3528-repro.test.ts: nine new pins covering the
auto-detection matrix, the wire emission on local Qwen + thinking, the
cloud-Qwen / reasoning-disabled negative cases, and the
explicit-override escape hatches in both directions.
- Backfilled qwenPreserveThinking: false on three hand-rolled
ResolvedOpenAICompat fixtures so the required field stays satisfied.
- AI + catalog changelog entries under ## [Unreleased].
Verification:
bun --cwd packages/ai test ./test/issue-3528-repro.test.ts ./test/openai-completions-compat.test.ts ./test/openai-completions-tool-result-images.test.ts ./test/issue-967-vision-guard.test.ts → 87 pass
bun --cwd packages/ai test ./test/issue-3434-repro.test.ts ./test/deepseek-reasoning-content.test.ts ./test/ollama-thinking-disable.test.ts ./test/openai-compat-policy.test.ts → 37 pass
bun --cwd packages/catalog test → 325/325 pass
bun --cwd packages/ai check:types && bun --cwd packages/catalog check:types → clean
Fixes#3541
StdioTransport.connect() unconditionally passed detached:true to
Bun.spawn. POSIX needs this so terminal job-control signals (SIGTSTP,
SIGTTIN) cannot stop stdio servers such as chrome-devtools-mcp; Windows
has no equivalent signals, but detached:true maps to
CreateProcess(DETACHED_PROCESS), which strips the parent's inherited
console. windowsHide:true (#3536) hides the direct child's window but
nothing else — when the direct child was a hidden cmd.exe wrapper that
later spawned a console grandchild (node wrapper, npx.cmd -y mcp-remote,
similar nested shells) the grandchild allocated a brand-new visible
conhost and its stdout no longer routed through OMP's pipe. The proxy
terminal reported the MCP bridge was healthy while OMP timed out
waiting for the MCP initialize response.
Move detached into StdioSpawnCommand alongside windowsHide so every
platform-derived spawn flag is resolved in one place:
resolveStdioSpawnCommand returns detached:false on every Windows return
shape (direct .exe, cmd.exe-wrapped batch/unresolvable command, npm
cmd-shim launched through node) and detached:true on POSIX. connect()
consumes the resolved flag. Add a regression test for the reporter's
exact shape (cmd.exe /C node wrapper) and assert detached on every
existing Windows / POSIX case so the contract cannot regress per-path.
Fixes#3544
- Surface the exception type and message in the error display by prepending the formatted error string to the traceback array.
- Prevent the host from hiding the actual error by ensuring the traceback is not empty, consistent with other language runners.
- Applied the oh-my-pi branding identity including a custom font family, dark color palette with OKLCH colors, and brand gradient wordmark.
- Replaced static background with an ascii grid pattern and radial light accents.
- Improved layout aesthetics with frosted-glass card effects, subtle drop shadows, and refined animations for success/error states.
- Updated the authentication page title and status display with improved accessibility and design coherence.
`discoverOAuthEndpoints` probes `/.well-known/oauth-authorization-server` at
the origin root before path-prefixed candidates and returns on the first hit.
Plane hosts a root issuer (`https://mcp.plane.so/`) at origin root and a
separate path-scoped issuer (`https://mcp.plane.so/http`) at the path-prefixed
well-known. The `/http/mcp` endpoint advertises only the path-scoped issuer
through protected-resource metadata, so discovery should follow that issuer's
metadata — instead it accepted the wrong origin-root document and routed the
grant to `https://mcp.plane.so/authorize`, which rejects every request with
`server_error=An unexpected error occurred` before the consent screen.
RFC 8414 §3.3 requires the metadata's `issuer` to equal the URL the client
used to construct the metadata URL. Validate it in the discovery loop: when
the queried well-known is the official authorization-server or OpenID Connect
document, skip metadata whose `issuer` doesn't match (after trailing-slash
normalization). Documents without an `issuer` field keep the existing
permissive behavior so legacy/nonstandard servers continue to work.
Verified live against `https://mcp.plane.so/http/authorize` with the same
client/PKCE: pre-fix `302 -> /callback?error=server_error`, post-fix
`302 -> /http/consent?txn_id=…`. Adds an `oauth-discovery.test.ts`
regression suite covering Plane's wrong-issuer origin-root document plus
trailing-slash and no-issuer paths.
Fixes#3537
Set windowsHide for every Windows stdio MCP spawn path so direct .exe servers no longer open a visible cmd.exe window.
Added a regression test covering direct Windows executable MCP server launch options.
Fixes#3535
`openAICompletionsReplaysUnsignedThinking` short-circuited to false
when `model.reasoning` was false, ahead of evaluating
`compat.replayReasoningContent`. Since `discoverLlamaCppModels` /
`discoverOpenAIModelsList` hardcode `reasoning: false`, a cross-API
switch into a discovered local llama.cpp Qwen / DeepSeek target
demoted the prior turn's thinking block to plain text, leaving
`convertMessages` nothing to surface as `reasoning_content` — exactly
the cache-stable `<think>` prefix the new compat flag is meant to
preserve.
Reorder the predicate so the local-host replay short-circuits to true
BEFORE the `model.reasoning` gate. Cloud-style flags
(`requiresReasoningContentForToolCalls`, `thinkingFormat === "zai"`)
still require `model.reasoning` — they govern provider-specific
validation that only fires in thinking-engaged requests. Add a
cross-API regression pin: Anthropic-source thinking block + opaque
continuation signature flowing into a discovered local target
(`reasoning: false`) must arrive as `reasoning_content` on the wire,
with the source signature stripped.
Addresses chatgpt-codex review on #3532.
The discovery paths for llama.cpp, LM Studio, and openai-models-list
hardcode reasoning: false because the upstream /models endpoints don't
advertise the capability. The original gate
Boolean(spec.reasoning) && (LOCAL_PROVIDER || loopback)
therefore left replayReasoningContent off for the common setup the bug
targets — a discovered Qwen / DeepSeek model on local llama.cpp — and
the stream parser still recorded the upstream's reasoning_content
deltas as thinking blocks, so #3528 reproduced unchanged.
Drop the spec.reasoning gate. The encoder's own
'if (nonEmptyThinkingBlocks.length > 0)' guard ensures the flag stays a
no-op for pure-text turns, so always-on for local hosts adds nothing to
non-reasoning histories. Flip the matching test and add an encoder pin
that mirrors the discovery setup (reasoning: false + actual thinking
block must still ride as reasoning_content).
Addresses chatgpt-codex review on #3532.
LiteLLM defaults to http://localhost:4000/v1 and is the only built-in
`openai-completions` provider that forwards to an unrelated upstream
(OpenAI, Anthropic, …) rather than running a chat-template renderer
itself. The loopback auto-detection from the original #3528 fix would
push `reasoning_content` to those upstreams, which gain no KV-cache
benefit and may 400 on the extra field.
Add a `PROXY_OPENAI_COMPAT_PROVIDERS` deny set (currently just
`litellm`) that excludes proxy ids from both the provider allow-list
and the loopback heuristic. Users running a custom proxy in front of a
llama.cpp-style backend can still opt in via
`compat.replayReasoningContent: true`.
Addresses chatgpt-codex review on #3532.
The new required field on ResolvedOpenAISharedCompat surfaced as a
TS error in three test fixtures that build a literal compat record
instead of resolving it through buildOpenAICompat. Backfill the false
default so the tests stay in lockstep with the shared type.
Local llama.cpp / LM Studio / vLLM / Ollama (openai-completions mode)
and custom providers on loopback/RFC1918 baseUrls re-tokenize the entire
chat-template prompt every request. Qwen3 / DeepSeek-R1 / GLM templates
reconstruct the prior assistant turn's `<think>…</think>` block from
`reasoning_content`; dropping the field re-renders the assistant turn
without thinking content, the rendered tokens diverge from the slot's
existing KV cache, and llama.cpp falls back to full prompt re-processing.
The auto-learn capture-at-stop nudge (#3504/#3505) made this reproduce
on every turn for thinking-enabled local models: the reporter's wire
captures show `cached_tokens` collapsing from 38650 (req11) to 0
(req12) the moment the assistant reply re-enters history as a context
message.
Add `OpenAICompat.replayReasoningContent`, auto-enabled in
`buildOpenAICompat` whenever `spec.reasoning` is set AND the model is on
a known local provider id or a loopback / RFC1918 / `*.local` host. The
`openai-completions` encoder gets a fourth thinking-block branch that
emits `reasoning_content` on every reasoning-engaged assistant turn
(not just tool-call turns), honoring the streamed signature when it
identifies a recognized wire field and falling back to the configured
`reasoningContentField` otherwise. `transformMessages` learns the new
flag so cross-API replays into local llama.cpp targets preserve
unsigned thinking blocks.
Fixes#3528
refreshMCPOAuthToken now takes a trailing { authorizationUrl, stripSameOriginResource } options object (issue #3502 follow-up). Update the two stale per-profile binding assertions to expect it, and assert the fallback resource is not persisted (resource: undefined) since it is re-derived from config.url on each refresh.
- Clarify the auto-advancement logic for the in-progress pointer in the documentation and output.
- Add an overall completion count summary to the task list view.
- Update the list format to use standard checkbox indicators and explicit tags for task statuses.
Mid-turn renderSessionContext (settings overlay close, focus attach during streaming) now hands the rebuilt todo snapshot back to the EventController via the new inheritDisplaceableTodo method instead of sealing it. Idle rebuilds keep the historic seal path.
Added a regression test that asserts the trailing todo snapshot is published to the controller and stays displaceable while session.isStreaming is true.
Fixes#3516