18666 Commits

Author SHA1 Message Date
Oleg Pulatov 71ebe88850 fix(coding-agent): explain external thinking prelude 2026-08-18 19:51:49 +02:00
qiyi71w 1213f52093 style(session): satisfy biome and SkillPromptDetails typing 2026-08-19 01:50:25 +08:00
qiyi71w ecd800ef8f fix(session): title user /skill invocations from name and args
Session titles ignored /skill:<name> <args> because the invocation is a
custom skill-prompt, not a user turn. Feed the chip or reconstructed
/skill line into first-title generation and replan context, never the
expanded SKILL.md body.
2026-08-19 00:59:31 +08:00
Samuel Reed 301c1879b2 fix: blocker advisory always steers a new turn, immune window exempts it
The post-interrupt immune-window downgraded end-of-turn blockers to non-interrupting asides, so a blocker that means the agent handed off broken work never woke a new turn. Concerns keep the cooldown; blockers now bypass it and steer a triggered turn, consistent with the #5628 blocker-after-terminal-answer exception.
2026-08-18 12:22:09 -04:00
Daniel Young fb7c2e1a52 [mcp] Expand extension package environment
Apply the same recursive placeholder expansion used by native MCP configs before extension-package servers are validated and surfaced. This prevents stdio credentials and remote headers from reaching servers as literal placeholders.\n\nSolves: Extension-package MCP environment expansion\nTests: bun test packages/coding-agent/test/discovery/omp-plugins.test.ts
2026-08-18 12:14:25 -04:00
Daniel Young 480d72ffd2 Merge branch 'can1357:main' into main 2026-08-18 12:10:24 -04:00
poorpaper 6862ac08f0 fix(settings): hide excluded search providers from summary 2026-08-18 22:30:03 +08:00
Daniel Young 1496a139ce [mcp] Revert plugin env expansion
Keep host-specific secret injection in the maintained plugin layer instead
of carrying a behavioral divergence in the OMP fork.

Solves: Unwanted fork maintenance for MCP injection
Tests: Reverts only drycode/oh-my-pi PR #1
2026-08-18 08:50:06 -04:00
Daniel Young a52c2fc2fb [mcp] Expand Claude plugin stdio env (#1)
Claude marketplace plugins may reference runtime secrets in stdio
server environment values. Resolve those placeholders after plugin-root
substitution so child processes receive credentials instead of literal
template strings.

Solves: Stdio MCP credentials remain unexpanded
Tests: Claude plugin discovery tests; coding-agent check; live CH auth
2026-08-18 08:49:10 -04:00
bnivanov f5976d7129 fix(session): resume Cursor idle stalls after unmarked MCP results
The idle watchdog aborts the request signal and cursor.ts closes
the Connect stream, so there is no in-flight server exec to race.
Unmarked MCP/todo blocks can continue once every emitted call has
a matching result, same as HTTP/2 RST.
2026-08-18 14:48:45 +02:00
bnivanov 271e7ba892 fix(cursor): answer interactionQuery so hosted fetch can continue
Cursor hosted web search / Exa / unnamed field-9 WebFetch send
interaction_query and block the Run RPC until the client writes
interaction_response. Dropping the frame leaves the HTTP/2 stream
alive on heartbeats that are not semantic progress, so the 300s
idle watchdog aborts with "Provider stream stalled while waiting
for the next event".

Approve network permission gates and reject interactive
ask / switch-mode / create-plan. Leave VM setup unanswered
rather than inventing a success result.
2026-08-18 14:48:45 +02:00
djdembeck 71d608aae6 fix(robomp): validate PR review comment anchors against the diff before submitting
A single comment on a line outside a diff hunk 422s the whole review on
GitHub, which the 422/500 fallback (ported here from the Forgejo work,
including payload shaping and commit_id) then degraded into plain issue
comments. Now we fetch /pulls/{n}/files, compute per-hunk anchorable line
sets, drop non-anchorable comments, fold them into the review summary, and
only fall back to issue comments for genuine 422/500s that survive
filtering. Files-fetch failure fails open (submit unfiltered).
2026-08-18 03:54:00 -05:00
can1357 8500092296 chore: bump version to 17.3.7
Retry: widened agent dequeue-hook deadline budgets from 25ms to 1s — the run loop checks the deadline before invoking dequeue hooks, so a cold or CPU-starved mock roundtrip expired the deadline first and the hooks never ran (deterministic failure in isolation, flaky under CI parallel load).
2026-08-18 11:34:10 +03:00
chuzui 21d8ef9fb3 fix(coding-agent): stop pinning worker subprocess cwd to the install dir
Workers spawned via resolveWorkerSpawnCmd ran with cwd anchored at the CLI
install directory and shared the agent's foreground process group. Terminal
cwd heuristics such as kitty's new_tab_with_cwd read the newest process in
that group, so new terminal tabs opened in
~/.bun/install/global/node_modules/@oh-my-pi/pi-coding-agent/dist while any
worker was alive. Spawn workers with the absolute host entry and inherit the
agent cwd instead; the bun-test fallback branch keeps its cwd-relative form.
2026-08-18 16:29:21 +08:00
can1357 adfa211bbf chore: bump version to 17.3.7
Retry: deflaked tui IME preedit border test (#5563) — VirtualTerminal.waitForRender's fixed 40ms sleep raced TUI's throttled render timer on starved CI runners, reading the pre-input frame; waitForRender now takes an optional settle predicate polled up to 2s and the test keys on the rendered content.
2026-08-18 11:24:23 +03:00
naruto cbd748046d fix(web-search): abort-protect browser fallback setup and teardown
browseHtmlPage wrapped its navigations in untilAborted(signal) but awaited
applyViewport, applyStealthPatches, and the finally-block page.close() raw.
When the shared headless daemon or the page target dies mid-setup, those
puppeteer calls never settle: the search hard timeout fires into a signal
with no listener at those await points, the provider promise hangs past
SEARCH_HARD_TIMEOUT_MS, and the turn never ends (only kill -9 recovers).

- Wrap applyViewport/applyStealthPatches in untilAborted(signal) so the
  existing hard timeout can abort a dead-session setup.
- Bound the teardown page.close() with a fresh 5s deadline; .catch() only
  covers rejection, not a hang, and the caller signal may already be fired.

Fixes #8865
2026-08-18 11:08:49 +03:00
Kenneth Watson 49543d798e fix(tui): detect SIXEL outside Windows Terminal
The startup graphics probe only ran on ConPTY hosts with WT_SESSION, so a
SIXEL-capable terminal that exposes no identifying environment variable
(foot exports TERM=foot and COLORTERM=truecolor only) resolved the
trueColor capability row, kept imageProtocol null, and rendered every
image as the "[Image: …]" text card.

The XTSMGRAPHICS branch also had its status inverted: per xterm ctlseqs a
reply of `CSI ? 2 ; Ps ; Pv S` carries Ps = 0 on success, and a terminal
without SIXEL reports a zero maximum geometry, so a successful reply was
read as unsupported.

Drop the dead DA1 half of the probe with it: ProcessTerminal swallows
every `CSI ? … c` reply for the whole session so a late one cannot leak
into the composer (#8542), which means the attribute list never reached
the probe's input listener on any platform. The bare `CSI c` it wrote was
also unaccounted for in the DA1 sentinel FIFO, so its reply consumed
another probe's sentinel.

PI_FORCE_IMAGE_PROTOCOL, including its off/none kill switch, still wins
over the probe.
2026-08-18 14:47:20 +08:00
唐鳳 a073cd1763 Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-18 13:42:47 +08:00
Vu Anh Nguyen cea6ecd3d7 fix(tui): render xdev tool images inline 2026-08-18 12:09:39 +07:00
Audrey Tang f7df5d4970 fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
2026-08-18 12:19:11 +08:00
ata 602c89ae80 docs(changelog): collapse HUD label notes into one Unreleased bullet
Keep the 17.3.6 notes verbatim and record the HUD/handle split as a
single Unreleased fix.
2026-08-18 13:48:44 +10:00
ata 96cb5b5d95 fix(task): treat case-insensitive Name-N labels as handle echoes
The exact-echo check used accent-insensitive compare, but the collision
suffix path was a case-sensitive startsWith. AuthLoader-3 vs authloader
therefore leaked through as a real description.
2026-08-18 13:48:00 +10:00
ata 4509c128df style: biome-format HUD and task-label files
CI lint failed on line wrapping and extra blank lines in the HUD
role/label changes. No behavior change.
2026-08-18 13:48:00 +10:00
ata 8a83fb0e5d fix(task): stop using the spawn handle as the HUD description
The first HUD commit hid Name: Name. The cause was earlier: task
name was copied into identity.label, which became progress.description
and skipped generateTaskLabel. Keep the handle for id allocation, but
only treat eval label as a real UI description so the tiny-model
summary can run.
2026-08-18 13:48:00 +10:00
ata 94417ec99b test(tui): use generic names in subagent HUD fixtures
Keep the role-badge and echoed-id cases, but drop the session-specific
spawn handle from the source tree.
2026-08-18 13:48:00 +10:00
ata 54ba7fa4ab fix(tui): treat Name-N HUD labels as echoed spawn ids
Collision suffixes such as HindsightMcpFunnel-3 were still shown as
Name-3: Name because the first pass only compared the raw id.
2026-08-18 13:48:00 +10:00
ata aacf42c7eb fix(tui): show subagent role and drop echoed HUD labels
The anchored Subagents list printed only `Id: description` and treated a
label that repeated the spawn handle as a real description. Show the same
⟨role⟩ badge as inline task rows and omit descriptions that only echo the id.
2026-08-18 13:48:00 +10:00
z80 e3475e6353 fix(config): make live settings reload safe 2026-08-17 21:56:21 -04:00
User 70857d1839 Preserve MCP tools across PlanYolo handoff 2026-08-17 18:28:23 -07:00
re2zero 00f5d4cfca fix: correct defaults.json blob (previous upload stored filename instead of content) 2026-08-18 09:19:21 +08:00
re2zero 1fefedce14 fix(biome): add .css to default fileTypes for CSS language support; add changelog entry
Addresses review feedback on the initial commit:
- restore the trailing newline at EOF stripped by the first edit
- add a Fixed entry for this change under ## [Unreleased] in the coding-agent CHANGELOG
2026-08-18 09:17:34 +08:00
ranxianglei 9aefdb5f81 Merge branch 'main' into fix/onpayload-replacement-completions-bedrock-cursor 2026-08-18 09:13:09 +08:00
z80 ced78801ba fix(task): refresh model roles before agent discovery 2026-08-17 20:43:06 -04:00
Andrey Kuznetsov 8d34d1bc2e fix(ai): accept Perplexity OTP challenge token 2026-08-17 23:08:54 +00:00
Sunil Srivatsa b8dbf68613 fix(tiny): match title stop string against generated tokens only
StopOnTextCriteria decoded the last STOP_DECODE_WINDOW_TOKENS of the whole
sequence, so prompt tokens were eligible for matching. A prompt that itself
contains the stop string stops generation at the first generated token and
yields an empty title.

Anchor the window to the generation boundary by recording the first
generated index per batch entry. Existing local title models are
unaffected: with the assistant-prefill prompt shape, the example `</title>`
tags sit outside the 32-token window for normal messages, so no shipping
model changes behavior. The bug becomes reachable with any chat-level
few-shot prompt that places the stop string near the generation boundary.
2026-08-17 15:19:06 -07:00
Sunil Srivatsa f294d1b177 fix(auth): rotate when a ChatGPT account lacks the requested Codex model
A Codex request to a model the signed-in ChatGPT account is not entitled
to fails with "The '<model>' model is not supported when using Codex with
a ChatGPT account." That was classified as a plain provider error, so the
request failed outright even when a sibling account was signed in and
entitled to the model.

Classify that exact denial as an account-policy error, the same category
`cyber_policy` already uses, so the existing credential-rotation path can
reach an entitled account.

The match is deliberately narrow: it fires only for provider
`openai-codex`, only when the denied model in the message is the model that
was requested, and only for a bounded, non-null model identity. A denial
naming some other model does not trigger rotation, so an unrelated mention
cannot burn sibling credentials.
2026-08-17 15:07:53 -07:00
Sunil Srivatsa 49415712f2 fix(memory): separate extraction instructions from user input
The memory-extraction prompt concatenated its instructions, few-shot
examples, and the user message into a single user turn, so a small local
model could not distinguish instructions from input and frequently echoed
the Globex/weather examples instead of extracting facts.

Send the instructions as a real system turn and the raw text as the user
turn. The tiny worker protocol gains a systemPrompt field, and Mnemopi
completion input carries task metadata so the backend selects the right
prompt per call.

Drop the code-built MEMORY_EXTRACTION_TEMPLATE rather than porting it:
prompt text belongs in .md files, and resolveMemoryCompletionInput already
overrides that template for every extraction call, so Mnemopi rendered it
only for the result to be discarded.

Measured on ONNX q4 CPU, LFM2.5-1.2B memory extraction improved from 1/8
to 5/8 once the roles were separated.
2026-08-17 15:06:37 -07:00
Daniel Young 4dbb11b190 fix(discovery): honor Claude Code enabledPlugins for marketplace plugins
The claude-plugins provider read installed_plugins.json but never the
`enabledPlugins` map Claude Code keeps in ~/.claude/settings.json and
<project>/.claude/settings(.local).json. Two consequences: a plugin the
user switched off for a project still loaded there, and a local-scope
install enabled for a project never loaded unless the project directory
matched the install's recorded projectPath exactly.

Merge enabledPlugins across the same layers Claude Code consults (user
settings, then the active project root's and cwd's .claude/settings.json
and settings.local.json; later wins) and apply it in
listClaudePluginRoots: `false` hides the plugin, `true` opts a
local-scope install in regardless of projectPath. Untouched ids keep the
existing behavior. Contributing settings files join the cache key.

Solves: Claude marketplace plugins loading in the wrong projects
Tests: claude-plugins.test.ts — per-project off switch (local wins over
settings.json, other projects unaffected) and enabledPlugins:true
opt-in for a local install recorded under a parent directory
2026-08-17 17:50:10 -04:00
Sunil Srivatsa 5e4757f187 refactor(read): sniff the bytes before decoding them
Independent review noted that a file at or below the snapshot cap with a
text-like extension but binary content was fully decoded into three string
views and a line array before the sniff rejected it — roughly three times
the file size in transient allocations for output that is thrown away.

Split the loader: read the bytes, sniff those bytes, and derive the views
only for what survives. A refused 4MiB binary now costs one read and no
decode.

No observable change: the 966-case differential against the base commit is
byte-identical to the run before this split, still differing only in the four
intended BOM snapshot tags.

Also name in BlockContextSource.text the one shape that would violate its
same-content contract — a source object reused across two different line
arrays — since the contract is documented rather than enforced.
2026-08-17 14:44:02 -07:00
Tommy Liu 42171bf16e fix(catalog): add deepseek-v4-pro-0813 discovery limits
Alibaba Token Plan advertises both dated DeepSeek V4 snapshots, but only
deepseek-v4-flash-0731 had an entry in ALIBABA_TOKEN_PLAN_DISCOVERED_MODEL_LIMITS.
deepseek-v4-pro-0813 is not in ALIBABA_TOKEN_PLAN_STATIC_MODELS either (only the
undated deepseek-v4-pro is), so it fell through to `contextWindow: null` /
`maxTokens: null` — the #7486 symptom, still live for this one id.

Reasoning already worked: the `normalizedId.startsWith("deepseek-v4")` branch
gives it reasoning: true and the high/max effort ladder. Only the limits were
missing, so this is a one-entry fix at 1M context / 384K output, matching both
deepseek-v4-pro and deepseek-v4-flash-0731.

Extends the existing discovery test to advertise the id and assert its limits
and thinking config. Verified the test fails without the source change
(contextWindow/maxTokens come back null) and passes with it.

Refs #8847
2026-08-17 17:27:17 -04:00
Sunil Srivatsa bce0b35e44 docs(changelog): record the read single-pass change 2026-08-17 14:15:16 -07:00
Sunil Srivatsa 301908560e perf(read): materialize a local file once per read
The local text read path opened the same file for every consumer. A ranged
read of a file within the snapshot cap cost four opens and three decodes:
an 8KiB binary sniff, a streaming scan for the rendered window, a whole-file
read for bracket context, and another whole-file read to hash the snapshot.
Whole-file reads under the structural summarizer paid a fifth. Two of those
readers also ran normalizeToLF over the same bytes.

Read the bytes once at or below SNAPSHOT_MAX_BYTES and derive every view
from them: sniff the leading 8KiB of the buffer, slice the rendered window
out of it under the identical line and byte budgets, index bracket context
into its addressable lines, and hand the normalized text to the snapshot
store and the summarizer. Past the cap nothing wants the whole file, so the
streaming reader stays.

Line byte lengths are walked out of the buffer rather than measured on the
decoded strings, so reported byte counts and the truncation boundary stay
exact for content that is not valid UTF-8.

The buffered text is BOM-stripped for hashing, matching the decoder the
patcher's live read uses. A whole-file read of a BOM file previously hashed
its tag from BOM-bearing text, so the following edit only applied through
stale-hash recovery and told the model the file had changed externally when
it had not.

Also stop rejoining lines into a fresh whole-file string on the way to
tree-sitter when the caller still holds that text, and drop the unread
selectedBytesTotal accounting from the streaming reader.

Measured on 966 differential cases across CRLF, BOM, lone-CR, invalid-UTF-8,
oversized-line, empty, no-trailing-newline, multi-range and raw shapes: the
spurious recovery warning is the only behavioral difference. Raw reads, which
skip the tree-sitter parse that dominates everything else, get 30-45% faster
(2.7MB: 10.3ms -> 5.7ms); non-raw reads 1-3%.
2026-08-17 14:15:15 -07:00
can1357 644ad30d6e chore: bump version to 17.3.7
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.
2026-08-17 23:55:09 +03:00
Sunil Srivatsa 0e9f4dab52 docs(changelog): record the tree cache and boundary prune 2026-08-17 13:03:16 -07:00
Sunil Srivatsa 9169ac5c52 perf(pi-ast): prune subtrees that cannot hold a block boundary
collect_boundaries walked every node in the file even though the answer is
bounded by the visible window, which cost roughly twice the parse: on an
81KB source the walk was 8.95ms against a 4.32ms parse, and on 1MB it was
138.8ms against 87.6ms.

A node contributes a boundary only when one of its own endpoint lines is
visible, and both of those lines lie inside its raw row span. Every
descendant's span is contained in its ancestor's, so a span holding no
visible line rules out that node and everything beneath it. Skip such
subtrees with a binary search over the merged visible ranges.

The test is the raw span, not endpoint visibility: a node whose span merely
straddles the window has both endpoints outside it yet can contain a child
that opens exactly on a visible line. The raw span is also conservative
relative to node_content_end_line, so the prune needs no reasoning about
that newline adjustment.

Equivalence is proven differentially rather than argued: the pre-prune walk
is retained under cfg(test) and compared for exact Option<Vec<u32>> equality
across 4827 .ts/.py/.rs files and 38,616 comparisons over eight window
shapes, including whole-file-visible, past-EOF, disjoint ranges, empty range
lists and files that fail to parse. Zero mismatches. root.has_error() is
still evaluated on the whole tree before the walk, so pruning cannot change
a None verdict.

Measured on the built addon with a mid-file 40-line window, medians of 20,
against the parse cache alone: 81KB 13.4ms -> 4.45ms cold and 9.04ms ->
0.149ms warm; 1.06MB 188.1ms -> 55.8ms cold and 131.7ms -> 0.440ms warm.
2026-08-17 12:57:59 -07:00
Sunil Srivatsa 010eda7834 perf(pi-ast): cache parsed trees by content and language
`enclosing_block_boundaries`, `block_range_at` and `summarize_code` each
re-parsed the whole file on every call. The results are not cacheable —
boundaries depend on the caller's visible ranges, which differ per call —
but the `tree_sitter::Tree` is, so cache that instead and hand out
`ts_tree_copy` clones.

Keyed on (xxh64 of the source, source length, language). The hash is a
bucket selector only: a hit re-verifies the stored source against the
request byte-for-byte before returning the tree, so a collision costs a
re-parse and can never yield a tree built from other content. Language is
in the key because the same bytes parsed as TypeScript and as Python are
different trees.

Bounded at 12 slots and 4 MiB of retained source with LRU eviction;
sources above 4 MiB are parsed but never retained. `Tree` is `Send` but
not `Sync`, so entries sit behind a `Mutex` that is held only for a map
probe, a byte compare and a refcount bump, never across a parse or walk.

Error trees are cached like any other: `has_error()` is a property of the
tree, so the callers' own checks reach an identical verdict from a cached
tree, and repeated "does this parse" probes get the speedup too.

Measured (M4 Max, bazel-built .node, median of 20, 1-40 visible):
read.ts 81 KB 13.34 ms -> 8.86 ms on repeat; 1 MB synthetic 225.7 ms ->
138.3 ms. Parser::new + set_language measured at 0.30 us against a
3.91 ms parse, so no parser pooling.
2026-08-17 12:39:22 -07:00
can1357 0a912cc467 chore: bump version to 17.3.7 2026-08-17 22:29:25 +03:00
Hayden Evan 02533ff770 ci: retrigger checks after unrelated runner flakes
The previous CI reds were a truncated rust-overlay fetch in Evaluate
flake and a Julia kernel shutdown in julia-prelude.test.ts. Neither
path is in this change.
2026-08-18 02:19:42 +07:00
Pedro Paulo Vezza Campos ae804857dd test(lsp): isolate config tests from the developer's user lsp.json
`lsp regressions > detects pyright and pylsp in Windows virtualenv Scripts for
Python-only roots` fails on any machine that has ~/.omp/agent/lsp.json:

  expect(config.servers[server]?.resolvedCommand).toBe(localBin)
  Expected: ".../.venv/Scripts/pyright-langserver.exe"
  Received: undefined

The test was not asserting against the packaged defaults at all. loadConfig
walks the user config dirs (~/.omp/agent, ~/.pi/agent, ~/.claude) via
getConfigDirPaths, which resolves from os.homedir(). Any user lsp.json with a
`servers` block sets hasOverrides, which takes loadConfig off its auto-detect
branch and onto the override branch, where the user's rootMarkers replace the
packaged ones.

On this machine that file overrides pyright with

  rootMarkers: ["pyproject.toml", "uv.lock", "requirements.txt", "setup.py"]

which contains neither `pyrightconfig.json` nor `setup.cfg` — precisely the
two markers the test creates. loadConfig therefore returned zero servers.
`ruff` is absent from that file, keeps the packaged rootMarkers, and its
sibling tests pass, which is why only this one failed.

Confirmed by probing inside the test: hasRootMarkers() true and
resolveCommand() returning the .exe, while loadConfig().servers was {} — the
detection helpers were fine, the branch was not.

Point os.homedir() at an empty directory for every test in the file so they
see a pristine environment. Bun's os.homedir() reads the passwd entry rather
than $HOME, so the env var alone does not redirect the walk; both are set.

Verification: 76 pass / 0 fail (was 75 / 1) on a machine with a user lsp.json.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 11:59:45 -07:00
Can Bölük 7913d70254 Merge pull request #8840 from Jaaneek/fix/openai-compat-user-agent
fix(ai): send omp User-Agent on xAI chat only
2026-08-17 20:52:46 +02:00