Commit Graph

3223 Commits

Author SHA1 Message Date
Vu Anh Nguyen 12c666f4c0 Fix mermaid markdown test expectations 2026-04-24 06:47:56 +02:00
Vu Anh Nguyen a7845b3b86 Fix mermaid assistant test settings init 2026-04-24 06:47:56 +02:00
Vu Anh Nguyen 81fd90ecc0 Fix Mermaid markdown rendering on non-image terminals 2026-04-24 06:47:56 +02:00
Can Bölük fd10e51b64 Merge pull request #682 from cgrossde/fix/657_pi-protocol-sourcePath
fix: grep tool silently mangles pi:// URLs and produces misleading Path not found error
2026-04-24 06:36:12 +02:00
can1357 6cfca06d8e fix(edit): mark stale-hash rejections as 'Edit rejected' instead of an info-shaped message
The previous mismatch text ('1 line has changed since last read. Use
the updated LINE#ID references shown below.') reads like a successful
edit followed by an informational note. Providers whose tool-result
plumbing does not surface `isError: true` to the model (e.g. Qwen via
the OpenAI-completions shim, where mitmproxy traces show the model only
receives the content text) treat the response as a success and proceed
with stale state.

Reword the message to start with 'Edit rejected:' and explicitly state
'The edit was NOT applied' so the failure is unambiguous in plaintext.
The structural payload (>>> markers, updated LINE#IDs) is unchanged.

Fixes #742
2026-04-24 06:33:20 +02:00
can1357 7f81179b40 feat(config): add commands.enableOpencode{User,Project} settings
Mirror the existing `commands.enableClaudeUser`/`commands.enableClaudeProject`
schema entries so the OpenCode discovery provider exposes the same
user/project toggle surface as Claude. Default remains true to preserve
current behavior.

Fixes #661
2026-04-24 06:33:20 +02:00
can1357 cafa86a6cf fix(tools/sqlite): reject comments, terminators, and pagination keywords in where=
The structured SQLite helper interpolates `where=` directly into SQL.
A crafted clause like `where=1=1 LIMIT 1000000 --` could comment out
the helper's bound `LIMIT ? OFFSET ?`, returning the full table in
violation of the documented pagination contract.

Validate where= at the selector boundary and reject SQL comments,
statement terminators, and pagination/attach/pragma keywords. Raw SQL
remains available via ?q=SELECT... for callers that need it.

Fixes #735
2026-04-24 06:33:20 +02:00
Can Bölük 98df8b72c7 Merge pull request #668 from robinmordasiewicz/fix/lsp-startup-noise
fix(lsp): filter informational log lines from LSP startup warnings
2026-04-24 06:31:53 +02:00
Can Bölük 84f6dc30f1 Merge pull request #740 from kagura-agent/fix/tools-equals-syntax
fix(cli): support --flag=value equals syntax for all CLI flags
2026-04-24 06:11:50 +02:00
can1357 bd07ebed53 fix(coding-agent/vim): force-emit kbd steps so pendingInput streams reliably 2026-04-24 06:10:01 +02:00
can1357 1fb22e1e9a chore: bump version to 14.2.1 2026-04-24 05:55:59 +02:00
can1357 661449df09 test(coding-agent/core): removed stale chunk-tree test suite from core
- Removed stale test suite file `packages/coding-agent/test/core/chunk-tree.test.ts`.
2026-04-24 05:54:10 +02:00
can1357 cb134f6446 refactor(find): consolidate path relativization and drop gitignore fallback
Extracts a single `formatMatchPath` helper used by both the
fast-glob/native code paths and the streaming onMatch callback so
relative paths, trailing-slash handling, and directory markers are
produced consistently. Also drops the retry-without-gitignore fallback
when the gitignored pass returns zero matches, so a broad hidden-file
pattern that is fully ignored stays fully ignored instead of silently
flipping gitignore off on the second attempt.
2026-04-24 05:49:16 +02:00
can1357 c5666e153f feat(grep): route comma-separated explicit files to exact-file grep
resolveMultiSearchPath now reports `exactFilePaths` when every token
resolves to a plain file (no globs, no suffix glob) and accepts a
single resolvable token so partially-missing lists still search the
resolvable subset. grep iterates those exact files individually
instead of collapsing them into a brace-union glob, which preserves
the user's explicit file set even when siblings share a basename.

Also adds a small `[grep] match lines use ':'; context lines use '-'`
banner when context lines are rendered, and splits the per-file
rendering helpers so files with no remaining matches no longer emit
empty headers.
2026-04-24 05:49:10 +02:00
can1357 e51f0b5321 fix(ast-edit): detect stale previews on apply
Apply now recomputes per-file replacement counts from the actual apply
pass and compares them against the preview. If totals or per-file
counts drift (file changed between preview and apply, or apply matched
nothing), the tool returns an isError result explaining that the
preview is stale instead of silently claiming success with mismatched
numbers.
2026-04-24 05:49:02 +02:00
can1357 3c8aced08c feat(resolve): pass through source result details and mark failed applies
The resolve tool now preserves the underlying tool result's details on
ResolveToolDetails.sourceResultDetails instead of dropping them, and
the renderer distinguishes a failed apply ("Failed") from a user
discard ("Discard") so errored applies are no longer mislabelled.
2026-04-24 05:48:57 +02:00
can1357 36b21cfb16 docs(ast-grep): note C++ qualified call semicolon footgun
tree-sitter-cpp parses `ns::doThing($ARG)` without a trailing
semicolon as declaration-like syntax, so ast_grep returns no matches.
Tell agents up-front to include the statement semicolon (or use a
looser `$CALLEE($ARG)` pattern) instead of debugging silently empty
results.
2026-04-24 05:48:52 +02:00
can1357 3add2b8a3a feat(bash): surface timeout-clamp notice in tool output
When a bash tool call requests a timeout outside the allowed 1-3600s
range, the effective clamped value and the originally requested value
are now emitted as a notice appended to the tool output and exposed on
BashToolDetails via requestedTimeoutSeconds. The renderer shows the
clamped+requested pair inline in the timeout badge.
2026-04-24 05:48:46 +02:00
can1357 ab0f50281c fix(coding-agent/dap): handled unhandled waitForEvent rejections in stop outcome setup
- Reworked #prepareStopOutcome to collect stopped, terminated, and exited event waits before racing them.
- Attached noop rejection handlers to each pending wait promise to prevent unhandled rejection noise after the race settles.
- Returned Promise.race(promises) so stop outcome preparation still waits for the first relevant lifecycle event.
2026-04-24 05:18:54 +02:00
can1357 6b976f6eeb fix(coding-agent/lsp): serialized LSP writes and retried references after project load
- Queued outbound JSON-RPC messages behind a per-client promise queue to serialize writes.
- Added a project-load gate for project-aware LSP operations before diagnostics and reference lookups.
- Retried declaration-only references with a short delay until project metadata is available, then proceeded with normal results.
2026-04-24 05:08:11 +02:00
can1357 1a3a3830b3 fix(coding-agent): validated bare-path chunk deletes in chunk mode
- Relaxed chunk-mode parameter validation to accept `{path}`-only edits as valid delete operations.
- Updated the invalid-parameters help text to document accepted chunk delete payloads.
- Added a test that verifies a bare `{path}` edit removes the targeted chunk when null values are stripped.
2026-04-24 04:35:59 +02:00
can1357 50e354c745 style: format Codex web search default model array 2026-04-24 01:20:06 +02:00
can1357 a7d78a459b fix(coding-agent): default Codex web search to gpt-5-codex-mini 2026-04-24 01:18:51 +02:00
can1357 06697b3f28 chore: bump version to 14.2.0 2026-04-24 01:05:06 +02:00
can1357 ffb5ad4880 fix(coding-agent): refresh Gemini web search credentials
Fixes #449
2026-04-24 01:02:36 +02:00
can1357 bf6b5f4989 fix(coding-agent): choose current Codex web search model
Fixes #723

Fixes #724
2026-04-24 01:02:30 +02:00
can1357 5bfdc8d46d fix(coding-agent): abort ask in headless sessions
Fixes #681
2026-04-24 01:02:24 +02:00
can1357 a9ca2daaed fix(coding-agent): fail structured subagents without submit_result
Fixes #729
2026-04-24 01:02:18 +02:00
can1357 fe076b8cd7 fix(coding-agent): cap memory rollout input
Fixes #764
2026-04-24 01:02:13 +02:00
can1357 6b8b46c410 fix(coding-agent/edit): fixed apply_patch streaming preview parsing behavior
- Added a streaming parser path for apply_patch envelopes that tolerates incomplete patch bodies.
- Updated apply patch preview expansion to return best-effort hunks when the renderer is in partial mode.
- Added a renderer test confirming streaming apply_patch input shows file paths without end-marker parse errors.
2026-04-24 00:48:03 +02:00
can1357 7e9d568346 fix(coding-agent/tools): replaced tabs with spaces in diagnostic rendering output
- Applied tab sanitization to LSP diagnostic parsing, fallback entries, and rendered diagnostics output.
- Updated tool diagnostics formatting to strip tabs from parsed file, source, message, code, and summary fields.
- Added regression tests confirming rendered diagnostics contain no tab characters and preserve message text after sanitization.
2026-04-24 00:43:55 +02:00
can1357 8a7b0a5c06 feat(ai): changed spark edit tool resolution
- Added built-in model entries for gpt-5.5 and gpt-image-2, including updated context windows, token limits, and pricing.
- Updated generated-model policy application to set or clear applyPatchToolType based on inferred GPT-5 freeform rules.
- Changed Spark edit-mode resolution to return apply_patch by default, honoring explicit replace and strict-mode overrides.
- Added tests for GPT-5 freeform policy inference and Spark edit-mode default, variant, and strict-mode behavior.
2026-04-24 00:42:19 +02:00
can1357 601c029c20 fix: stabilize apply_patch custom tool integration 2026-04-24 00:25:21 +02:00
Hans Josephsen 2a367bf043 feat(coding-agent/edit): add codex apply_patch as a new edit mode
Slots a new "apply_patch" variant alongside the existing edit modes
(replace, patch, hashline, chunk, vim). The mode accepts a single input
string containing a Codex *** Begin Patch / *** End Patch envelope,
parses it with a new lenient parser (heredoc-tolerant), and fans each
file-op out to the existing executePatchSingle so LSP writethrough,
plan-mode guards, fs-cache invalidation and diagnostics are shared
with the patch mode.

Exposes both tool shapes from the spec: the JSON function-tool variant
(§1.2, {input: string}) and the OpenAI custom-tool / Lark-grammar
"freeform" variant (§1.1, raw patch string). The edit tool advertises
a Lark grammar via customFormat and a wire name via customWireName;
openai-responses emits it as a grammar-constrained custom tool when a
model opts in with applyPatchToolType: "freeform" in models.json.
custom_tool_call / custom_tool_call_output are plumbed end-to-end
through the shared responses code (emission, streaming, history
replay), and the agent-loop dispatcher matches tool calls by either
name or customWireName so returned calls route correctly.

Also threads preview/diff rendering for apply_patch through the TUI
(tool-execution + edit renderer) so streaming patches show per-file
diffs like the other edit modes.

Default edit mode is unchanged (hashline); opt in via edit.mode or
PI_EDIT_VARIANT=apply_patch.
2026-04-24 00:15:17 +02:00
can1357 ee05fb6c2d fix: system prompt regressions 2026-04-24 00:09:09 +02:00
can1357 31294d6b0f Merged PR #763 2026-04-24 00:04:30 +02:00
can1357 2e6dac791d Merged PR #761 2026-04-24 00:04:27 +02:00
can1357 5a0c33a2db Merged PR #757 2026-04-24 00:03:59 +02:00
can1357 46f5d9d63a Merged PR #756 2026-04-24 00:03:59 +02:00
can1357 8413b406c7 docs(coding-agent): document darwin binary signing fix 2026-04-23 23:05:35 +02:00
can1357 820776d0cc test(coding-agent): align prompt wording assertions 2026-04-23 23:05:24 +02:00
Miroslav Drbal ea113ea1d2 Restore MUST weight and variadic-capture guardrail in tool prompts
Review-pass findings:

1. python.md (P3) — "Put workflow explanations in assistant message or
   cell title" lost its **MUST** weight during compression. Agents were
   observed embedding explanations as in-cell comments, polluting the
   kernel and wasting tokens. Restored as "You **MUST** put workflow
   explanations in the assistant message or cell title — never inside
   cell code."

2. ast-grep.md + ast-edit.md (P3) — The variadic-capture warning that
   $$$NAME (three dollars) is correct and $$NAME (two dollars) is
   invalid was dropped during compression. LLMs trained on shell/regex
   conventions where $$ is common are prone to this mistake; ast-grep
   emits a generic parse error rather than a "bad metavariable"
   diagnostic, so the negative example has direct reproduction value.
   Added "Use $$$NAME, **NOT** $$NAME" to both files.

All 6/6 template tests pass; bun check clean.
2026-04-23 22:59:38 +02:00
Miroslav Drbal e04d2620df Compress and tighten system prompt; relocate SECTION_SEPARATOR helper
Further compresses packages/coding-agent/src/prompts/system/system-prompt.md
from 21,530 B to 16,545 B (−23% on top of the prior compression;
−9,060 B / ~−2,265 tok per turn vs origin/main) while restoring
RFC 2119 weight on every inviolable rule and reorganizing the
pre-yield discipline.

system-prompt.md

RFC 2119 keyword restoration

All rules under `# Contract`, Procedure §§2/6, and the tail `<critical>`
block use `**MUST**` / `**MUST NOT**` keywords that the file's preamble
pins to RFC 2119. The previous draft had downgraded 29 of these to
`Do **NOT**`, which reads as below-MUST-NOT against the same preamble.
Restored sites:
- `# Contract` — all 7 inviolable bullets
- `<critical>` tail block — all 4 safety rules
- `## 6. Verification` — mock ban and no-proof yield rule
- Procedure §2 — "search for existing examples" rule regains the
  PROHIBITED parallel-convention clause it lost
- `<dir-context>` — AGENTS.md read requirement regains MUST

Structural dedup

- `<source-of-truth>` and `<instruction-priority>` were two priority
  lists with overlapping scope. Merged into a single
  `<instruction-priority>` covering all 5 conflict-resolution levels.
- `<self-check>` and `<scope-check>` were two pre-yield checklists;
  `## 6. Verification` also had a third numbered "Before yielding,
  verify..." list with overlapping checks. Collapsed all three into
  one `<pre-yield-check>` section (incl. "output format matches the
  ask" from the old verification list). `## 6. Verification` now
  focuses on verification discipline (test rigor, mocks, run-scope).

Signal restoration

Added four anchors to `<design-checklist>` that were dropped from the
earlier `<code-integrity>` block and had no home in the new structure:
- Adversarial-caller / tired-maintainer self-questioning
- Cost-of-easy-path framing
- Inhabit-the-call-site self-review
- Persistence on hard problems (do not punt half-solved work)

Handlebars fix

The `### Tool priority` header was emitted unconditionally but its
body was wrapped in `{{#ifAny python|bash}}`, producing a bare heading
with no content when neither tool was available. Header now sits
inside the conditional.

SECTION_SEPARATOR helper relocation

The `SECTION_SEPARATOR` Handlebars helper is a generic section-header
formatter that was registered in coding-agent's
`config/prompt-templates.ts`. This coupled every template consumer to
a side-effect import of that module; a prior fix added
`import "./config/prompt-templates"` to `system-prompt.ts` so the
`/system-prompt` sub-path export would register the helper, but
sibling consumers (`task/executor.ts`, `task/template.ts`) worked only
by accident because the parent agent happened to load
`system-prompt.ts` first.

- Moved `sectionSeparator` function and helper registration to
  `packages/utils/src/prompt.ts` next to the other generic helpers
  (`xml`, `codeblock`, `ifAny`, `includes`, `not`, `jsonStringify`).
- `packages/coding-agent/src/config/prompt-templates.ts` now
  re-exports `sectionSeparator` from `@oh-my-pi/pi-utils/prompt` for
  the test that imports it via the coding-agent path.
- Removed the side-effect import from
  `packages/coding-agent/src/system-prompt.ts`.

Template typo fix

Also fixes a latent `SECTION_SEPERATOR` spelling (→ `SECTION_SEPARATOR`)
in the two subagent templates (`subagent-system-prompt.md`,
`subagent-user-prompt.md`). Before the move the typo was masked
because every call site and the helper shared the wrong spelling.
Post-move, the helper is canonical `SECTION_SEPARATOR` in pi-utils
and all templates match.

Verification

- `bun check` clean (TS + Rust, all 9 packages)
- `bun test packages/coding-agent/test/system-prompt-templates.test.ts`
  6/6 pass (was 4/6 failing pre-existing before the infrastructure fix)
- `bun test packages/coding-agent/test/tools/task-template.test.ts`
  4/4 pass (exercises the `sectionSeparator` re-export)
- `bun run format-prompts` clean
2026-04-23 22:59:37 +02:00
Miroslav Drbal 416c5494fd Compress system prompt and tool descriptions for token efficiency
Reduces per-turn token cost of prompt and tool metadata by ~2,160
tokens (~8.6 KB across 12 prompt files) without losing instructional
signal. Moves literal default values from description text into
TypeBox's native `default:` keyword, and updates the strict-mode
sanitizer to preserve default info for providers that strip the
keyword.

Prompt compression

|File                                      |before |after  |delta |
|------------------------------------------|------:|------:|-----:|
|system-prompt.md                          |25,605 |21,530 |−16%  |
|prompts/tools/ast-grep.md                 | 4,892 | 3,880 |−21%  |
|prompts/tools/ast-edit.md                 | 4,141 | 3,549 |−14%  |
|prompts/tools/bash.md                     | 3,770 | 3,093 |−18%  |
|prompts/tools/task.md                     | 7,297 | 6,711 |−8%   |
|prompts/tools/read.md                     | 2,701 | 2,309 |−15%  |
|prompts/tools/debug.md                    | 2,791 | 2,412 |−14%  |
|prompts/tools/find.md                     |   860 |   561 |−35%  |
|prompts/tools/grep.md                     | 1,226 | 1,038 |−15%  |
|prompts/tools/todo-write.md               | 2,600 | 2,402 |−8%   |
|prompts/tools/python.md                   | 2,013 | 1,916 |−5%   |
|prompts/tools/hashline.md                 | 4,280 | 4,135 |−3%   |

Changes are textual compression only — grammar-scaffolding removed,
redundant bullets stripped, repeated facts consolidated. All
behavioral contracts, safety rules, worked examples, and
tool-precedence directives are preserved. The system prompt retains
RFC 2119 invocation, XML tag semantics, adversarial-caller guidance,
persistence doctrine, unit-of-change and no-forwarding-addresses
rules, completeness contract, DRY-at-2, earn-every-line, trust-
internal-code, tool precedence, AST tool priority, pattern-syntax
cheatsheet (via ast-grep.md/ast-edit.md), tool persistence,
outside-in code integrity, default follow-through, and procedural
steps 1-7. AST tool docs retain the class-wrapper +
method_definition sel example (the most error-prone usage pattern).
bash.md critical section retains MUST-weight on the ast_grep /
ast_edit directives.

Schema: defaults as first-class metadata

Moved 18 literal default values from `(default: X)` description
text into TypeBox's native `default:` keyword across:
ast-edit.ts, ast-grep.ts, bash.ts, browser.ts, find.ts, gh.ts,
grep.ts, python.ts, read.ts, ssh.ts. Runtime-resolved placeholders
(cwd, pr-<number>) remain as text since they cannot be literal
TypeBox defaults.

Added:
- `minItems: 1` on ast-edit.ts `ops` array — machine-enforces what
  the .md previously stated only in prose; handler already rejects
  empty ops arrays, this adds the schema-level constraint.
- `find.ts` `pattern` description enriched with facts previously
  only in find.md (comma-separated lists, simple patterns recurse
  from cwd).

Strict-mode sanitizer: inline `default` into `description`

OpenAI's Structured Outputs strict mode rejects schemas containing
`default` with HTTP 422 ("default is not permitted"). Affects
openai, azure, github-copilot, openrouter, cerebras, together,
zenmux, and deepseek providers.

`sanitizeSchemaForStrictMode` in packages/ai/src/utils/schema/
strict-mode.ts now appends ` (default: X)` to the sibling
`description` before stripping the `default` keyword. Non-strict
providers (Anthropic, Google) still see the native keyword.

Rules:
- Inline is skipped when description already contains `(default:`
  (prevents double-inlining on recursive calls)
- Inline is skipped when no sibling description exists (no
  synthesis)
- Formatting: strings as-is (`cwd`), other values via
  `JSON.stringify` (matches the conventional text form)

CONSTRAINTS.md documents the inlining rule alongside the existing
keyword-strip rule.

Regression tests

packages/ai/test/schema-strict-mode.test.ts gains 7 `it` blocks:
- number/bool/string default types inline correctly
- falsy defaults (`false`, `""`, `0`) are not confused with absent
- `null` default goes through JSON.stringify branch
- double-inline prevention when description already says `(default:`
- no synthesis when no description exists
- nested object property with default (recursion + cache path)
- type-array `[T, null]` branch with default on outer schema
  (variant-materialization path)

23/23 schema tests pass, 541/541 ai-package tests pass,
`bun check` clean.

Rationale

|Metric                          |Value          |
|--------------------------------|--------------:|
|Per-turn prompt savings         |~2,160 tok     |
|Files touched                   |25             |
|Lines changed                   |+325 / −296    |
|Schemas migrated to `default:`  |18             |
|New regression tests            |7              |
2026-04-23 22:59:37 +02:00
Kagura f9522de33c fix(auth): let OAuth credentials override keyless provider flag from stale models.yml
When a user models.yml declares a provider with auth: none (e.g. from
an older version), getApiKey() returns the literal 'N/A' without ever
consulting authStorage — even when valid OAuth credentials exist from
a successful /login.

Check authStorage.hasAuth() before returning kNoAuth. If real
credentials are present, fall through to the normal auth flow.

Fixes #749
2026-04-23 22:59:33 +02:00
fettpl 14b08f89d4 fix: restore runnable darwin compiled binaries
fixes #754
2026-04-23 22:59:27 +02:00
can1357 4490332822 test(coding-agent): avoid inline imports in claude plugin discovery tests 2026-04-23 22:54:59 +02:00
can1357 5072d3df11 docs: add compiled binary autoload changelog 2026-04-23 22:54:40 +02:00
fcremo (Filippo Cremonese) 17df89086b fix: disable bunfig.toml and .env autoloading in compiled binary
Compiled Bun binaries unconditionally load bunfig.toml and .env from
the current working directory at runtime, before any application code
runs. Since omp is a coding agent that runs from arbitrary project
directories, it picks up foreign project configs -- most critically
preload directives that cause immediate crashes, but also a potential
security issue since preloads execute arbitrary code.

Add --no-compile-autoload-bunfig and --no-compile-autoload-dotenv to
the bun build --compile invocations (available since Bun v1.3.3).

Note: source installs via 'bun install -g' are still affected because
'bun run' has no equivalent flag. That remains an open issue.
2026-04-23 22:51:25 +02:00
Parsifa1 782a309f6f fix(coding-agent): fix untrusted path resolve 2026-04-23 22:51:01 +02:00