Commit Graph

3202 Commits

Author SHA1 Message Date
can1357 3add2b8a3a feat(bash): surface timeout-clamp notice in tool output
When a bash tool call requests a timeout outside the allowed 1-3600s
range, the effective clamped value and the originally requested value
are now emitted as a notice appended to the tool output and exposed on
BashToolDetails via requestedTimeoutSeconds. The renderer shows the
clamped+requested pair inline in the timeout badge.
2026-04-24 05:48:46 +02:00
can1357 ab0f50281c fix(coding-agent/dap): handled unhandled waitForEvent rejections in stop outcome setup
- Reworked #prepareStopOutcome to collect stopped, terminated, and exited event waits before racing them.
- Attached noop rejection handlers to each pending wait promise to prevent unhandled rejection noise after the race settles.
- Returned Promise.race(promises) so stop outcome preparation still waits for the first relevant lifecycle event.
2026-04-24 05:18:54 +02:00
can1357 6b976f6eeb fix(coding-agent/lsp): serialized LSP writes and retried references after project load
- Queued outbound JSON-RPC messages behind a per-client promise queue to serialize writes.
- Added a project-load gate for project-aware LSP operations before diagnostics and reference lookups.
- Retried declaration-only references with a short delay until project metadata is available, then proceeded with normal results.
2026-04-24 05:08:11 +02:00
can1357 1a3a3830b3 fix(coding-agent): validated bare-path chunk deletes in chunk mode
- Relaxed chunk-mode parameter validation to accept `{path}`-only edits as valid delete operations.
- Updated the invalid-parameters help text to document accepted chunk delete payloads.
- Added a test that verifies a bare `{path}` edit removes the targeted chunk when null values are stripped.
2026-04-24 04:35:59 +02:00
can1357 50e354c745 style: format Codex web search default model array 2026-04-24 01:20:06 +02:00
can1357 a7d78a459b fix(coding-agent): default Codex web search to gpt-5-codex-mini 2026-04-24 01:18:51 +02:00
can1357 06697b3f28 chore: bump version to 14.2.0 2026-04-24 01:05:06 +02:00
can1357 ffb5ad4880 fix(coding-agent): refresh Gemini web search credentials
Fixes #449
2026-04-24 01:02:36 +02:00
can1357 bf6b5f4989 fix(coding-agent): choose current Codex web search model
Fixes #723

Fixes #724
2026-04-24 01:02:30 +02:00
can1357 5bfdc8d46d fix(coding-agent): abort ask in headless sessions
Fixes #681
2026-04-24 01:02:24 +02:00
can1357 a9ca2daaed fix(coding-agent): fail structured subagents without submit_result
Fixes #729
2026-04-24 01:02:18 +02:00
can1357 fe076b8cd7 fix(coding-agent): cap memory rollout input
Fixes #764
2026-04-24 01:02:13 +02:00
can1357 6b8b46c410 fix(coding-agent/edit): fixed apply_patch streaming preview parsing behavior
- Added a streaming parser path for apply_patch envelopes that tolerates incomplete patch bodies.
- Updated apply patch preview expansion to return best-effort hunks when the renderer is in partial mode.
- Added a renderer test confirming streaming apply_patch input shows file paths without end-marker parse errors.
2026-04-24 00:48:03 +02:00
can1357 7e9d568346 fix(coding-agent/tools): replaced tabs with spaces in diagnostic rendering output
- Applied tab sanitization to LSP diagnostic parsing, fallback entries, and rendered diagnostics output.
- Updated tool diagnostics formatting to strip tabs from parsed file, source, message, code, and summary fields.
- Added regression tests confirming rendered diagnostics contain no tab characters and preserve message text after sanitization.
2026-04-24 00:43:55 +02:00
can1357 8a7b0a5c06 feat(ai): changed spark edit tool resolution
- Added built-in model entries for gpt-5.5 and gpt-image-2, including updated context windows, token limits, and pricing.
- Updated generated-model policy application to set or clear applyPatchToolType based on inferred GPT-5 freeform rules.
- Changed Spark edit-mode resolution to return apply_patch by default, honoring explicit replace and strict-mode overrides.
- Added tests for GPT-5 freeform policy inference and Spark edit-mode default, variant, and strict-mode behavior.
2026-04-24 00:42:19 +02:00
can1357 601c029c20 fix: stabilize apply_patch custom tool integration 2026-04-24 00:25:21 +02:00
Hans Josephsen 2a367bf043 feat(coding-agent/edit): add codex apply_patch as a new edit mode
Slots a new "apply_patch" variant alongside the existing edit modes
(replace, patch, hashline, chunk, vim). The mode accepts a single input
string containing a Codex *** Begin Patch / *** End Patch envelope,
parses it with a new lenient parser (heredoc-tolerant), and fans each
file-op out to the existing executePatchSingle so LSP writethrough,
plan-mode guards, fs-cache invalidation and diagnostics are shared
with the patch mode.

Exposes both tool shapes from the spec: the JSON function-tool variant
(§1.2, {input: string}) and the OpenAI custom-tool / Lark-grammar
"freeform" variant (§1.1, raw patch string). The edit tool advertises
a Lark grammar via customFormat and a wire name via customWireName;
openai-responses emits it as a grammar-constrained custom tool when a
model opts in with applyPatchToolType: "freeform" in models.json.
custom_tool_call / custom_tool_call_output are plumbed end-to-end
through the shared responses code (emission, streaming, history
replay), and the agent-loop dispatcher matches tool calls by either
name or customWireName so returned calls route correctly.

Also threads preview/diff rendering for apply_patch through the TUI
(tool-execution + edit renderer) so streaming patches show per-file
diffs like the other edit modes.

Default edit mode is unchanged (hashline); opt in via edit.mode or
PI_EDIT_VARIANT=apply_patch.
2026-04-24 00:15:17 +02:00
can1357 ee05fb6c2d fix: system prompt regressions 2026-04-24 00:09:09 +02:00
can1357 31294d6b0f Merged PR #763 2026-04-24 00:04:30 +02:00
can1357 2e6dac791d Merged PR #761 2026-04-24 00:04:27 +02:00
can1357 5a0c33a2db Merged PR #757 2026-04-24 00:03:59 +02:00
can1357 46f5d9d63a Merged PR #756 2026-04-24 00:03:59 +02:00
can1357 8413b406c7 docs(coding-agent): document darwin binary signing fix 2026-04-23 23:05:35 +02:00
can1357 820776d0cc test(coding-agent): align prompt wording assertions 2026-04-23 23:05:24 +02:00
Miroslav Drbal ea113ea1d2 Restore MUST weight and variadic-capture guardrail in tool prompts
Review-pass findings:

1. python.md (P3) — "Put workflow explanations in assistant message or
   cell title" lost its **MUST** weight during compression. Agents were
   observed embedding explanations as in-cell comments, polluting the
   kernel and wasting tokens. Restored as "You **MUST** put workflow
   explanations in the assistant message or cell title — never inside
   cell code."

2. ast-grep.md + ast-edit.md (P3) — The variadic-capture warning that
   $$$NAME (three dollars) is correct and $$NAME (two dollars) is
   invalid was dropped during compression. LLMs trained on shell/regex
   conventions where $$ is common are prone to this mistake; ast-grep
   emits a generic parse error rather than a "bad metavariable"
   diagnostic, so the negative example has direct reproduction value.
   Added "Use $$$NAME, **NOT** $$NAME" to both files.

All 6/6 template tests pass; bun check clean.
2026-04-23 22:59:38 +02:00
Miroslav Drbal e04d2620df Compress and tighten system prompt; relocate SECTION_SEPARATOR helper
Further compresses packages/coding-agent/src/prompts/system/system-prompt.md
from 21,530 B to 16,545 B (−23% on top of the prior compression;
−9,060 B / ~−2,265 tok per turn vs origin/main) while restoring
RFC 2119 weight on every inviolable rule and reorganizing the
pre-yield discipline.

system-prompt.md

RFC 2119 keyword restoration

All rules under `# Contract`, Procedure §§2/6, and the tail `<critical>`
block use `**MUST**` / `**MUST NOT**` keywords that the file's preamble
pins to RFC 2119. The previous draft had downgraded 29 of these to
`Do **NOT**`, which reads as below-MUST-NOT against the same preamble.
Restored sites:
- `# Contract` — all 7 inviolable bullets
- `<critical>` tail block — all 4 safety rules
- `## 6. Verification` — mock ban and no-proof yield rule
- Procedure §2 — "search for existing examples" rule regains the
  PROHIBITED parallel-convention clause it lost
- `<dir-context>` — AGENTS.md read requirement regains MUST

Structural dedup

- `<source-of-truth>` and `<instruction-priority>` were two priority
  lists with overlapping scope. Merged into a single
  `<instruction-priority>` covering all 5 conflict-resolution levels.
- `<self-check>` and `<scope-check>` were two pre-yield checklists;
  `## 6. Verification` also had a third numbered "Before yielding,
  verify..." list with overlapping checks. Collapsed all three into
  one `<pre-yield-check>` section (incl. "output format matches the
  ask" from the old verification list). `## 6. Verification` now
  focuses on verification discipline (test rigor, mocks, run-scope).

Signal restoration

Added four anchors to `<design-checklist>` that were dropped from the
earlier `<code-integrity>` block and had no home in the new structure:
- Adversarial-caller / tired-maintainer self-questioning
- Cost-of-easy-path framing
- Inhabit-the-call-site self-review
- Persistence on hard problems (do not punt half-solved work)

Handlebars fix

The `### Tool priority` header was emitted unconditionally but its
body was wrapped in `{{#ifAny python|bash}}`, producing a bare heading
with no content when neither tool was available. Header now sits
inside the conditional.

SECTION_SEPARATOR helper relocation

The `SECTION_SEPARATOR` Handlebars helper is a generic section-header
formatter that was registered in coding-agent's
`config/prompt-templates.ts`. This coupled every template consumer to
a side-effect import of that module; a prior fix added
`import "./config/prompt-templates"` to `system-prompt.ts` so the
`/system-prompt` sub-path export would register the helper, but
sibling consumers (`task/executor.ts`, `task/template.ts`) worked only
by accident because the parent agent happened to load
`system-prompt.ts` first.

- Moved `sectionSeparator` function and helper registration to
  `packages/utils/src/prompt.ts` next to the other generic helpers
  (`xml`, `codeblock`, `ifAny`, `includes`, `not`, `jsonStringify`).
- `packages/coding-agent/src/config/prompt-templates.ts` now
  re-exports `sectionSeparator` from `@oh-my-pi/pi-utils/prompt` for
  the test that imports it via the coding-agent path.
- Removed the side-effect import from
  `packages/coding-agent/src/system-prompt.ts`.

Template typo fix

Also fixes a latent `SECTION_SEPERATOR` spelling (→ `SECTION_SEPARATOR`)
in the two subagent templates (`subagent-system-prompt.md`,
`subagent-user-prompt.md`). Before the move the typo was masked
because every call site and the helper shared the wrong spelling.
Post-move, the helper is canonical `SECTION_SEPARATOR` in pi-utils
and all templates match.

Verification

- `bun check` clean (TS + Rust, all 9 packages)
- `bun test packages/coding-agent/test/system-prompt-templates.test.ts`
  6/6 pass (was 4/6 failing pre-existing before the infrastructure fix)
- `bun test packages/coding-agent/test/tools/task-template.test.ts`
  4/4 pass (exercises the `sectionSeparator` re-export)
- `bun run format-prompts` clean
2026-04-23 22:59:37 +02:00
Miroslav Drbal 416c5494fd Compress system prompt and tool descriptions for token efficiency
Reduces per-turn token cost of prompt and tool metadata by ~2,160
tokens (~8.6 KB across 12 prompt files) without losing instructional
signal. Moves literal default values from description text into
TypeBox's native `default:` keyword, and updates the strict-mode
sanitizer to preserve default info for providers that strip the
keyword.

Prompt compression

|File                                      |before |after  |delta |
|------------------------------------------|------:|------:|-----:|
|system-prompt.md                          |25,605 |21,530 |−16%  |
|prompts/tools/ast-grep.md                 | 4,892 | 3,880 |−21%  |
|prompts/tools/ast-edit.md                 | 4,141 | 3,549 |−14%  |
|prompts/tools/bash.md                     | 3,770 | 3,093 |−18%  |
|prompts/tools/task.md                     | 7,297 | 6,711 |−8%   |
|prompts/tools/read.md                     | 2,701 | 2,309 |−15%  |
|prompts/tools/debug.md                    | 2,791 | 2,412 |−14%  |
|prompts/tools/find.md                     |   860 |   561 |−35%  |
|prompts/tools/grep.md                     | 1,226 | 1,038 |−15%  |
|prompts/tools/todo-write.md               | 2,600 | 2,402 |−8%   |
|prompts/tools/python.md                   | 2,013 | 1,916 |−5%   |
|prompts/tools/hashline.md                 | 4,280 | 4,135 |−3%   |

Changes are textual compression only — grammar-scaffolding removed,
redundant bullets stripped, repeated facts consolidated. All
behavioral contracts, safety rules, worked examples, and
tool-precedence directives are preserved. The system prompt retains
RFC 2119 invocation, XML tag semantics, adversarial-caller guidance,
persistence doctrine, unit-of-change and no-forwarding-addresses
rules, completeness contract, DRY-at-2, earn-every-line, trust-
internal-code, tool precedence, AST tool priority, pattern-syntax
cheatsheet (via ast-grep.md/ast-edit.md), tool persistence,
outside-in code integrity, default follow-through, and procedural
steps 1-7. AST tool docs retain the class-wrapper +
method_definition sel example (the most error-prone usage pattern).
bash.md critical section retains MUST-weight on the ast_grep /
ast_edit directives.

Schema: defaults as first-class metadata

Moved 18 literal default values from `(default: X)` description
text into TypeBox's native `default:` keyword across:
ast-edit.ts, ast-grep.ts, bash.ts, browser.ts, find.ts, gh.ts,
grep.ts, python.ts, read.ts, ssh.ts. Runtime-resolved placeholders
(cwd, pr-<number>) remain as text since they cannot be literal
TypeBox defaults.

Added:
- `minItems: 1` on ast-edit.ts `ops` array — machine-enforces what
  the .md previously stated only in prose; handler already rejects
  empty ops arrays, this adds the schema-level constraint.
- `find.ts` `pattern` description enriched with facts previously
  only in find.md (comma-separated lists, simple patterns recurse
  from cwd).

Strict-mode sanitizer: inline `default` into `description`

OpenAI's Structured Outputs strict mode rejects schemas containing
`default` with HTTP 422 ("default is not permitted"). Affects
openai, azure, github-copilot, openrouter, cerebras, together,
zenmux, and deepseek providers.

`sanitizeSchemaForStrictMode` in packages/ai/src/utils/schema/
strict-mode.ts now appends ` (default: X)` to the sibling
`description` before stripping the `default` keyword. Non-strict
providers (Anthropic, Google) still see the native keyword.

Rules:
- Inline is skipped when description already contains `(default:`
  (prevents double-inlining on recursive calls)
- Inline is skipped when no sibling description exists (no
  synthesis)
- Formatting: strings as-is (`cwd`), other values via
  `JSON.stringify` (matches the conventional text form)

CONSTRAINTS.md documents the inlining rule alongside the existing
keyword-strip rule.

Regression tests

packages/ai/test/schema-strict-mode.test.ts gains 7 `it` blocks:
- number/bool/string default types inline correctly
- falsy defaults (`false`, `""`, `0`) are not confused with absent
- `null` default goes through JSON.stringify branch
- double-inline prevention when description already says `(default:`
- no synthesis when no description exists
- nested object property with default (recursion + cache path)
- type-array `[T, null]` branch with default on outer schema
  (variant-materialization path)

23/23 schema tests pass, 541/541 ai-package tests pass,
`bun check` clean.

Rationale

|Metric                          |Value          |
|--------------------------------|--------------:|
|Per-turn prompt savings         |~2,160 tok     |
|Files touched                   |25             |
|Lines changed                   |+325 / −296    |
|Schemas migrated to `default:`  |18             |
|New regression tests            |7              |
2026-04-23 22:59:37 +02:00
Kagura f9522de33c fix(auth): let OAuth credentials override keyless provider flag from stale models.yml
When a user models.yml declares a provider with auth: none (e.g. from
an older version), getApiKey() returns the literal 'N/A' without ever
consulting authStorage — even when valid OAuth credentials exist from
a successful /login.

Check authStorage.hasAuth() before returning kNoAuth. If real
credentials are present, fall through to the normal auth flow.

Fixes #749
2026-04-23 22:59:33 +02:00
fettpl 14b08f89d4 fix: restore runnable darwin compiled binaries
fixes #754
2026-04-23 22:59:27 +02:00
can1357 4490332822 test(coding-agent): avoid inline imports in claude plugin discovery tests 2026-04-23 22:54:59 +02:00
can1357 5072d3df11 docs: add compiled binary autoload changelog 2026-04-23 22:54:40 +02:00
fcremo (Filippo Cremonese) 17df89086b fix: disable bunfig.toml and .env autoloading in compiled binary
Compiled Bun binaries unconditionally load bunfig.toml and .env from
the current working directory at runtime, before any application code
runs. Since omp is a coding agent that runs from arbitrary project
directories, it picks up foreign project configs -- most critically
preload directives that cause immediate crashes, but also a potential
security issue since preloads execute arbitrary code.

Add --no-compile-autoload-bunfig and --no-compile-autoload-dotenv to
the bun build --compile invocations (available since Bun v1.3.3).

Note: source installs via 'bun install -g' are still affected because
'bun run' has no equivalent flag. That remains an open issue.
2026-04-23 22:51:25 +02:00
Parsifa1 782a309f6f fix(coding-agent): fix untrusted path resolve 2026-04-23 22:51:01 +02:00
Parsifa1 e56a33c857 fix(coding-agent): honor claude plugin manifest paths 2026-04-23 22:51:01 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 359318f931 fix(coding-agent/edit): updated edit rendering to read path and metadata from first edit entry
- Updated path resolution to derive a file path from the first edit entry via filePathFromEditEntry when top-level fields are absent.
- Updated single-file result rendering to fall back to first-edit entry values for op and rename metadata before detail-level metadata.
2026-04-19 08:47:49 +02:00
can1357 7453c10801 feat(coding-agent): added optional inline read preview toggle
- Added a new read.toolResultPreview boolean setting defaulting to false to control inline read result rendering.
- Updated read tool group rendering to show inline previews only when enabled and to hide duplicate summary rows for previewed entries.
- Passed the setting through event and UI helper constructors and added tests for default-off behavior and the duplicate-preview summary case.
2026-04-18 23:37:47 +02:00
can1357 23c66791bc fix: stabilize merged review follow-ups 2026-04-18 23:24:00 +02:00
can1357 c3524ea98e merge: PR #719 2026-04-18 23:22:54 +02:00
can1357 3581569ef1 merge: PR #714 2026-04-18 23:22:52 +02:00
can1357 214600265e merge: PR #706
# Conflicts:
#	packages/coding-agent/CHANGELOG.md
2026-04-18 23:22:51 +02:00
djdembeck e15a4ba283 fix: resolve local: URI scheme on Windows
The previous check only validated 'local:' prefix, which caused
'local:PLAN.md' to incorrectly resolve to 'LAN.md' (the colon was
interpreted as a Windows drive letter).

Now requires 'local:/' or 'local://' prefix to ensure proper URI parsing.
2026-04-18 22:41:09 +02:00
djdembeck 5e45bee73f refactor: improve path handling with normalization refactor
- Refactor path normalization to combine expandPath and normalizeLocalScheme
- Add validation in utils.ts to reject local:// paths as filesystem paths
- Fix bash-skill-urls regex to handle hyphen-prefixed local:/ patterns
- Add tests for hyphen-prefixed and @local: patterns
2026-04-18 22:41:09 +02:00
djdembeck c7c3d4a82d fix: avoid matching local:/ in filesystem paths
- Add negative lookbehind to regex in bash-skill-urls to prevent matching local:/
  inside paths like /repo/local:/PLAN.md
- Normalize local scheme before expanding paths in path-utils
- Add test cases for both changes
2026-04-18 22:41:09 +02:00
djdembeck 91998d326d fix: handle local:/ single-slash URL pattern
Expands the regex pattern to match local:/ (single-slash) URLs in addition to local:// (triple-slash), preventing potential Linux path leaks.

- Add regex patterns for single-quoted, double-quoted, and unquoted local:/ URLs
- Add test coverage for all three quote styles
2026-04-18 22:41:08 +02:00
djdembeck 90ba65ac85 docs: add changelog entry for local:// URL path leak fix 2026-04-18 22:41:08 +02:00
djdembeck 3dc7fce583 refactor: remove redundant isInternalUrlPath guards from edit preview paths 2026-04-18 22:40:07 +02:00
djdembeck f353873b75 refactor: extract local:// URL normalization to shared utility
Extract duplicate normalizeLocalScheme regex pattern into a shared function in path-utils.ts. Updated interactive-mode.ts, approved-plan.ts, agent-session.ts, bash-skill-urls.ts, and plan-mode-guard.ts to use the shared utility. Also fixed error message formatting (removed extra backslashes).
2026-04-18 22:40:07 +02:00
djdembeck 5da806671e fix: prevent local:// URI from creating local: directory on Linux
On Linux, Node's path.normalize() collapses the double slash in
local://PLAN.md to local:/PLAN.md, creating a directory called local:
in the project root instead of routing through the local:// protocol handler.

Defense-in-depth fixes across 5 layers:

1. resolveToCwd() now throws if a path starts with any internal URL
   scheme prefix (local:, agent:, skill:, etc.), preventing all 59
   call sites from treating URIs as relative filesystem paths.

2. resolvePlanPath() now matches on local: prefix (not just local://)
   and normalizes local:/ to local:// before resolution, catching
   all slash variants.

3. Bash URL expansion regex and early-exit checks now also match
   local:/ (single slash), and normalize before resolution.

4. Edit preview/diff functions now gracefully skip internal URL paths
   instead of crashing via the resolveToCwd guard.

5. All startsWith('local://') checks updated to startsWith('local:')
   with normalization in agent-session, interactive-mode, and
   approved-plan modules.

Also adds local: to .gitignore to prevent accidental commits of the
leaked directory.
2026-04-18 22:40:07 +02:00
Sam Biggins 22f429235e fix(web-search): decoupled Tavily topic from recency filter
Removed unconditional topic:news coupling in buildRequestBody that scoped
Tavily index to news publications whenever recency was set. Technical queries
with --recency now search the general index filtered by time only.

Tightened SearchParams.recency contract in base.ts: providers MUST interpret
recency as a pure time filter and MUST NOT change topic scope as a side effect.
2026-04-18 22:36:33 +02:00