Commit Graph
7102 Commits
Author SHA1 Message Date
roboomp fbb7255e8e fix(eval): flipped isolated default to strict opt-in
Per maintainer ruling on #3196, eval agent() now defaults to non-isolated regardless of task.isolation.mode, mirroring the task tool. isolated=true is the only way to turn it on; isolated=true while task.isolation.mode === "none" still throws the same clear error.

Updated tests, workflow-notice.md, and Python agent() docstring to reflect the strict opt-in contract. Existing isolation tests now pass isolated:true explicitly; the inherit-from-settings assertion is replaced with a default-off + isolated=true opt-in regression.

Fixes #3196
2026-06-22 18:29:28 +00:00
roboomp a48f1ea813 fix(eval): persisted nested patches before throwing apply failures
When an isolated apply fails the bridge throws a ToolError and never returns details, so the nested-patch payload that previously lived in details.nestedPatches was unrecoverable after the isolation worktree was torn down.

The bridge now writes each captured nested patch to a file under the per-call artifacts dir (e.g. <agentId>.nested-<index>-<slug>.patch) before throwing and includes the resolved paths in the error message so the caller can apply them manually.

Added a regression test verifying the persisted file exists with the original patch contents and that the path is surfaced in the thrown error.

Fixes #3196
2026-06-22 04:21:11 +00:00
roboomp c1b01f1cb9 fix(eval): surfaced isolated apply failures as ToolError
A failed isolated apply (changesApplied === false) previously only set details.isolationSummary and returned the subagent text. Schema-backed agent() calls then parsed the JSON and returned the object, so workflows saw a successful structured result while none of the edits had landed.

Throw a ToolError when mergeIsolatedChanges reports a failed apply, with the merge summary plus a recovery hint pointing at the preserved patch/branch/nested artifacts so the caller can apply manually.

Added regression tests for the schema and non-schema apply-failure paths.

Fixes #3196
2026-06-22 04:13:43 +00:00
roboomp 3677e71e9d fix(eval): exposed nested apply=false patches
Branch-mode isolation can capture nested repository changes without creating a root branch. Eval agent() with apply=false previously treated that shape as no captured changes and returned no recoverable nested patch payload after the isolation worktree was removed.

Expose captured nested patches in EvalAgentResult details and copy them onto JS/Python returnHandle nodes (nestedPatches / nested_patches). Document the return_handle escape hatch and add regression coverage for branch-mode nested-only apply=false runs.

Fixes #3196
2026-06-22 04:08:56 +00:00
roboomp 636fbb4fcc style: bun run fix 2026-06-22 03:59:27 +00:00
roboomp a6ee95df08 fix(eval): preserved returnHandle artifacts and nested branch patches
Eval preludes now forward returnHandle to the bridge so no-session eval runs can preserve the temp artifacts backing returned agent:// handles. The bridge keeps those temporary artifact directories whenever returnHandle is requested, including non-isolated runs and successful isolated applies.

Branch-mode isolation now treats nested-only changes as merge-eligible even when no root branch was produced, letting callers apply nested patches instead of dropping them when the root repo had no diff.

Added regression coverage for returnHandle artifact preservation and nested-only branch isolation.

Fixes #3196
2026-06-22 03:59:17 +00:00
roboomp 54f4652573 fix(eval): exposed apply=false artifacts on the agent() return_handle node
When agent() ran with schema and apply=false, the bridge correctly returned the captured patch/branch in details, but the preludes only forwarded id/agent/handle/data on the returnHandle node. Structured workflows had no way to recover the artifact for a manual apply.

Both runtimes now copy isolated, patchPath/branchName, changesApplied, and isolationSummary onto the returnHandle node (snake_case in Python, camelCase in JS), keeping null changesApplied so apply=false stays distinguishable from a successful apply. Updated the workflow notice and the Python agent() docstring to point callers at return_handle as the artifact escape hatch for isolated+apply=false runs. Added prelude tests locking the new node shape in both runtimes.

Fixes #3196
2026-06-21 17:35:26 +00:00
roboomp 9b33e0e644 style: bun run fix 2026-06-21 17:28:23 +00:00
roboomp 068a27b07f fix(eval): preserved temp artifacts when isolated agent() runs with apply=false
The cleanup gate was treating changesApplied===null (apply=false) the same as a clean apply, deleting the temp artifacts dir before returning details.patchPath. Sessions without a session file (which fall back to a per-call tmp dir) ended up with a patchPath pointing at a removed file, defeating the documented manual-apply path.

Tightened the cleanup condition to remove the temp dir only on a confirmed clean apply (changesApplied===true); apply=false and failed applies both keep the artifact for the caller.

Added regression tests for the apply=false preserve case and the apply-succeeds cleanup case.

Fixes #3196
2026-06-21 17:28:16 +00:00
roboomp 91cf5ec97c fix(eval): kept isolated schema output parseable
Kept isolation merge/apply summaries out of agent() text when a schema is supplied so Python and JS eval helpers can still parse the JSON payload. The summary now lands in details.isolationSummary for callers that need the human-readable apply state.

Added a regression test that exercises an isolated schema-backed eval agent with a merge summary.

Fixes #3196
2026-06-21 17:25:49 +00:00
roboomp 6a22499dcd style: bun run fix 2026-06-21 17:20:14 +00:00
roboomp eeed4b93b6 feat(eval): added isolated/apply/merge options to agent() helper
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.

Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.

Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.

Fixes #3196
2026-06-21 17:20:07 +00:00
can1357 35508c48f7 chore: bump version to 16.1.10 2026-06-21 08:40:20 +02:00
can1357 076fd0de81 test(coding-agent): realign discovery and TTSR fixtures with current tool invariants
- sdk-mcp-discovery: `find` became an essential tool (2eef88978), so it can no longer be hidden/rediscovered under `tools.discoveryMode: all`. Switch the discoverable-tool assertions to `search` (still `loadMode: discoverable`).
- agent-session-concurrent: the agent loop now drops tool calls that never reached `toolcall_end` from an aborted turn (0890b2be6, partial args are unsafe to replay). Emit `toolcall_end` before the TTSR rule-driven abort so the labeled placeholder result is minted.
2026-06-21 08:37:37 +02:00
can1357 5207a98fd8 refactor(coding-agent): simplified temporary model status message
- Inlined the temporary model status formatting logic directly into the controller.
- Removed the unused `formatTemporaryModelStatus` utility function and its associated test.
2026-06-21 07:42:46 +02:00
can1357 c5cf47aa7f fix(coding-agent/utils): handled EISDIR and ENOTDIR errors during ref resolution
- Update `shouldRetry` to treat `EISDIR` and `ENOTDIR` as terminal errors, preventing unnecessary retries when encountering Git reference directory conflicts.
- Add a test suite to verify graceful resolution of branches in scenarios where a packed ref conflicts with a directory path in the filesystem.
2026-06-21 07:36:25 +02:00
can1357 07c910210b feat(coding-agent): allowed toggling workspace tree in system prompt
- Added `includeWorkspaceTree` configuration setting to optionally render the workspace directory tree.
- Configured the system prompt template and SDK to respect this toggle, allowing users to disable the tree to prevent prompt cache invalidation.
2026-06-21 07:03:45 +02:00
can1357 2eef88978b feat(coding-agent): made write and find tools essential
- Promoted `write` and `find` tools to `essential` status to ensure they are always available regardless of discovery mode.
- Updated `DEFAULT_ESSENTIAL_TOOL_NAMES` to include these tools by default.
- Updated documentation and tests to reflect the change in default essential tool availability.

Fixes #3165
2026-06-21 07:02:44 +02:00
can1357 151561cb7c fix(terminal-bench): improved docker cleanup error reporting
- Added error handling to docker container and network removal commands.
- Print yellow warning messages to stdout when cleanup commands fail to exit cleanly.
2026-06-21 06:52:21 +02:00
can1357 984c8dd2f6 fix(coding-agent): reconciled tool arguments on execution start
- Synchronize tool arguments with the component state upon receipt of `tool_execution_start` to ensure visual consistency when final update events are missed.
- Terminate active argument reveal streams to prevent late ticks from overwriting valid, fully-materialized tool arguments with stale partial data.
- Add test coverage to verify that tool UI components render finalized arguments even in the absence of intermediate streaming updates.
2026-06-21 06:51:37 +02:00
can1357 e649017322 feat(coding-agent): added ARIA snapshot support for browser tools
- Implemented `tab.ariaSnapshot()` to capture and represent page structures as ARIA-tree YAML.
- Introduced `tab.ref()` and ref-based selector parsing to enable precise element interaction via unique ARIA identifiers.
- Integrated automated script bundling for cross-environment evaluation of ARIA snapshot logic.
- Updated browser action methods to resolve and target elements using ARIA-ref handles.
2026-06-21 06:39:14 +02:00
can1357 d7d15b94a9 test(coding-agent): extended test suite to mock process rows
- Mocked stdout rows for the AgentTranscriptViewer tests to ensure consistent terminal height.
- Added an afterEach cleanup to restore the original property descriptor for process.stdout.rows.
2026-06-21 04:19:53 +02:00
can1357 5ba474656c chore: bump version to 16.1.9 2026-06-21 04:16:49 +02:00
can1357 45410e865e chore: fix changelogs 2026-06-21 04:16:35 +02:00
can1357 57cb7447e0 feat(coding-agent): enhanced browser stealth and automation capabilities
- Overhauled stealth spoofing mechanisms for WebGL, screen dimensions, Web workers, and iframe contexts using prototype-aware injection.
- Centralized function string representation patching to improve mimicry of native browser behavior across global objects.
- Patched puppeteer-core to remove detectable evaluation markers and implement lazy, pull-style execution context management.
- Enabled support for capturing LLM request JSON dumps and adjusted launcher flags to improve organic request patterns.
2026-06-21 04:14:26 +02:00
can1357 4d96bcf6b6 feat: replaced manual debug request path with automated llm dump
- Replaced the `/debug dump-next-request` command with an updated `/dump` command that exports LLM request context to JSON sidecar files.
- Removed persistent debug path state and manual path configuration in favor of automated generation.
- Updated session logic to handle serializing LLM request context to temporary directories.
- Refactored testing suites to remove path-based debug tests and verify dynamic request file generation.
2026-06-21 03:18:20 +02:00
can1357 320f514f77 fix(ai): adjusted llama-cpp base url and refactor token validation
- Update `llama.cpp` base URL to remove the `/v1` suffix.
- Consolidate local provider token validation in `coding-agent` using a `Set`.
- Add test coverage to ensure catalog model IDs are preserved verbatim on the wire.
2026-06-21 02:45:48 +02:00
oldschoolaandcan1357 13bfc7b9c0 Remove Wafer Pass provider 2026-06-21 02:45:48 +02:00
can1357 7fcb85f812 chore: updated changelogs & added moonshot tests 2026-06-21 02:29:57 +02:00
can1357 e31bfa4b96 Merge remote-tracking branch 'origin/farm/7a22e7a4/mcp-toggle-reconnects-others' 2026-06-21 02:18:29 +02:00
can1357 0f976084b8 fix(coding-agent/cli): route leading global flags before a subcommand to that subcommand
`omp --approval-mode=yolo acp` was rewritten to `launch --approval-mode=yolo
acp`, swallowing `acp` as a launch prompt so the yolo override never reached the
ACP command path (the ACP permission gate from #2097 stayed in always-ask).

`resolveCliArgv` only inspected `argv[0]`, so any leading global option flag hid
the real subcommand. It now scans past leading flags using the launch parser's
value-consumption contract (a flag's value is never mistaken for the subcommand,
e.g. `--model acp`) and hoists the recognized subcommand to the front with the
flags preserved as its own argv. Genuine launch prompts are untouched.

The flag value-consumption rule is factored into `cli/flag-tables.ts`
(`flagConsumesValue` + the shared `isUnknownLongValueCandidate`) so the resolver
and the profile bootstrap share one source of truth.

Fixed permission mode not respected in ACP mode ([#2970](https://github.com/can1357/oh-my-pi/issues/2970))

Fixes #2970
2026-06-21 02:17:54 +02:00
can1357 725aa2222b fix(coding-agent/lsp): drain config pulls on lazy cold start
The background message reader matched every incoming message against the
pending client-request map by id before checking for a `method`. Server
request ids live in the server's own id space and routinely collide with
the client's in-flight request ids, so a server-originated
`workspace/configuration` pull whose id matched a pending request (e.g. a
basedpyright pull landing while a `documentSymbol` request with the same
id was open) was swallowed as a bogus response: the client request
resolved with `undefined` and the pull was never answered, wedging
servers that gate analysis on configuration.

Route any message carrying a `method` as a server request (or
notification) before id-matching responses, so every config pull is
answered under lazy init -- parity with the warmup/reload path, which
escaped the bug only because it issues no concurrent semantic request
while the cold-start pulls drain. lsp.lazy default is unchanged.

Fixes #3001
2026-06-21 02:16:26 +02:00
can1357 b0c6ad825c fix(coding-agent): hint at omp plugin list/uninstall instead of leaking bare plugin verbs to launch
The plugins docs advertise `omp list`/`omp remove` as top-level commands,
but only `omp install` is registered. `resolveCliArgv` rewrote any
unregistered first-arg to `launch`, so `omp list` silently started an
interactive agent session with "list" as the LLM prompt instead of
managing plugins.

Reserve the bare, documented-but-unregistered plugin verbs `list` and
`remove` with a helpful hint pointing at the real `omp plugin list` /
`omp plugin uninstall <name>` commands (same `extensions` mechanism /
same class as the `install` leak fixed in #1496/#1498). Multi-word
prompts that merely begin with those words still route to `launch`, so
genuine prompts are unaffected.

Fixes #2935
2026-06-21 02:16:17 +02:00
roboomp c18d400bdf fix(cli): scoped mcp toggles to one server
Updated /mcp enable and /mcp disable so they connect or disconnect only the named server instead of reloading every MCP server in the session. Added regression coverage for both toggle directions and updated the coding-agent changelog.

Fixes #3157
2026-06-20 23:58:03 +00:00
can1357 12cfb530ad chore: bump version to 16.1.8 2026-06-21 00:39:58 +02:00
can1357 e750fb257c test(coding-agent): updated compaction strategy in test configuration
- Updated the test configuration to specify an explicit compaction strategy.
- Enabled context-full compaction to match expected behavior during tests.
2026-06-21 00:35:07 +02:00
can1357 29f7b57fde feat: migrated core rendering and context transformation to async
- Updated `render`, `renderMany`, and native snapcompact methods to return promises, ensuring scalable async execution.
- Refactored `transformProviderContext` and `buildSideRequestContext` to support asynchronous operations in agent loops.
- Integrated `Promise.all` for improved concurrency when processing frame rendering and rendering batch operations.
- Updated all internal call sites, SDK hooks, and test suites to accommodate the asynchronous API signatures.
2026-06-21 00:32:13 +02:00
can1357 51f5919486 feat(coding-agent): improved advisor communication constraints
- Updated the system prompt to explicitly instruct the advisor never to send the same advice twice.
2026-06-21 00:14:06 +02:00
can1357 78a94c5a93 feat(snapcompact): improved archive text preservation and continuity
- Unify `previousText` resolution to correctly concatenate `textHead` and `textTail` during re-compaction.
- Ensure summary fallback logic correctly handles non-text legacy archives.
- Add test coverage for cross-compaction text retention and legacy archive continuity.
- Enable snapcompact strategy in agent plan reference re-injection tests.
2026-06-21 00:11:02 +02:00
can1357 829d6c66dc Merge PR #2903: fix(mnemopi): make proactive linking configurable from host settings (@wolfiesch) 2026-06-21 00:09:20 +02:00
can1357 aa9709c47f feat(coding-agent): added snapcompact safety checks and UI integration
- Added validation to scan for non-ASCII characters before performing snap-compaction, falling back to LLM-based summarization if the unrenderable ratio is too high.
- Updated event handling and status reporting to explicitly support snapcompact actions, including specific error warnings and cancellation states in the UI.
- Updated session logic to default to snapcompact strategy when auto-compaction is enabled.
2026-06-21 00:08:12 +02:00
can1357 3c356fc0e0 fix(coding-agent/session): improved auto-compaction strategy selection
- Adjusted the fallback strategy when enabling auto-compaction to respect the system default instead of hardcoding a specific value.
2026-06-21 00:06:36 +02:00
can1357 850d55e2ea config(coding-agent/config): updated default agent settings
- Increase default immune turns for advisor concerns from 1 to 3.
- Update default compaction strategy from context-full to snapcompact.
2026-06-21 00:05:43 +02:00
can1357 aa87848f1a fix(coding-agent): prevented accidental cancellation of background maintenance
- Stop the Esc key from aborting active background maintenance (compaction, handoff, or retry) while a subagent is focused.
- Remove "(esc to cancel)" hints from maintenance loaders when a subagent is active to avoid false affordance.
- Ensure that main-session maintenance remains cancellable via Esc when no subagent is focused.

Fixes #2819
2026-06-21 00:00:35 +02:00
can1357 9f34c45d94 fix(coding-agent/modes): restricted branch key input to focused editor
- Added a check to ensure the editor is focused before allowing the branch keyboard shortcut.
2026-06-20 23:57:29 +02:00
can1357 aaac4e414f Merge PR #2950: feat(coding-agent): add copy affordance to completed /btw answers (@wolfiesch)
Adds 'c copy' to the completed /btw panel footer (alongside b branch / Esc
dismiss), copying the sanitized visible answer to the clipboard. The copy
shortcut is guarded by canCopyBtw + main-editor focus + empty editor.
2026-06-20 23:56:24 +02:00
can1357 330a37e289 test(coding-agent): fixed type error in event controller tests
- Added missing settings property to the InteractiveModeContext mock in event controller tests.
2026-06-20 23:43:38 +02:00
can1357 9fc7ab9fe0 chore: update changelogs 2026-06-20 23:43:16 +02:00
roboompandcan1357 75b21611f7 fix(mcp): sanitized startup server names
Sanitized MCP server names before status formatting so configured keys cannot leak home paths, tabs, newlines, or oversized text into the TUI.

Fixes #3150
2026-06-20 23:43:16 +02:00
can1357 63655a5b71 Merge PR #3050: feat: show live slash command status in autocomplete (@wolfiesch)
Shows live per-command status (plan/goal/loop/model/advisor/collab/jobs/
context/...) in the slash-command autocomplete.

Adjustments on merge:
- Exclude the two .github/pr-assets/*.png screenshots.
- Resolve the autocomplete.ts conflict keeping main's skill-command empty-prefix
  boost alongside the new static-vs-display description split.
- Compute the live display description lazily — only once a command actually
  matches (name or alias) — instead of for every command on each keystroke, and
  guard the getter with a typeof check. getAutocompleteDescription reads live
  session state, so the eager call was O(commands) work per refresh.
2026-06-20 23:36:17 +02:00