Commit Graph

89 Commits

Author SHA1 Message Date
can1357 71af89cbaa refactor(coding-agent): extracted generic kernel session registry
- Python, Ruby and Julia each carried their own copy of the same session
  maps, acquire/reset/replace/dispose lifecycle and executeOnSession; one
  generic registry now owns it, parameterized by a per-language descriptor.
- Julia additionally re-implemented seven executor-base helpers locally; those
  copies are gone and executor-base gained sparse managed-env and timeout
  resolver hooks so Julia's differing behavior survives unchanged.
- All twelve exported entry points keep their names and signatures.
2026-08-08 06:32:00 +02:00
metaphorics 85177a29ec fix(coding-agent): fail loudly when no POSIX bash exists; add stdin-isolation regression test 2026-08-05 10:42:36 +00:00
metaphorics f03581dc22 fix(coding-agent): stop eval shell children inheriting the runner control channel; resolve Git Bash on Windows 2026-08-05 10:24:24 +00:00
Slava Zavadsky b8779dae63 fix(coding-agent/eval): remove per-call model override from agent() (#6438)
Completes the maintainer's removal of per-call model selection from
subagent spawns (9f8aa87dbf removed it from the task tool and the
model-facing agent() docs/prompt, but the eval agent() runtime and all
four preludes still accepted and forwarded a per-call model).

Subagents now always resolve through the selected agent's frontmatter
model and settings, so an explicit model: "default" can no longer
silently route children onto the parent session model.

- agent-bridge: drops "model?" from agentArgsSchema and the request
  forward; adds "+": "delete" so a legacy model argument is stripped
  (same contract as the task wire schemas).
- JS/Python/Ruby/Julia preludes: remove the model parameter from
  agent(); completion()'s tier selector is unchanged.
- docs (tools/eval.md, python-repl.md) updated to the removed surface.

Refs #6438
2026-08-04 08:35:01 -04:00
can1357 b48aeacf25 fix(coding-agent/eval): made cells own every bridge call they start
- Split the Python bridge signal: the raw abort reaches tools (so
  subagents die with the turn) while a shielded signal governs how long
  the host waits, so a cancel can no longer settle a cell on top of a
  still-running isolation merge.
- Mirrored the contract in the JS runtime: abort in-flight tool calls
  immediately, then drain any deferExternalAbort phase before killing
  the worker, and refuse new bridge calls once cancelled.
- Held a finished worker result until the run's tool calls drain. A
  floated agent() previously settled the cell at once, dropping the
  run's abort listener and leaving the subagent running with nothing
  able to cancel it.
- Added regression coverage for all three, each verified to fail
  without its fix.
2026-08-03 19:55:20 +02:00
can1357 c7a3010d7c fix(coding-agent/eval): ensured tool bridge receives unshielded abort signal
- Prevent eval agent bridge calls from inheriting the kernel abort shield.
- Unset maxRuntimeMs in runEvalAgent to properly inherit task.maxRuntimeMs.
2026-08-03 18:57:22 +02:00
can1357 46c66ac03a fix(setup): always resolve interpreter during setup probe 2026-08-02 20:53:34 +02:00
roboomp 3fab985c14 fix(eval): preserved console for conpty python kernels
Combined native HWND and stdio TTY evidence so compiled Windows Terminal launches no longer set CREATE_NO_WINDOW for the Python eval kernel.

Fixes #7343
2026-08-02 05:05:52 +00:00
can1357 ccc9b900cf fix(eval): forked subagent kernel resets away from shared sessions
- Subagents inherit the parent's eval session id, so a child's
  reset: true destroyed the co-owned kernel and every sibling's
  interpreter state mid-session.
- resolveOwnerScopedSessionKey now routes a reset from a non-exclusive
  owner onto a deterministic per-owner fork key: the requester gets a
  fresh private kernel, co-owners keep the shared one, and the fork
  stays sticky for that owner until its teardown reaps it.
- Applied across Python, JavaScript, Ruby, and Julia executors; JS
  contexts gained an owner registry plus disposeVmContextsByOwner,
  wired into EvalRunner.disposeKernels and SDK session teardown.
- Covered by pure key-resolution contracts and an end-to-end JS test:
  co-owner reset forks, shared state survives, fork is sticky, and
  per-owner dispose reaps only the fork.
2026-07-31 20:13:23 +02:00
can1357 14dc5d4fcc fix(coding-agent/eval): bypassed environment proxies for python bridge calls
- Configure the python eval prelude to use a custom urllib opener that ignores environment proxies.
- Add test coverage verifying parallel tool bridge calls succeed when proxy variables are set.
2026-07-30 16:19:48 +02:00
can1357 d16a251777 chore: reorg tests 2026-07-27 16:43:53 +02:00
can1357 4eb94125b2 fix(coding-agent/eval): filtered internal runner frames from python cell error tracebacks
- Filter out runner-internal frames from runtime exception tracebacks to start at user code.
- Omit full tracebacks for cell syntax errors to render only the caret display with `<cell>` filename.
2026-07-27 07:28:22 +02:00
roboomp 03489d1ebe fix(eval): coalesced python kernel replacement
Tracked retained Python kernel generations and shared one replacement promise per dead generation. Reset and disposal now invalidate and drain replacement work before allowing a new session to take ownership.

Added deterministic fake-kernel coverage for concurrent callers, cancellation, reset, owner/global disposal, and independent cwd keys.

Fixes #6367
2026-07-23 19:00:08 +00:00
can1357 b6e68fd243 fix(coding-agent): updated mentions of explore -> scout 2026-07-23 00:12:07 +02:00
vmcall d944879f21 feat(task): unified structured subagent execution
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.

Fixes #5279
2026-07-17 17:36:59 +02:00
can1357 55f5ebec49 chore: reformat 2026-07-15 00:08:50 +02:00
can1357 a9c038818d feat(tools): removed separate selector args from read and grep APIs
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
2026-07-15 00:04:02 +02:00
can1357 37b1263a26 Merge PR #5494: fix(eval): delegate Python URI reads to host resolver (@roboomp) 2026-07-14 23:11:12 +02:00
roboomp 2bfdf0ab07 fix(eval): passed uri selectors separately
- Kept opaque MCP resource paths unchanged during pagination.
- Sent line ranges through the read tool selector field.
- Covered paged artifact and MCP reads in the Python prelude test.

Fixes #5353
2026-07-14 20:20:42 +00:00
roboomp 5d6fcff2f1 fix(eval): delegated python uri reads to host resolver
- Routed non-local URI reads through the session read tool.
- Preserved offset and limit as host line selectors.
- Covered artifact delegation with the shipped Python prelude.

Fixes #5353
2026-07-14 19:12:06 +00:00
can1357 5d28a319fd Merge PR #5015: fix(eval): shield Python agent bridge aborts (@roboomp) 2026-07-14 18:45:03 +02:00
Colin Mason c40ccdc684 isolate eval runtimes from terminal 2026-07-13 09:57:19 -04:00
roboomp a86c1ec46d fix(eval): shield agent bridge aborts
- Deferred eval cancellation while Python bridge calls are paused so already-started agent() subagents can finish and persist output.
- Rejected new Python bridge calls after an external abort is pending to prevent post-abort fan-out waves.
- Added regression coverage for the bridge shield and Python parallel agent() interruption path.

Fixes #5005
2026-07-10 01:21:39 +00:00
can1357 db8c79cc93 style: apply formatter and fix devin test enum
biome + cargo fmt over merge-sweep eval-fix commits; correct StopReason.END_TURN (nonexistent) to StopReason.FUNCTION_CALL in devin streaming test
2026-07-01 04:45:28 +02:00
can1357 5888bccba0 fix(eval): truncate oversized Python shell output 2026-07-01 04:40:55 +02:00
roboomp e348e8d9b5 fix(eval): bound python shell helper output
Stream Python eval shell helper stdout in fixed-size chunks instead of buffering through subprocess.run or newline-bound text iteration.

Added regression coverage for !cmd and newline-free %%bash streaming.

Fixes #3950
2026-07-01 01:15:38 +00:00
can1357 1dd78b207e feat(coding-agent): removed unused eval helper functions
- Removed deprecated eval prelude helpers `append`, `tree`, `diff`, `sort`, `uniq`, and `counter` from all supported runtimes.
- Cleaned up runtime implementations, protocol definitions, and UI rendering logic associated with the removed helpers.
- Updated project documentation, prompts, and test suites to reflect the reduced helper API surface.
- Recorded functional changes in the package changelog.
2026-06-23 01:39:24 +02:00
can1357 9e6eb98db5 refactor(coding-agent/eval): removed unused sh cell magic
- Removed the `_magic_cell_sh` function as `sh` cell magic was no longer required.
2026-06-23 00:27:45 +02:00
can1357 24d58033bb refactor(eval): rename agent() params agent_type→agent, return_handle→handle
The eval agent() helper used `agent_type`/`return_handle` (snake_case) in
Python/Ruby/Julia and `agentType`/`returnHandle` (camelCase) in JS, forcing
the prelude docs to repeat every option twice ("JS same but camelcased").
Both are now single lowercase words identical across all four runtimes, and
`agent` matches the `task` tool's existing agent-selection parameter.

- Renamed across py/js/rb/jl preludes (signatures, forwarding, docstrings).
- Renamed the `__agent__` bridge wire protocol + `EvalAgentArgs` (`agentType`
  → `agent`, `returnHandle` → `handle`) so no prelude-side remap is needed.
- Updated prompt docs (workflow-notice.md, tools/eval.md), repo docs
  (docs/tools/eval.md, docs/python-repl.md), and all bridge/prelude tests.
- CHANGELOG: Breaking Changes entry under [Unreleased].
2026-06-22 22:12:24 +02:00
can1357 2f2acaf928 Merge pull request #3205: feat(eval): isolated/apply/merge options for agent() helper
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.

Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
  forwards isolated/apply/merge (as booleans) plus returnHandle, and the
  return_handle node carries isolated/patch_path/branch_name/
  nested_patches/changes_applied/isolation_summary.

Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
  task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
  the stale "defaults track task.isolation.mode" wording to the final
  strict opt-in behavior, and note all four runtimes.

Fixes #3196
2026-06-22 21:56:45 +02:00
roboomp fbb7255e8e fix(eval): flipped isolated default to strict opt-in
Per maintainer ruling on #3196, eval agent() now defaults to non-isolated regardless of task.isolation.mode, mirroring the task tool. isolated=true is the only way to turn it on; isolated=true while task.isolation.mode === "none" still throws the same clear error.

Updated tests, workflow-notice.md, and Python agent() docstring to reflect the strict opt-in contract. Existing isolation tests now pass isolated:true explicitly; the inherit-from-settings assertion is replaced with a default-off + isolated=true opt-in regression.

Fixes #3196
2026-06-22 18:29:28 +00:00
can1357 a94cadf723 refactor(coding-agent): standardized evaluation kernel and executor architecture
- Extracted common kernel and executor logic into `BaseKernel` and `executor-base` to eliminate duplicated implementations for Julia, Python, and Ruby.
- Migrated shared operational workflows--including session namespacing, environment filtering, and result mapping--to centralized backend helpers.
- Consolidated runtime discovery and resolution logic into a unified `runtime-env` utility module.
- Simplified language-specific modules by delegating subprocess lifecycle, IPC, and configuration management to the newly established base classes.
2026-06-22 07:22:10 +02:00
roboomp 3677e71e9d fix(eval): exposed nested apply=false patches
Branch-mode isolation can capture nested repository changes without creating a root branch. Eval agent() with apply=false previously treated that shape as no captured changes and returned no recoverable nested patch payload after the isolation worktree was removed.

Expose captured nested patches in EvalAgentResult details and copy them onto JS/Python returnHandle nodes (nestedPatches / nested_patches). Document the return_handle escape hatch and add regression coverage for branch-mode nested-only apply=false runs.

Fixes #3196
2026-06-22 04:08:56 +00:00
roboomp a6ee95df08 fix(eval): preserved returnHandle artifacts and nested branch patches
Eval preludes now forward returnHandle to the bridge so no-session eval runs can preserve the temp artifacts backing returned agent:// handles. The bridge keeps those temporary artifact directories whenever returnHandle is requested, including non-isolated runs and successful isolated applies.

Branch-mode isolation now treats nested-only changes as merge-eligible even when no root branch was produced, letting callers apply nested patches instead of dropping them when the root repo had no diff.

Added regression coverage for returnHandle artifact preservation and nested-only branch isolation.

Fixes #3196
2026-06-22 03:59:17 +00:00
roboomp 54f4652573 fix(eval): exposed apply=false artifacts on the agent() return_handle node
When agent() ran with schema and apply=false, the bridge correctly returned the captured patch/branch in details, but the preludes only forwarded id/agent/handle/data on the returnHandle node. Structured workflows had no way to recover the artifact for a manual apply.

Both runtimes now copy isolated, patchPath/branchName, changesApplied, and isolationSummary onto the returnHandle node (snake_case in Python, camelCase in JS), keeping null changesApplied so apply=false stays distinguishable from a successful apply. Updated the workflow notice and the Python agent() docstring to point callers at return_handle as the artifact escape hatch for isolated+apply=false runs. Added prelude tests locking the new node shape in both runtimes.

Fixes #3196
2026-06-21 17:35:26 +00:00
roboomp eeed4b93b6 feat(eval): added isolated/apply/merge options to agent() helper
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.

Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.

Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.

Fixes #3196
2026-06-21 17:20:07 +00:00
can1357 7d28c60c86 refactor: renamed intent field from _i to i
- Renamed the global `INTENT_FIELD` constant from `_i` to `i`.
- Updated documentation strings, type annotations, and test expectations across packages to reflect the new field name.
- Ensured consistent usage of the constant in tool schema construction and intent serialization.
2026-06-19 16:42:33 +02:00
can1357 4f03180ae6 refactor(deps): moved intent field constant to pi-wire
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
2026-06-19 04:58:29 +02:00
can1357 dc915dfef1 feat(coding-agent/eval): tracked and forwarded session id and working directory
- Added `sessionId` and `cwd` tracking to JS evaluator session instances.
- Updated `acquireSession` in the JS context manager to update existing session identity and directory metadata on reuse.
- Forwarded the current working directory from `executePerCall` and `executeOnSession` callers to the underlying Python kernel tasks.
2026-06-18 00:59:56 +02:00
can1357 e4a87fa3ce feat(coding-agent): added matplotlib rendering and image persistence across session reload
- Added Matplotlib figure PNG rendering and display tracking in Python runner to emit PNG output immediately when figures are displayed via display(fig).
- Extended session persistence to externalize oversized image payloads in both content and details.images, enabling tool result images to survive session reload.
- Enhanced session loader to resolve image data payloads and blob references across content and details.images during session reconstruction.
- Added image cache invalidation in TUI image component when image protocol, cell dimensions, or Kitty Unicode placeholder mode changes.
- Added comprehensive test coverage for Matplotlib display, image persistence across reload, and TUI image rendering with protocol and dimension changes.
2026-06-17 13:19:28 +02:00
metaphorics 5d95423515 feat(coding-agent/eval): added return_handle to agent() for DAG handle piping
- Added a return_handle (Python) / returnHandle (JS) option to the eval agent() helper that returns a DAG node dict { text, output, handle, id, agent } instead of bare text, where handle is the spawned agent's recoverable agent:// URI.
- Enabled downstream pipeline/parallel stages to reference a large transcript by handle/output instead of re-inlining it; the default path stays backward compatible (bare text, or the parsed object under schema).
- Documented return_handle and the acyclic DAG-wiring pattern in the eval tool description and added a VM-level prelude regression test for the node shape and the no-details fallback.
2026-06-15 21:54:39 +09:00
can1357 a92d2ce989 feat(coding-agent): removed context argument from eval agent() spawn
Shared background now flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each prompt instead of a context string forwarded into the subagent's system prompt. The JS and Python preludes drop the context kwarg from agent(), the subagent system prompt drops the {{#if context}} block and the conversation-context file pointer, and runEvalAgent no longer writes a per-call conversation context file. AgentSession sheds the now-unused formatCompactContext() helper that supplied the file's body, and ToolSession.getCompactContext is removed alongside it.
2026-06-10 17:49:01 +02:00
can1357 4068bfc304 fix(coding-agent): key python kernel sessions by resolved interpreter 2026-06-10 09:51:45 +02:00
can1357 24cbc93913 fix(eval): thread python.interpreter from session settings and expand ~
Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.

Addresses review feedback on #2204.
2026-06-10 08:26:00 +02:00
danzaio caff395d07 feat(eval): add python interpreter setting 2026-06-10 08:26:00 +02:00
can1357 d7eae06830 fix(coding-agent): fixed eval artifact double-writes and kernel I/O capture
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
2026-06-10 01:27:41 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
can1357 cfd72ffb40 fix(coding-agent): fixed eval helpers treating local:// as plain paths
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
2026-06-08 18:28:06 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
can1357 088fb7fb75 fix(eval): resolved JS/Python resets by awaiting in-flight operations
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
2026-06-08 02:04:59 +02:00