Commit Graph

36 Commits

Author SHA1 Message Date
can1357 3ce33d436b feat(catalog): added Gemini Flash model definitions and routing
- Added Gemini 3.7 Flash model definitions and reasoning effort routing configurations.
- Implemented geminiLevelFlashFamily in variant-collapse.ts to manage 3.6+ flash thinking levels.
- Added test coverage for Gemini 3.7 Flash variant collapse and discovery routing.
2026-08-13 19:36:51 +02:00
can1357 b279db1790 test: refactored test suites to eliminate time-based sleeps and polling loops
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
2026-08-13 19:32:22 +02:00
metaphorics 85177a29ec fix(coding-agent): fail loudly when no POSIX bash exists; add stdin-isolation regression test 2026-08-05 10:42:36 +00:00
metaphorics f03581dc22 fix(coding-agent): stop eval shell children inheriting the runner control channel; resolve Git Bash on Windows 2026-08-05 10:24:24 +00:00
Slava Zavadsky b8779dae63 fix(coding-agent/eval): remove per-call model override from agent() (#6438)
Completes the maintainer's removal of per-call model selection from
subagent spawns (9f8aa87dbf removed it from the task tool and the
model-facing agent() docs/prompt, but the eval agent() runtime and all
four preludes still accepted and forwarded a per-call model).

Subagents now always resolve through the selected agent's frontmatter
model and settings, so an explicit model: "default" can no longer
silently route children onto the parent session model.

- agent-bridge: drops "model?" from agentArgsSchema and the request
  forward; adds "+": "delete" so a legacy model argument is stripped
  (same contract as the task wire schemas).
- JS/Python/Ruby/Julia preludes: remove the model parameter from
  agent(); completion()'s tier selector is unchanged.
- docs (tools/eval.md, python-repl.md) updated to the removed surface.

Refs #6438
2026-08-04 08:35:01 -04:00
can1357 b48aeacf25 fix(coding-agent/eval): made cells own every bridge call they start
- Split the Python bridge signal: the raw abort reaches tools (so
  subagents die with the turn) while a shielded signal governs how long
  the host waits, so a cancel can no longer settle a cell on top of a
  still-running isolation merge.
- Mirrored the contract in the JS runtime: abort in-flight tool calls
  immediately, then drain any deferExternalAbort phase before killing
  the worker, and refuse new bridge calls once cancelled.
- Held a finished worker result until the run's tool calls drain. A
  floated agent() previously settled the cell at once, dropping the
  run's abort listener and leaving the subagent running with nothing
  able to cancel it.
- Added regression coverage for all three, each verified to fail
  without its fix.
2026-08-03 19:55:20 +02:00
can1357 c7a3010d7c fix(coding-agent/eval): ensured tool bridge receives unshielded abort signal
- Prevent eval agent bridge calls from inheriting the kernel abort shield.
- Unset maxRuntimeMs in runEvalAgent to properly inherit task.maxRuntimeMs.
2026-08-03 18:57:22 +02:00
can1357 d8881e3fd6 Merge PR #7350: fix(eval): preserve console for ConPTY Python kernels (@roboomp) 2026-08-02 20:53:19 +02:00
can1357 ef852af986 feat(computer): implemented modular desktop backend and script workflows
- Replaced monolithic desktop native bindings and action batching with a modular cross-platform backend structure supporting Wayland, X11, macOS, and Win32.
- Updated the computer tool schema and supervisor to execute persistent JavaScript script runs with timeout clamping and asynchronous tool calling.
- Integrated accessibility (AX) tree snapshotting, node querying, and bounds-based hit testing across platform desktop layers.
- Added native clipboard bindings and updated coding-agent prompts, renderers, and tests to validate script-based computer workflows.
2026-08-02 19:46:25 +02:00
roboomp 3fab985c14 fix(eval): preserved console for conpty python kernels
Combined native HWND and stdio TTY evidence so compiled Windows Terminal launches no longer set CREATE_NO_WINDOW for the Python eval kernel.

Fixes #7343
2026-08-02 05:05:52 +00:00
can1357 ccc9b900cf fix(eval): forked subagent kernel resets away from shared sessions
- Subagents inherit the parent's eval session id, so a child's
  reset: true destroyed the co-owned kernel and every sibling's
  interpreter state mid-session.
- resolveOwnerScopedSessionKey now routes a reset from a non-exclusive
  owner onto a deterministic per-owner fork key: the requester gets a
  fresh private kernel, co-owners keep the shared one, and the fork
  stays sticky for that owner until its teardown reaps it.
- Applied across Python, JavaScript, Ruby, and Julia executors; JS
  contexts gained an owner registry plus disposeVmContextsByOwner,
  wired into EvalRunner.disposeKernels and SDK session teardown.
- Covered by pure key-resolution contracts and an end-to-end JS test:
  co-owner reset forks, shared state survives, fork is sticky, and
  per-owner dispose reaps only the fork.
2026-07-31 20:13:23 +02:00
Larry Gordon 959ad3451f fix(test): gave CLI-spawning tests explicit timeouts
Four tests spawn the full CLI entry graph (or compile a standalone binary)
and declared no timeout, so they inherited Bun's 5s default. Spawning
`src/cli.ts` costs ~900ms warm on a fast machine and ~3.1s cold, so the
budget is spent almost entirely on transpile. When CI runs the native
bucket with OMP_TEST_CONCURRENCY=4, a cold spawn on a contended runner
crosses 5s and the test fails with `timed out after 5000ms` plus a
trailing `killed 1 dangling process` - the subprocess was still alive when
the timeout fired.

Reproduced locally by oversubscribing the box (48 concurrent runs of the
same chunk): 48/48 failed with the identical signature, while 4-way
concurrency - what CI actually configures - passed every time. Only the
subprocess tests starve; the pure-unit tests in the same files pass.

Timeouts are sized to the work, matching existing subprocess tests
(read-cli-mcp-resource 30_000, acp-stdout-hygiene 60_000): 30s for CLI
spawns, 60s for the `bun build --compile` case. After the change, 64
concurrent runs of all three files produce zero timeouts.

`profile-cli.test.ts` gets the timeout on both spawn tests. Only the first
was observed failing, because it warms the transpile cache for its
sibling - that ordering is incidental and would flake if it changed.

These are drift and wiring assertions, not latency assertions, so a
generous ceiling costs nothing on a healthy run.
2026-07-30 11:27:10 -07:00
can1357 14dc5d4fcc fix(coding-agent/eval): bypassed environment proxies for python bridge calls
- Configure the python eval prelude to use a custom urllib opener that ignores environment proxies.
- Add test coverage verifying parallel tool bridge calls succeed when proxy variables are set.
2026-07-30 16:19:48 +02:00
can1357 d16a251777 chore: reorg tests 2026-07-27 16:43:53 +02:00
vmcall d944879f21 feat(task): unified structured subagent execution
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.

Fixes #5279
2026-07-17 17:36:59 +02:00
Colin Mason ffa879ba2c keep process cwd while other cells are live 2026-07-13 09:57:20 -04:00
Colin Mason 9c9da2387d harden js eval subprocess 2026-07-13 09:57:19 -04:00
can1357 70754dfa01 fix(coding-agent): hardened same-realm runtime guards from PR review
- setCwd now updates the saved __omp_session__ stack entry so a deferred
  cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
  a first init during another runtime's live run fails via init-failed
  instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
  setup throw under an already-aborted signal cannot become an unhandled
  rejection (review P2)
- credited #4907 in the changelog entry
2026-07-10 12:33:53 +02:00
ben 71c9cc0324 fix(coding-agent): keep same-realm setCwd from killing TUI sessions
When the JS eval worker falls back to the in-process inline path, concurrent
JsRuntime instances share one realm. setCwd used to throw on exclusive-owner
conflicts, and the microtask delivery path turned that into a fatal
unhandledRejection that postmortem exited on. Stamp local cwd without
stealing the active realm, report init failures over the worker protocol,
and cover process survival with in-process and child-process regressions.
2026-07-09 16:02:23 +08:00
ben 9cb051a4b1 Fix JS worker cwd conflict handling 2026-07-08 18:32:05 +08:00
can1357 24d58033bb refactor(eval): rename agent() params agent_type→agent, return_handle→handle
The eval agent() helper used `agent_type`/`return_handle` (snake_case) in
Python/Ruby/Julia and `agentType`/`returnHandle` (camelCase) in JS, forcing
the prelude docs to repeat every option twice ("JS same but camelcased").
Both are now single lowercase words identical across all four runtimes, and
`agent` matches the `task` tool's existing agent-selection parameter.

- Renamed across py/js/rb/jl preludes (signatures, forwarding, docstrings).
- Renamed the `__agent__` bridge wire protocol + `EvalAgentArgs` (`agentType`
  → `agent`, `returnHandle` → `handle`) so no prelude-side remap is needed.
- Updated prompt docs (workflow-notice.md, tools/eval.md), repo docs
  (docs/tools/eval.md, docs/python-repl.md), and all bridge/prelude tests.
- CHANGELOG: Breaking Changes entry under [Unreleased].
2026-06-22 22:12:24 +02:00
can1357 db58a163b2 fix(coding-agent): isolated JS eval runtime globals across runtime lifetimes
- Made `JsRuntime` track the realm globals it installs (`#globalOwner`, `#ownedGlobalKeys`, `PRELUDE_GLOBAL_KEYS`) and reclaim them in `dispose()`, re-binding ownership through `#activateGlobals()` on cwd/run-scope/run; disposing an older inline/direct runtime no longer deletes a newer runtime's helper globals, and overlapping same-realm runs are serialized via `enterGlobalRun`.
- Added `WorkerCore.dispose()` (rejects pending tool calls with `ToolError` and disposes the runtime) and invoked it from `spawnInlineWorker` teardown in `context-manager.ts`.
- Updated the `js-static-import-rewrite` test to snapshot and restore any pre-existing `__omp_import__` global instead of unconditionally deleting it.
- Added `runtime-global-dispose.test.ts` covering newer-runtime global survival, older-runtime reactivation, and cross-runtime mutation rejection.
- Clarified the `runWorkerEntrypoint` comment in `cli.ts` about `parentPort` sync-prefix buffering for tab/eval workers.
2026-06-17 01:21:33 +02:00
can1357 1fa4f2eee8 fix(subagent-session-spawning): resolved subagent parent id propagation
- Forwarded `parentAgentId` through task and eval launch paths when spawning subagents.
- Mapped `parentAgentId` to `parentId` in `createAgentSession`.
- Passed each caller's session `getAgentId` (or `MAIN_AGENT_ID`) as the parent for spawned agents.
2026-06-16 23:03:54 +02:00
can1357 3f82589ec1 fix: fixed OAuth and profile-boundary regressions across CLI and env handling
- Fixed OAuth credentials to keep unknown fields in schema while preserving existing shape checks.
- Fixed MCP OAuth IDs to be profile-scoped and avoid deleting credentials from non-active profiles.
- Fixed string-flag parsing so PROFILE_BOOTSTRAP_BOUNDARY tokens are not consumed as values.
- Fixed active-profile directory resolution to refresh after env updates so profile .env overrides apply.
2026-06-15 03:20:45 +02:00
can1357 9d99ae1af0 feat(coding-agent): rewrote the task tool to spawn one persistent subagent per call
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
2026-06-10 17:54:47 +02:00
can1357 9d457f73d9 test: migrated test imports to package subpath exports
- Replaced relative `../src` imports with `@oh-my-pi/pi-ai` and `@oh-my-pi/pi-agent-core` subpaths.
2026-06-08 19:03:55 +02:00
can1357 cf621d0abf feat(coding-agent-eval): added runEvalAgent bridge for agent plan checks
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
2026-05-31 06:57:02 +02:00
can1357 5d7a452f11 test(coding-agent/eval): updated eval tests to inject runtime hooks explicitly
- Refactored console-table tests to build explicit RuntimeHooks and pass them to JsRuntime.run.
- Refactored image coercion tests to pass explicit RuntimeHooks into JsRuntime.displayValue instead of constructor hooks.
2026-05-26 15:28:10 +02:00
can1357 7f0208ac83 feat(coding-agent/eval): added console.table bridge to runtime text output
- Added a `console.table` helper in the JS prelude that forwards calls to the runtime `__omp_table__` hook.
- Implemented `__omp_table__` in the runtime to render tables through `node:console.Console` and emit text via `onText`.
- Added tests verifying `console.table` produced formatted table output and respected the optional columns filter.
2026-05-26 13:16:59 +02:00
can1357 226fe87345 fix(coding-agent/eval): hardened image display value coercion to strict base64
- Implemented strict base64 validation and normalization for image `displayValue` payloads in the JS runtime, supporting strict strings, `Uint8Array`, `Buffer`, `ArrayBuffer`, typed-array views, and JSON `Buffer` objects.
- Dropped unrecognized image payloads while emitting a warning and fallback text instead of forwarding malformed data.
- Added tests covering successful coercions and invalid image data rejection paths.
2026-05-19 04:36:30 +02:00
can1357 84ec8fba49 feat(coding-agent/eval): implemented JSON cell-based eval tool inputs
- Removed the legacy `parseEvalInput` parser module and `eval.lark`, eliminating `*** Cell` stream parsing.
- Replaced eval tool arguments from single `input` strings to ordered `cells` arrays in tool calls and schema.
- Updated execution to resolve language explicitly, map `py` to `python`, and apply timeout/reset defaults.
- Removed backend sniffing and `ABORT_WARNING` suffix handling, then updated docs and tests to the new JSON cells format.
2026-05-16 19:33:44 +02:00
can1357 028a442f18 feat(coding-agent): added parser support for t/rst in Cell eval
- Introduced canonical `*** Cell` headers with `t:` and `rst` attributes in eval prompts, schema, and docs.
- Updated parser and grammar to parse `*** Cell` blocks, stop on `*** End`/next header/EOF, and handle invalid `rst` with errors.
- Added quote-aware attribute tokenizers and split HTML eval parsing into `Cell` and legacy `Begin` handlers with `py` defaults.
- Expanded parsing behavior and tests for `rst` booleans, title aliases, abort boundaries, and stray-line skips between cells.
2026-05-12 10:01:56 +02:00
can1357 c562f4dc2a fix(coding-agent/eval): hardened eval parser against stray non-marker lines
- Updated parseEvalInput to verify begin-cell markers before dereferencing regex matches.
- Skipped stray non-marker lines between and after cells, preserving valid cells when model output is noisy.
- Added eval parse regression tests for stray content and trailing chatter, and kept abort-line handling explicit.
2026-05-12 09:41:06 +02:00
can1357 8b92ec937e feat: gpt-5 harmony errata fixes
- Replaced `===== ... =====` eval cell headers with `*** Begin ` / `*** End ` markers; legacy format remains renderable in HTML exports.
- Replaced hashline patch grammar with `*** Begin Patch` / `*** End Patch` envelope; old inputs without the envelope are still accepted.
- Extracted `sniffEvalLanguage` into a shared `sniff.ts` module reused by the parser and tool.
- Added `docs/ERRATA-GPT5-HARMONY.md` and `scripts/session-stats/harmony_backtest.py` documenting and backtesting the GPT-5 Harmony-header leak defect.
2026-05-10 19:52:44 +02:00
can1357 d064170c56 fix: chatgpt really doesnt like sandwitching lark 2026-05-02 04:54:56 +02:00
can1357 cf60e6df51 feat(coding-agent): implemented eval framework and replaced python tool
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
2026-04-30 18:08:37 +02:00