Commit Graph

11 Commits

Author SHA1 Message Date
can1357 9d457f73d9 test: migrated test imports to package subpath exports
- Replaced relative `../src` imports with `@oh-my-pi/pi-ai` and `@oh-my-pi/pi-agent-core` subpaths.
2026-06-08 19:03:55 +02:00
can1357 cf621d0abf feat(coding-agent-eval): added runEvalAgent bridge for agent plan checks
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
2026-05-31 06:57:02 +02:00
can1357 5d7a452f11 test(coding-agent/eval): updated eval tests to inject runtime hooks explicitly
- Refactored console-table tests to build explicit RuntimeHooks and pass them to JsRuntime.run.
- Refactored image coercion tests to pass explicit RuntimeHooks into JsRuntime.displayValue instead of constructor hooks.
2026-05-26 15:28:10 +02:00
can1357 7f0208ac83 feat(coding-agent/eval): added console.table bridge to runtime text output
- Added a `console.table` helper in the JS prelude that forwards calls to the runtime `__omp_table__` hook.
- Implemented `__omp_table__` in the runtime to render tables through `node:console.Console` and emit text via `onText`.
- Added tests verifying `console.table` produced formatted table output and respected the optional columns filter.
2026-05-26 13:16:59 +02:00
can1357 226fe87345 fix(coding-agent/eval): hardened image display value coercion to strict base64
- Implemented strict base64 validation and normalization for image `displayValue` payloads in the JS runtime, supporting strict strings, `Uint8Array`, `Buffer`, `ArrayBuffer`, typed-array views, and JSON `Buffer` objects.
- Dropped unrecognized image payloads while emitting a warning and fallback text instead of forwarding malformed data.
- Added tests covering successful coercions and invalid image data rejection paths.
2026-05-19 04:36:30 +02:00
can1357 84ec8fba49 feat(coding-agent/eval): implemented JSON cell-based eval tool inputs
- Removed the legacy `parseEvalInput` parser module and `eval.lark`, eliminating `*** Cell` stream parsing.
- Replaced eval tool arguments from single `input` strings to ordered `cells` arrays in tool calls and schema.
- Updated execution to resolve language explicitly, map `py` to `python`, and apply timeout/reset defaults.
- Removed backend sniffing and `ABORT_WARNING` suffix handling, then updated docs and tests to the new JSON cells format.
2026-05-16 19:33:44 +02:00
can1357 028a442f18 feat(coding-agent): added parser support for t/rst in Cell eval
- Introduced canonical `*** Cell` headers with `t:` and `rst` attributes in eval prompts, schema, and docs.
- Updated parser and grammar to parse `*** Cell` blocks, stop on `*** End`/next header/EOF, and handle invalid `rst` with errors.
- Added quote-aware attribute tokenizers and split HTML eval parsing into `Cell` and legacy `Begin` handlers with `py` defaults.
- Expanded parsing behavior and tests for `rst` booleans, title aliases, abort boundaries, and stray-line skips between cells.
2026-05-12 10:01:56 +02:00
can1357 c562f4dc2a fix(coding-agent/eval): hardened eval parser against stray non-marker lines
- Updated parseEvalInput to verify begin-cell markers before dereferencing regex matches.
- Skipped stray non-marker lines between and after cells, preserving valid cells when model output is noisy.
- Added eval parse regression tests for stray content and trailing chatter, and kept abort-line handling explicit.
2026-05-12 09:41:06 +02:00
can1357 8b92ec937e feat: gpt-5 harmony errata fixes
- Replaced `===== ... =====` eval cell headers with `*** Begin ` / `*** End ` markers; legacy format remains renderable in HTML exports.
- Replaced hashline patch grammar with `*** Begin Patch` / `*** End Patch` envelope; old inputs without the envelope are still accepted.
- Extracted `sniffEvalLanguage` into a shared `sniff.ts` module reused by the parser and tool.
- Added `docs/ERRATA-GPT5-HARMONY.md` and `scripts/session-stats/harmony_backtest.py` documenting and backtesting the GPT-5 Harmony-header leak defect.
2026-05-10 19:52:44 +02:00
can1357 d064170c56 fix: chatgpt really doesnt like sandwitching lark 2026-05-02 04:54:56 +02:00
can1357 cf60e6df51 feat(coding-agent): implemented eval framework and replaced python tool
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
2026-04-30 18:08:37 +02:00