8 Commits

Author SHA1 Message Date
can1357 060f4004e7 feat(coding-agent): refactored eval tool to single-step execution
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
2026-06-23 00:59:58 +02:00
can1357 6bdba2d3c9 test(coding-agent): cover opt-in eval backend schema gating
- Assert the eval tool hides disabled backends from the model-facing wire
  schema (language enum + field descriptions), summary, and description by
  default (rb/jl off), and advertises them once enabled — including the
  enabled-subset case.
- Update the env-flag fallback test for the new rb/jl opt-in defaults and
  extend the env guard to PI_RB/PI_JL so the suite is shell-independent.
- Note the opt-in default and dynamic advertising in the changelog.
2026-06-22 06:22:29 +02:00
can1357 1f3f3cf5d1 feat: added ruby and julia language support to coding-agent
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
2026-06-22 06:13:01 +02:00
can1357 6efe86c07b fix(coding-agent): honored boolean env flag overrides over settings
- Made PI_INTENT_TRACING, PI_AUTO_QA, PI_PY, and PI_JS take precedence when set.
- Fell back to config when the env flag is unset instead of ORing.
- Surfaced PI_PY=0/PI_JS=0 in the disabled-backend error messages.
2026-06-06 16:30:09 +02:00
can1357 5c931ac135 refactor(ai/schema): merged strict-mode into normalize, removed strict-mode.ts
- Moved sanitizeSchemaForStrictMode, enforceStrictSchema, and tryEnforceStrictSchema into normalize.ts.
- Moved sanitizeSchemaForOpenAIResponses/rewriteOneOfToAnyOf into normalize.ts alongside other normalizers.
- Removed strict-mode.ts and its public re-export; adapt.ts now only exposes NO_STRICT and adaptSchemaForStrict.
- Dropped StringEnum helper from strict-mode.ts (already removed from public API).
2026-05-16 19:41:18 +02:00
can1357 84ec8fba49 feat(coding-agent/eval): implemented JSON cell-based eval tool inputs
- Removed the legacy `parseEvalInput` parser module and `eval.lark`, eliminating `*** Cell` stream parsing.
- Replaced eval tool arguments from single `input` strings to ordered `cells` arrays in tool calls and schema.
- Updated execution to resolve language explicitly, map `py` to `python`, and apply timeout/reset defaults.
- Removed backend sniffing and `ABORT_WARNING` suffix handling, then updated docs and tests to the new JSON cells format.
2026-05-16 19:33:44 +02:00
can1357 b0a31a5956 fix: outdated tests 2026-04-30 18:40:26 +02:00
can1357 cf60e6df51 feat(coding-agent): implemented eval framework and replaced python tool
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
2026-04-30 18:08:37 +02:00