Commit Graph

5 Commits

Author SHA1 Message Date
can1357 af1832af1b feat(coding-agent/prompts): refined tool prompts for shell, browser, and eval workflows
- Simplified `bash` guidance to tighten allowed command patterns, pipeline limits, and launch-based process handling.
- Reworked `browser` instructions into grouped helper sections while preserving selector restrictions and key action semantics.
- Harmonized `eval`, `irc`, `read`, and `todo` prompt wording around state reuse, messaging, selector formats, and task operations.
2026-07-15 03:28:21 +02:00
roboomp 8e6d26b1e8 fix(eval): honored unlimited cell timeouts
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.

Fixes #5250
2026-07-14 17:33:56 +00:00
can1357 060f4004e7 feat(coding-agent): refactored eval tool to single-step execution
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
2026-06-23 00:59:58 +02:00
can1357 6bdba2d3c9 test(coding-agent): cover opt-in eval backend schema gating
- Assert the eval tool hides disabled backends from the model-facing wire
  schema (language enum + field descriptions), summary, and description by
  default (rb/jl off), and advertises them once enabled — including the
  enabled-subset case.
- Update the env-flag fallback test for the new rb/jl opt-in defaults and
  extend the env guard to PI_RB/PI_JL so the suite is shell-independent.
- Note the opt-in default and dynamic advertising in the changelog.
2026-06-22 06:22:29 +02:00
can1357 49aa6e5839 fix(eval): corrected eval LLM calls and spawn-aware tool descriptions
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.
2026-06-08 02:04:59 +02:00