- Simplified `bash` guidance to tighten allowed command patterns, pipeline limits, and launch-based process handling.
- Reworked `browser` instructions into grouped helper sections while preserving selector restrictions and key action semantics.
- Harmonized `eval`, `irc`, `read`, and `todo` prompt wording around state reuse, messaging, selector formats, and task operations.
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.
Fixes#5250
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
- Assert the eval tool hides disabled backends from the model-facing wire
schema (language enum + field descriptions), summary, and description by
default (rb/jl off), and advertises them once enabled — including the
enabled-subset case.
- Update the env-flag fallback test for the new rb/jl opt-in defaults and
extend the env guard to PI_RB/PI_JL so the suite is shell-independent.
- Note the opt-in default and dynamic advertising in the changelog.
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.